Facebook
TwitterThese datasets include ratings as well as social (or trust) relationships between users. Data are from LibraryThing (a book review website) and epinions (general consumer reviews).
Metadata includes
reviews
price paid (epinions)
helpfulness votes (librarything)
flags (librarything)
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset provides a comprehensive collection of features for building a content-based recommendation system in an ecommerce environment. Content filtering, which relies on users' interests and past activities, is a prevalent method for suggesting products tailored to individual preferences.
Each entry in the dataset represents a product along with various attributes that can be leveraged for recommendation purposes. Here's an overview of the features included:
1) Number of clicks on similar products: Indicates the popularity or engagement level of similar items. 2) Number of similar products purchased so far: Reflects the conversion rate of similar products. 3) Average rating given to similar products: Offers insight into the perceived quality of comparable items. 4) Gender: Allows for gender-specific recommendations. 5) Median purchasing price (in rupees): Provides pricing information for segmentation or pricing strategy analysis. 6) Rating of the product: The rating of the product itself, indicating its overall quality. 7)**Brand of the product**: Brand loyalty or preference can influence recommendations. 8) Customer review sentiment score (overall): Sentiment analysis of customer reviews, indicating overall satisfaction. 9) Price of the product: The actual price of the product. 10) Holiday: Seasonal or holiday-specific buying patterns. 11) Season: Seasonal preferences may influence product choices. 12) Geographical locations: Regional preferences or availability may impact recommendations. 13) Probability for the product to be recommended to the person: The likelihood of recommending the product to a specific user based on their profile and past behavior.
With this rich set of features, businesses can implement sophisticated recommendation algorithms to personalize the shopping experience for users, ultimately leading to increased customer satisfaction, engagement, and sales.
Facebook
TwitterThis dataset was created by Raihan Sikdar
Facebook
TwitterThis is a common Zenodo repository for both lastfm-360K and lastfm-1K datasets. See below the details of both datasets, including license, acknowledgements, contact, and instructions to cite.
LASTFM-360K (version 1.2, March 2010).
user-mboxsha1 \t musicbrainz-artist-id \t artist-name \t plays
user-mboxsha1 \t gender (m|f|empty) \t age (int|empty) \t country (str|empty) \t signup (date|empty)
000063d3fe1cf2ba248b9e3c3f0334845a27a6be \t a3cb23fc-acd3-4ce0-8f36-1e5aa6a18432 \t u2 \t 31 ...
000063d3fe1cf2ba248b9e3c3f0334845a27a6be \t m \t 19 \t Mexico \t Apr 28, 2008 ...
LASTFM-1K (version 1.0, March 2010).
userid \t timestamp \t musicbrainz-artist-id \t artist-name \t musicbrainz-track-id \t track-name
userid \t gender ('m'|'f'|empty) \t age (int|empty) \t country (str|empty) \t signup (date|empty)
user_000639 \t 2009-04-08T01:57:47Z \t MBID \t The Dogs D'Amour \t MBID \t Fall in Love Again? user_000639 \t 2009-04-08T01:53:56Z \t MBID \t The Dogs D'Amour \t MBID \t Wait Until I'm Dead ...
user_000639 \t m \t Mexico \t Apr 27, 2005 ...
LICENSE OF BOTH DATASETS. The data contained in both datasets is distributed with permission of Last.fm. The data is made available for non-commercial use. Those interested in using the data or web services in a commercial context should contact:
partners [at] last [dot] fm
For more information see Last.fm terms of service
ACKNOWLEDGEMENTS. Thanks to Last.fm for providing the access to this data via their web services. Special thanks to Norman Casagrande.
REFERENCES. When using this dataset you must reference the Last.fm webpage. Optionally (not mandatory at all!), you can cite Chapter 3 of this book:
@book{Celma:Springer2010,
author = {Celma, O.},
title = {{Music Recommendation and Discovery in the Long Tail}},
publisher = {Springer},
year = {2010}
}
CONTACT: This data was collected by Òscar Celma @ MTG/UPF
Facebook
Twitterthuychang404/job-recommendation-system dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
🎬 Movie Recommendation Dataset 📖 Overview
This dataset is created to help researchers, students, and practitioners build movie recommendation systems. It contains 20,000 rows of realistic, user-movie interaction data, blending popular movie titles, genres, and user behavior metrics.
It is ideal for experimenting with collaborative filtering, content-based filtering, hybrid approaches, and deep learning-based recommenders.
📂 Dataset Structure
The dataset has the following columns:
movie_title 🎥 The name of the movie (e.g., The Shawshank Redemption, Inception, Titanic). Titles are based on realistic movie names from popular cinema.
user_id 👤 A unique identifier for each user (e.g., User_1023). Simulates multiple users interacting with various movies.
genres 🎭 One or more genres for each movie (e.g., Action, Comedy, Drama). Useful for content-based recommendation.
watch_time ⏱️ Time spent (in minutes) by the user watching the movie. Higher watch time can imply greater engagement.
imdb_score ⭐ A score between 1.0 and 10.0, inspired by IMDb ratings. Acts as a measure of quality or popularity.
movie_likes ❤️ The number of likes a movie has (0–10,000). Works as an implicit popularity measure.
📊 Use Cases
Building Recommendation Systems (Collaborative Filtering, Content-based, Hybrid, Deep Learning models)
Exploratory Data Analysis (EDA) on user preferences
Data visualization (genre distribution, rating patterns, engagement analysis)
Practicing data cleaning, preprocessing, and feature engineering
Testing machine learning algorithms for personalization tasks
🌍 Source
The dataset is synthetically generated but uses realistic movie titles and attributes.
Movie names are inspired by popular films on IMDb and general cinema culture.
User interactions (watch time, likes, ratings) were simulated with statistical distributions to mimic real-world behavior.
⚠️ Disclaimer: This dataset was not scraped from IMDb, Netflix, or any external service. It was created for educational and research purposes only.
📌 Column Summary Column Type Description movie_title String The movie name (popular real-world titles included) user_id String Unique identifier for each user (e.g., User_1234) genres String One or more genres, comma-separated watch_time Integer Time in minutes user watched the movie (10–180) imdb_score Float IMDb-like score between 1.0 and 10.0 movie_likes Integer Number of likes the movie received (0–10,000) 🙌 Acknowledgements
This dataset was created for the data science and ML community to:
Build and test recommendation systems
Explore movie analytics
Practice ML modeling on structured data
If you use this dataset, a citation or link back is appreciated. ❤️
✨ Happy analyzing & recommending movies! ✨
Facebook
TwitterAttribution-NonCommercial-NoDerivs 4.0 (CC BY-NC-ND 4.0)https://creativecommons.org/licenses/by-nc-nd/4.0/
License information was derived automatically
Wyze Rule Recommendation Dataset
Dataset Summary
The Wyze Rule dataset is a new large-scale dataset designed specifically for smart home rule recommendation research. It contains over 1 million rules generated by 300,000 users from Wyze Labs, offering an extensive collection of real-world automation rules tailored to users' unique smart home setups. The goal of the Wyze Rule dataset is to advance research and development of personalized rule recommendation systems for… See the full description on the dataset page: https://huggingface.co/datasets/wyzelabs/RuleRecommendation.
Facebook
TwitterThe CDEI has been tasked with researching the ways in which algorithmically driven recommendation systems have impacted music consumption, including how creators are being affected (see Recommendation 18 in the government’s response to the economics of music streaming Committee’s Second Report). The CDEI will be carrying out a survey to take the views of creators into consideration as part of our research, as well as begin to understand if and how algorithmically driven recommendation systems affect different categories of creators, creators across different genres, and whether there are any apparent differences in their effect by region, age, gender identity, or ethnic group. This privacy notice explains who the CDEI are, the personal data the CDEI collects, how the CDEI uses it, who the CDEI shares it with, and what your legal rights are.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
What does this dataset contain?
This dataset contains over 700 million time-stamped listening events collected from 3.4M anonymised users on the music streaming service Deezer, occurred between March and August 2022. It includes 50k anonymised songs, among the most popular ones on the service as well as their pre-trained embedding vectors, calculated by our internal model. All files are in parquet format which could be read by using pandas.read_parquet function.
What could this dataset be used for?
This dataset could be used for collaborative filtering as well as sequential recommendation (including both next-item and next-session recommendations).
Citation
If you use this dataset, please cite following paper:
@inproceedings{tran-recsys2024, title={Transformers Meet ACT-R: Repeat-Aware and Sequential Listening Session Recommendation}, author={Viet-Anh Tran, Guillaume Salha-Galvan, Bruno Sguerra and Romain Hennequin}, booktitle = {Proceedings of the 18th ACM Conference on Recommender Systems}, year = {2024} }
Facebook
Twitterhttps://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy
Content-Based Recommendation System market was valued at $5.8 billion in 2025, projected to reach $18.6 billion by 2034, growing at 13.8% CAGR. Comprehensive analysis of software, services, e-commerce, media, and healthcare applications.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This is a synthetic dataset containing 5,000 rows for career path recommendation based on education level, specialization, skills, certifications, and CGPA.
Facebook
TwitterDuring a March 2024 survey among adults in the United States, around ** percent of respondents selected friends and family as a trustworthy product recommendation source. Recommendations from an expert reviewer followed with a share of ** percent, whereas artificial intelligence (AI) applications such as ChatGPT and Bard ranked fifth, chosen by less than ** percent of the interviewees.
Facebook
TwitterThe confluence of Search and Recommendation (S&R) services is a vital aspect of online content platforms like Kuaishou and TikTok. The integration of S&R modeling is a highly intuitive approach adopted by industry practitioners. However, there is a noticeable lack of research conducted in this area within the academia, primarily due to the absence of publicly available datasets. Consequently, a substantial gap has emerged between academia and industry regarding research endeavors in this field. To bridge this gap, we introduce the first large-scale, real-world dataset KuaiSAR of integrated Search And Recommendation behaviors collected from Kuaishou, a leading short-video app in China with over 300 million daily active users. Previous research in this field has predominantly employed publicly available datasets that are semi-synthetic and simulated, with artificially fabricated search behaviors. Distinct from previous datasets, KuaiSAR records genuine user behaviors, the occurrence of each interaction within either search or recommendation service, and the users’ transitions between the two services. This work aids in joint modeling of S&R, and the utilization of search data for recommenders (and recommendation data for search engines). Additionally, due to the diverse feedback labels of user-video interactions, KuaiSAR also supports a wide range of other tasks, including intent recommendation, multi-task learning, and long sequential multi-behavior modeling etc. We believe this dataset will facilitate innovative research and enrich our understanding of S&R services integration in real-world applications.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Research Domain/Project:
This dataset is part of the Tour Recommendation System project, which focuses on predicting user preferences and ratings for various tourist places and events. It belongs to the field of Machine Learning, specifically applied to Recommender Systems and Predictive Analytics.
Purpose:
The dataset serves as the training and evaluation data for a Decision Tree Regressor model, which predicts ratings (from 1-5) for different tourist destinations based on user preferences. The model can be used to recommend places or events to users based on their predicted ratings.
Creation Methodology:
The dataset was originally collected from a tourism platform where users rated various tourist places and events. The data was preprocessed to remove missing or invalid entries (such as #NAME? in rating columns). It was then split into subsets for training, validation, and testing the model.
Structure of the Dataset:
The dataset is stored as a CSV file (user_ratings_dataset.csv) and contains the following columns:
place_or_event_id: Unique identifier for each tourist place or event.
rating: Rating given by the user, ranging from 1 to 5.
The data is split into three subsets:
Training Set: 80% of the dataset used to train the model.
Validation Set: A small portion used for hyperparameter tuning.
Test Set: 20% used to evaluate model performance.
Folder and File Naming Conventions:
The dataset files are stored in the following structure:
user_ratings_dataset.csv: The original dataset file containing user ratings.
tour_recommendation_model.pkl: The saved model after training.
actual_vs_predicted_chart.png: A chart comparing actual and predicted ratings.
Software Requirements:
To open and work with this dataset, the following software and libraries are required:
Python 3.x
Pandas for data manipulation
Scikit-learn for training and evaluating machine learning models
Matplotlib for chart generation
Joblib for saving and loading the trained model
The dataset can be opened and processed using any Python environment that supports these libraries.
Additional Resources:
The model training code, README file, and performance chart are available in the project repository.
For detailed explanation and code, please refer to the GitHub repository (or any other relevant link for the code).
Dataset Reusability:
The dataset is structured for easy use in training machine learning models for recommendation systems. Researchers and practitioners can utilize it to:
Train other types of models (e.g., regression, classification).
Experiment with different features or add more metadata to enrich the dataset.
Data Integrity:
The dataset has been cleaned and preprocessed to remove invalid values (such as #NAME? or missing ratings). However, users should ensure they understand the structure and the preprocessing steps taken before reusing it.
Licensing:
The dataset is provided under the CC BY 4.0 license, which allows free usage, distribution, and modification, provided that proper attribution is given.
Facebook
TwitterFriends and family were the most trusted source of product recommendations according to consumers surveyed in the United States as of March 2024. Nearly ** percent of respondents mentioned the source. Around **** out of 10 customers strongly or somewhat agreed that they trust product recommendations from AI applications, such as ChatGPT and Bard.
Facebook
Twitterhttps://bottleneckcalculator.to/en/terms/https://bottleneckcalculator.to/en/terms/
Upgrade actions and priorities derived from the bottleneck model for Ryzen 7 7800X3D with RTX 4070 Super.
Facebook
Twitterhttps://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy
The Recommendation Engine market was valued at $3.2 billion in 2025 and is projected to reach $18.5 billion by 2034, growing at 21.8% CAGR.
Facebook
Twitterhttps://bottleneckcalculator.to/en/terms/https://bottleneckcalculator.to/en/terms/
Upgrade actions and priorities derived from the bottleneck model for Core i5-12400F with RX 7800 XT.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This dataset is for a Crop Recommendation Machine Learning Model This dataset contains the following Data Fields - • N - ratio of Nitrogen content in soil • P - ratio of Phosphorous content in soil • K - ratio of Potassium content in soil • temperature - temperature in degree Celsius • humidity - relative humidity in % • ph - ph value of the soil • rainfall - rainfall in mm
Facebook
Twitterhttps://mockdatafaker.com/https://mockdatafaker.com/
Free basket dataset for recommendation-system practice — co-purchase patterns for item-to-item and collaborative methods. Reproducible by seed.
Facebook
TwitterThese datasets include ratings as well as social (or trust) relationships between users. Data are from LibraryThing (a book review website) and epinions (general consumer reviews).
Metadata includes
reviews
price paid (epinions)
helpfulness votes (librarything)
flags (librarything)