datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
criteogoodreads-projectRetailRocket-Recommender-Datafashion-recommender-dataMIND
MIND
Microsoft News Dataset (MIND) is a large-scale dataset for news recommendation research. It was collected from anonymized behavior logs of Microsoft News website. The mission of MIND is to serve as a benchmark dataset for news recommendation and facilitate the research in news recommendation and recommender systems area.
MIND contains about 160k English news articles and more than 15 million impression logs generated by 1 million users. Every news article contains rich textual… See the full description on the dataset page: https://huggingface.co/datasets/Recommenders/MIND.MovieLens
MovieLens
To acknowledge use of the dataset in publications, please cite the following paper:
F. Maxwell Harper and Joseph A. Konstan. 2015. The MovieLens Datasets: History and Context. ACM Transactions on Interactive Intelligent Systems (TiiS) 5, 4: 19:1–19:19. https://doi.org/10.1145/2827872
BOOK-RECOMMENDER-DATASET
Book Recommender Dataset
CSV exports from my Book Recommender pipeline. Includes cleaned metadata, category labels, emotion tags, and a tagged description file.
Files
books_cleaned.csv: Core cleaned book metadata.
books_with_categories.csv: Adds multi-label categories column.
books_with_emotions.csv: Adds emotion_* columns (one-hot or scores).
tagged_description.txt: Preprocessed descriptions (one per line, or TSV).
Column Schema (example)
book_id (str)… See the full description on the dataset page: https://huggingface.co/datasets/svastikkka/BOOK-RECOMMENDER-DATASET.tomplay-classic-recommender
Tomplay - Processed for Classic Recommenders
Dataset Description
This is the processed version of the tomplay dataset, specifically prepared for classic recommendation algorithms like SVD (Singular Value Decomposition) and NMF (Non-negative Matrix Factorization).
Processing Pipeline
The original dataset has been processed with the following steps:
Data Cleaning: Removed invalid entries and outliers
ID Mapping: Created sequential user and item IDs starting from… See the full description on the dataset page: https://huggingface.co/datasets/OloriBern/tomplay-classic-recommender.movie-recommender-artifacts
🎬 Two-Tower + FAISS + XGBoost Recommender — Inference Artifacts
This dataset contains the serialized inference artifacts for the Two-Stage Movie Recommendation System built using TensorFlow Recommenders, FAISS, and XGBoost.
These artifacts support low-latency retrieval and ranking for the deployed Hugging Face Space:
🔗 Space: two-tower-faiss-xgb-recommender
📌 Overview
The recommendation system follows a two-stage architecture:
Retrieval (Stage 1)… See the full description on the dataset page: https://huggingface.co/datasets/dmckinney-ml/movie-recommender-artifacts.als-recommendermovielens-small-completedThis dataset are combining by:
The Movielens latest dataset: https://grouplens.org/datasets/movielens/latest/
The Movies dataset: https://www.kaggle.com/datasets/rounakbanik/the-movies-dataset
The Tag Genome dataset for movies: https://grouplens.org/datasets/movielens/tag-genome-2021/
The Tag Genome dataset for books: https://grouplens.org/datasets/book-genome
References
[Kotkov et al., 2021] Kotkov, D., Maslov, A., and Neovius, M. (2021). Revisiting the tag relevance… See the full description on the dataset page: https://huggingface.co/datasets/recommender-system/movielens-small-completed.movie-recommender-datasetbook-recommender-artifactsassistments-classic-recommender
assistments
Processed dataset for the LLM as Recommender project.
visual-product-recommender-datafashion-recommender-imagesdrugbank-recommender-datasteam-review-and-bundle-dataset
About
This dataset is copying from https://cseweb.ucsd.edu/~jmcauley/datasets.html in Steam Video Game and Bundle Data.
We are using it for our research. These datasets contain reviews from the Steam video game platform, and information about which games were bundled together.
Statistics
Reviews : 7,793,069
Users Write Reviews : 25,799
Users Has Games : 88,310
Steam Games : 32,135
Steam Bundles : 615
Samples
For… See the full description on the dataset page: https://huggingface.co/datasets/recommender-system/steam-review-and-bundle-dataset.book-recommender-dataBOOK-RECOMMENDER-DATASET
Book Recommender Dataset
CSV exports from my Book Recommender pipeline. Includes cleaned metadata, category labels, emotion tags, and a tagged description file.
Files
books_cleaned.csv: Core cleaned book metadata.
books_with_categories.csv: Adds multi-label categories column.
books_with_emotions.csv: Adds emotion_* columns (one-hot or scores).
tagged_description.txt: Preprocessed descriptions (one per line, or TSV).
Column Schema (example)… See the full description on the dataset page: https://huggingface.co/datasets/orbitk/BOOK-RECOMMENDER-DATASET.movielens-32m-sequential-recommender
MovieLens 32M Sequential Recommender Dataset
This dataset is a processed version of the MovieLens 32M dataset, specifically formatted for sequential recommendation tasks. It contains user-item interaction sequences, enriched with rating and timestamp information, split into training, validation, and test sets.
Dataset Structure
The dataset is provided as a DatasetDict with three splits: train, validation, and test. Each split contains:
input_sequence: A string… See the full description on the dataset page: https://huggingface.co/datasets/krishnakamath/movielens-32m-sequential-recommender.letterboxd-recommender-datasetRecommender-System-Datasetbundle-instacart-datasetbook-recommender-dataset
Book Recommender Dataset
CSV exports from my Book Recommender pipeline. Includes cleaned metadata, category labels, emotion tags, and a tagged description file.
Files
books_cleaned.csv: Core cleaned book metadata.
books_with_categories.csv: Adds multi-label categories column.
books_with_emotions.csv: Adds emotion_* columns (one-hot or scores).
tagged_description.txt: Preprocessed descriptions (one per line, or TSV).
Column Schema (example)
book_id (str)… See the full description on the dataset page: https://huggingface.co/datasets/swayista/book-recommender-dataset.movie-recommender-datamovie-recommender-datamusic-recommender-datamovie-recommender-filesdiscogs-recommender-model
