datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recsys-papers-2025-2026
📚 Recommender Systems Papers 2025–2026
A curated library of 3,151 recent recommender-systems papers spanning 2025-01-02 → 2026-09-17, each with the original PDF and a structured, section-by-section Markdown analysis (research problem, prior work, method, math, experiments, strengths & weaknesses, …). Includes a self-contained Apple-style HTML browser (index.html).
🔑 Browse by meeting (Data Viewer subsets)
The Dataset Viewer above has a subset dropdown keyed by… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-papers-2025-2026.REC_SYScross-domain-recsys-interactions
Cross-Domain RecSys Interactions — 15 sources, integer-indexed
One unified, integer-indexed interaction corpus glued from 15 recommendation sources for
foundation / sequential recommenders. Built in 3 additive versions that share one schema and
one global index space, so reading all of them together = the full corpus, collision-free.
Interactions only — no model, no embeddings. Item text is provided separately as title+description
metadata, joinable 1:1 by item_idx.
Full corpus: 882… See the full description on the dataset page: https://huggingface.co/datasets/TOPAPEC/cross-domain-recsys-interactions.recsys-genrec-dataset-final
SIDReasoner final training dataset
Consolidated training data for the Video Games, Office Products, and Industrial
and Scientific domains.
Included data
Video_Games_catalog, Video_Games_reasoning, and Video_Games_seqrec are
sourced from
yufan/recsys-genrec-dataset-refresh-gpt5.4-candidateV2.
Video_Games_catalog.retrieval_summary contains compact GPT-5.4-generated
product summaries for semantic retrieval while preserving every original
catalog field.… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-genrec-dataset-final.CityPOIsrecsys-genrec-checkpoints-finalRecSys_ACM_2026_Hallucinated
RecSys 2026 ACM challenge, Team: Hallucinated
Step 1: Download the full challenge datasets to obtail the following structure
data/talkpl-ai/TalkPlayData-Challenge-Blind-A
data/talkpl-ai/TalkPlayData-Challenge-Blind-B
data/talkpl-ai/TalkPlayData-Challenge-Dataset
data/talkpl-ai/TalkPlayData-Challenge-Track-Embeddings
data/talkpl-ai/TalkPlayData-Challenge-Track-Metadata # Both all_tracks and test_tracks
data/talkpl-ai/TalkPlayData-Challenge-User-Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/NicoloLocatelli/RecSys_ACM_2026_Hallucinated.recsys-genrec-dataset
🧠 Amazon Semantic-ID Recommendation + Reasoning
Data for the three-stage SIDReasoner pipeline · Reasoning over Semantic IDs Enhances Generative Recommendation
Three Amazon-review categories (5-core, 2016-10 ~ 2018-11), each an independent item
universe with its own Semantic-ID codebook <a_x><b_y><c_z>. Every item maps to a
3-token semantic ID; the model learns to reason over these IDs before recommending.
<cat> below is any of Video_Games, Office_Products… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-genrec-dataset.seq-recsys-simrecsys_slates_dataset
FINN.no Slate Dataset for Recommender Systems
Data and helper functions for FINN.no slate dataset containing both viewed items and clicks from the FINN.no second hand marketplace.
Note: The dataset is originally hosted at https://github.com/finn-no/recsys_slates_dataset and this is a copy of the readme until this repo is properly created "huggingface-style".
We release the FINN.no slate dataset to improve recommender systems research.
The dataset includes both search and… See the full description on the dataset page: https://huggingface.co/datasets/simeneide/recsys_slates_dataset.music-crs-recsys2026
Music-CRS — RecSys 2026 Challenge (Team volart)
A conversational music recommender for the RecSys Challenge 2026 — Music
Conversational Recommendation System.
Given a conversation, it predicts an ordered list of 20 catalog tracks and a
natural-language reply.
This repository reproduces our final Blind-B submission. It is a clean, 3-stage
retrieve → rerank → respond pipeline. The only trained component is a small
LightGBM reranker, trained solely on the provided challenge train… See the full description on the dataset page: https://huggingface.co/datasets/artvolgin/music-crs-recsys2026.amazon-recsys-dataset
Amazon RecSys Dataset — Tools & Home Improvement
Processed dataset derived from the Amazon Reviews 2023 corpus (Tools & Home
Improvement category), used to train a two-stage two-tower recommender system
with cross-encoder re-ranking.
Dataset splits
Split
Rows
Users
Description
train
4,436,875
625,140
Historical interactions per user
val
625,140
625,140
Second-to-last interaction per user
test
625,140
625,140
Last interaction per user
Split… See the full description on the dataset page: https://huggingface.co/datasets/chaturg/amazon-recsys-dataset.shikimori-recsys
Dataset Information
Shikimori RecSys is a recommendation-oriented dataset built from publicly available data on shikimori.io. It is designed specifically for recommender systems, machine learning, and data science tasks. It combines rich anime metadata, genres, and real user rating behavior.
Use Cases:
Anime recommendation systems
User preference modeling
Rating prediction
Educational ML projects
Research & benchmarking
Data Collection & Parsing:
All data was collected using a… See the full description on the dataset page: https://huggingface.co/datasets/kdduha/shikimori-recsys.recsys-test
loader.py
Dataset Summary
A industrial dataset with image audio modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: auto ml
Augmentation: randaugment
Splits & Sampling
Split strategy: random 90 10
Sampling: random
Quality & Labeling
Quality filtering: adaptive
Labeling: semi auto
Files
loader.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/rina-sari/recsys-test.otto-recsysrecsys-practice
preprocess.py
Dataset Summary
A weather dataset with text tabular modality, stored in webdataset format.
Preprocessing & Augmentation
Preprocessing: standard
Augmentation: none
Splits & Sampling
Split strategy: leave one out
Sampling: stratified
Quality & Labeling
Quality filtering: moderate
Labeling: weak supervision
Files
preprocess.py — main artifact of this repository
License… See the full description on the dataset page: https://huggingface.co/datasets/ARTHURDRODRIGUES/recsys-practice.resnet-recsys
build_dataset.py
Dataset Summary
A charts dataset with audio video modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: progressive
Augmentation: randaugment
Splits & Sampling
Split strategy: stratified 90 10
Sampling: balanced
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
build_dataset.py — main artifact of this repository… See the full description on the dataset page: https://huggingface.co/datasets/garciajohn/resnet-recsys.mlp-recsys
dataset.py
Dataset Summary
A math dataset with image text modality, stored in csv format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: random 90 10
Sampling: hard negative
Quality & Labeling
Quality filtering: moderate
Labeling: weak supervision
Files
dataset.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/jorgeherna/mlp-recsys.seq-recsys-simrecsys
loader.py
Dataset Summary
A nlp qa dataset with audio text modality, stored in hdf5 format.
Preprocessing & Augmentation
Preprocessing: standard
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: random 90 10
Sampling: curriculum
Quality & Labeling
Quality filtering: lenient
Labeling: manual
Files
loader.py — main artifact of this repository
License
See the license field… See the full description on the dataset page: https://huggingface.co/datasets/sevramirez/recsys.Million_Song_Datasetbook-recsysdatasets for Book RecSys team project
recsys-mini43
dataloader.py
Dataset Summary
A food dataset with multimodal3 modality, stored in parquet format.
Preprocessing & Augmentation
Preprocessing: domain specific
Augmentation: autoaugment
Splits & Sampling
Split strategy: kfold 5
Sampling: active
Quality & Labeling
Quality filtering: moderate
Labeling: manual
Files
dataloader.py — main artifact of this repository
License
See the license… See the full description on the dataset page: https://huggingface.co/datasets/romanosimone/recsys-mini43.RecSys_largerecsys-proto
loader.py
Dataset Summary
A ocean dataset with text tabular modality, stored in jsonl format.
Preprocessing & Augmentation
Preprocessing: domain specific
Augmentation: none
Splits & Sampling
Split strategy: stratified 90 10
Sampling: active
Quality & Labeling
Quality filtering: strict
Labeling: semi auto
Files
loader.py — main artifact of this repository
License
See the license… See the full description on the dataset page: https://huggingface.co/datasets/davidjqzo/recsys-proto.recsys-challenge-2026-artifacts
recsys-challenge-2026-artifacts
Intermediate artifacts of team FPMs_UMONS's pipeline for the
RecSys Challenge 2026. They let you run
Blind-B inference without recomputing the expensive preprocessing steps.
These are pipeline intermediates, not a usable dataset on their own. They
contain track ids and derived tags keyed to the official challenge catalog, and
only mean something to the code that consumes them. There is no standalone
loading recipe here on purpose:… See the full description on the dataset page: https://huggingface.co/datasets/MaximeM/recsys-challenge-2026-artifacts.recsys
dataset.py
Dataset Summary
A architecture dataset with audio video modality, stored in lmdb format.
Preprocessing & Augmentation
Preprocessing: auto ml
Augmentation: light
Splits & Sampling
Split strategy: leave one out
Sampling: active
Quality & Labeling
Quality filtering: lenient
Labeling: manual
Files
dataset.py — main artifact of this repository
License
See the license field… See the full description on the dataset page: https://huggingface.co/datasets/volkovdaniil/recsys.RecSys-Challenge-2022recsys-datasetaurora-recsys-v2-4108
Customer Feedback Corpus
Derived dataset published by the Data Platform team. Aggregates customer feedback records for sentiment modeling.
Provenance
Derived from TianfuXinqu/acme-sentiment-511pin4k.
