datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rejection-sampling-replace-repeat
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling-replace-repeat.novel-repeat-dataset
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/novel-repeat-dataset.genrec_reasoning
GenRec Reasoning
Regenerated reasoning data for generative recommendation experiments.
igc1-top50-rejection-sampling
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/igc1-top50-rejection-sampling.rejection-sampling-V4
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling-V4.V2-rejection-sampling-dataset
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/V2-rejection-sampling-dataset.diversity-reward-analysis
Diversity Reward Analysis
Video Games constrained-beam evaluation outputs for diversity-reward and no-diversity-reward models. The Video_Games_SBS_metrics subset provides their metrics side by side.
genrec_reasoning_new
GenRec Reasoning
Regenerated reasoning data for generative recommendation experiments.
recsys-genrec-dataset-gpt5.4
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/recsys-genrec-dataset-gpt5.4.rejection-sampling
🧠 Amazon Semantic-ID Recommendation + Reasoning — GPT-5.4 regenerated
GPT-5.4 enrichment variant of yufan/recsys-genrec-dataset · Reasoning over Semantic IDs Enhances Generative Recommendation
This repo mirrors the format of the original yufan/recsys-genrec-dataset, but the
LLM-enriched fields are regenerated with GPT-5.4 (Azure OpenAI). The Video Games domain is fully populated.
<cat> below is Video_Games.
✅ Currently available
config… See the full description on the dataset page: https://huggingface.co/datasets/budgiesarecooliguess/rejection-sampling.igc1-top50-thirdsV2-rejection-sampling-checkpointsrejection-sampling-checkpoints
