datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ds-coder-instruct-v1
Dataset Card for DS Coder Instruct Dataset
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python.
The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.WorldReward-Bench
WorldReward-Bench
A human-annotated preference benchmark for camera-conditioned world models.
760 pairs of videos, each pair generated by two different models from the same
source image and the same camera-action sequence, with human verdicts on three
independent axes.
📰 Paper: https://arxiv.org/abs/2609.03952
🪐 Project Page: https://codegoat24.github.io/WorldReward
🤗 Model Collections: https://huggingface.co/CodeGoat24/WorldReward-9B
🚀 Github:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/WorldReward-Bench.ImageGen-CoT-Reward-5K
ImageGen_Reward_Cold_Start
Dataset Summary
This dataset is distilled from GPT-4o for our UnifiedReward-Think-7b cold-start training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2505.03318
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/Think
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/ImageGen-CoT-Reward-5K.OIP
OIP
Dataset Summary
This dataset is derived from open-image-preferences-v1-binarized for our UnifiedReward-7B training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2503.05236
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/OIP.HPD
HPD
Dataset Summary
This dataset is derived from 700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3 for our UnifiedReward-7B training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2503.05236
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/HPD.codette_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.bots_ultime_datos_proyect
