datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.text-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.text-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.text-2-video-human-preferences-veo3.1
Rapidata Video Generation Veo 3.1 Human Preference
In this dataset, ~74k human responses from ~23k human annotators were collected to evaluate the Veo 3.1 video generation model on our benchmark. This dataset was collected using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it ❤️… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.1.text-2-video-human-preferences-veo2
Rapidata Video Generation Google DeepMind Veo2 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Google DeepMind Veo2 video generation model on our benchmark. The up to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo2.text-2-video-human-preferences-genmo-mochi-1
Rapidata Video Generation Genmo Mochi-1 Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate mochi-1 video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-genmo-mochi-1.text-2-video-human-preferences-pika2.2
Rapidata Video Generation Pika 2.2 Human Preference
In this dataset, ~756k human responses from ~29k human annotators were collected to evaluate Pika 2.2 video generation model on our benchmark. This dataset was collected in ~1 day total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-pika2.2.fold_pants_preferences
fold_pants — pairwise preferences on a Franka Panda
Real-robot trajectories for "fold the shorts" with human pairwise preference
labels on multiple judgment axes. Built for reward-model / preference-learning
research: every label is a comparison of two trajectories on one named axis, not a
scalar score.
The trajectory data is a standard LeRobot
v2.1 dataset, so it also loads directly as an imitation-learning dataset.
Contents
Episodes
536
Frames
530… See the full description on the dataset page: https://huggingface.co/datasets/MarcelTorne/fold_pants_preferences.setup_table_preferences
setup_table — pairwise preferences on a Franka Panda
Real-robot trajectories for "set up the table" with human pairwise preference
labels on multiple judgment axes. Built for reward-model / preference-learning
research: every label is a comparison of two trajectories on one named axis, not a
scalar score.
The trajectory data is a standard LeRobot
v2.1 dataset, so it also loads directly as an imitation-learning dataset.
Contents
Episodes
467
Frames… See the full description on the dataset page: https://huggingface.co/datasets/MarcelTorne/setup_table_preferences.plate_toast_preferences
plate_toast — pairwise preferences on a Franka Panda
Real-robot trajectories for "put the toast in the plate" with human pairwise preference
labels on multiple judgment axes. Built for reward-model / preference-learning
research: every label is a comparison of two trajectories on one named axis, not a
scalar score.
The trajectory data is a standard LeRobot
v2.1 dataset, so it also loads directly as an imitation-learning dataset.
Contents
Episodes
271… See the full description on the dataset page: https://huggingface.co/datasets/MarcelTorne/plate_toast_preferences.put_cube_in_bowl_preferences
put_cube_in_bowl — pairwise preferences on a Franka Panda
Real-robot trajectories for "put the cube in the bowl" with human pairwise preference
labels on multiple judgment axes. Built for reward-model / preference-learning
research: every label is a comparison of two trajectories on one named axis, not a
scalar score.
The trajectory data is a standard LeRobot
v2.1 dataset, so it also loads directly as an imitation-learning dataset.
Contents
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/MarcelTorne/put_cube_in_bowl_preferences.image-2-video-human-preferences-large
I2V Human Preferences (Large)
Human preference dataset for image-to-video (I2V) generation quality. Each row contains a reference image, two generated videos (one from Pika and one from CogVideoX), and 10 human preference annotations aggregated via majority vote.
This is the large (3,000-row) subset — the complete dataset. See also: small (1,000 rows), medium (2,000 rows).
Dataset Summary
Metric
Value
Total rows
3,000
Annotations per row
10
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/image-2-video-human-preferences-large.text-2-video-human-preferences-motion
Human Preferences for AI-Generated Video: Motion Quality
29,283 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from 4,349 real annotators via Datapoint AI.
This is the largest publicly available human preference dataset focused specifically on human motion in AI-generated video.
Why This Dataset
Video generation models are improving fast, but evaluating human motion remains… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion.text-2-video-ranking-human-preferences
T2V Ranking Human Preferences
~91,000 human ranking labels across 18 text-to-video models on 3 quality dimensions, collected from real annotators via Datapoint AI.
This is the first public ranking-based (not pairwise) human preference dataset for text-to-video generation. Each datapoint contains 5 videos generated from the same prompt by different models, ranked 1st through 5th by 15 annotators on each dimension.
Why This Dataset
Existing video preference… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-ranking-human-preferences.text-2-video-human-preferences-motion-v2-large
Human Preferences for AI-Generated Video: Motion Quality v2 (large)
115,732 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from real annotators via Datapoint AI.
This is an expanded version of the motion quality dataset with 417 unique prompts (up from 60) and 11 motion categories (up from 6).
Why This Dataset
Video generation models are improving fast, but evaluating human motion… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion-v2-large.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/yangge10/text-2-video-human-preferences.image-2-video-human-preferences-medium
I2V Human Preferences (Medium)
Human preference dataset for image-to-video (I2V) generation quality. Each row contains a reference image, two generated videos (one from Pika and one from CogVideoX), and 10 human preference annotations aggregated via majority vote.
This is the medium (2,000-row) subset. See also: small (1,000 rows), large (3,000 rows).
Dataset Summary
Metric
Value
Total rows
2,000
Annotations per row
10
Total annotations
20,000
Unique… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/image-2-video-human-preferences-medium.text-2-video-human-preferences-motion-v2-medium
Human Preferences for AI-Generated Video: Motion Quality v2 (medium)
57,866 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from real annotators via Datapoint AI.
This is an expanded version of the motion quality dataset with 417 unique prompts (up from 60) and 11 motion categories (up from 6).
Why This Dataset
Video generation models are improving fast, but evaluating human motion… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion-v2-medium.image-2-video-human-preferences-small
I2V Human Preferences (Small)
Human preference dataset for image-to-video (I2V) generation quality. Each row contains a reference image, two generated videos (one from Pika and one from CogVideoX), and 10 human preference annotations aggregated via majority vote.
This is the small (1,000-row) subset. See also: medium (2,000 rows), large (3,000 rows).
Dataset Summary
Metric
Value
Total rows
1,000
Annotations per row
10
Total annotations
10,000
Unique… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/image-2-video-human-preferences-small.
