datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.text-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.world-model-physics
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same
real-world starting frame and scene description from Physics-IQ and asked
them to predict what happens next. ~283,000 human votes, collected with the
Rapidata Python SDK, decided which continuation is more realistic — with the
real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics.text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.text-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.sora-video-generation-style-likert-scoring
Rapidata Video Generation Preference Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~6000 human evaluators were asked to rate AI-generated videos based on their visual appeal, without seeing the prompts used to generate them. The specific… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-style-likert-scoring.text-2-video-human-preferences-veo2
Rapidata Video Generation Google DeepMind Veo2 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Google DeepMind Veo2 video generation model on our benchmark. The up to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo2.text-2-video-human-preferences-veo3.1
Rapidata Video Generation Veo 3.1 Human Preference
In this dataset, ~74k human responses from ~23k human annotators were collected to evaluate the Veo 3.1 video generation model on our benchmark. This dataset was collected using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it ❤️… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.1.camera-movement
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 281,738 human responses, collected with the
Rapidata Python SDK, comparing how well 14 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement.sora-video-generation-physics-likert-scoring
Rapidata Video Generation Physics Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~6000 human evaluators were asked to rate AI-generated videos based on if gravity and colisions make sense, without seeing the prompts used to generate them.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-physics-likert-scoring.text-2-video-human-preferences-genmo-mochi-1
Rapidata Video Generation Genmo Mochi-1 Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate mochi-1 video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-genmo-mochi-1.text-2-video-human-preferences-pika2.2
Rapidata Video Generation Pika 2.2 Human Preference
In this dataset, ~756k human responses from ~29k human annotators were collected to evaluate Pika 2.2 video generation model on our benchmark. This dataset was collected in ~1 day total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-pika2.2.sora-video-generation-alignment-likert-scoring
Rapidata Video Generation Prompt Alignment Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~6000 human evaluators were asked to evaluate AI-generated videos based on how well the generated video matches the prompt. The specific question… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-alignment-likert-scoring.text-2-video-human-preferences-sora-2
Rapidata Video Generation Sora 2 Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate the Sora 2 video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-sora-2.multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.text-2-video-human-preferences-kling-v2.1-master
Rapidata Video Generation Kling v2.1 Master Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Kling v2.1 Master video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-kling-v2.1-master.text-2-video-human-preferences-sora-2-pro
Rapidata Video Generation Sora 2 Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate the Sora 2 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-sora-2-pro.text-2-video-human-preferences-luma-ray2
Rapidata Video Generation Luma Ray2 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Luma's Ray 2 video generation model on our benchmark. The up to date benchmark can be… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-luma-ray2.text-2-video-human-preferences-runway-alpha
Rapidata Video Generation Runway Alpha Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~30'000 human annotations were collected to evaluate Runway's Alpha video generation model on our benchmark. The up to date benchmark can… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-runway-alpha.sora-video-generation-time-flow
Rapidata Video Generation Time flow Annotation Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~3700 human evaluators were asked to evaluate AI-generated videos based on how time flows in the video. The specific question posed was: "How… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-time-flow.text-2-video-Rich-Human-Feedback
Rapidata Video Generation Rich Human Feedback Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~4 hours total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~22'000 human annotations were collected to evaluate AI-generated videos (using Sora) in 5 different categories.
Prompt - Video… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-Rich-Human-Feedback.1k-ranked-videos-coherence
1k Ranked Videos
This dataset contains approximately one thousand videos, ranked from most preferred to least preferred based on human feedback from over 25k pairwise comparisons. The videos are rated solely on coherence as evaluated by human annotators, without considering the specific prompt used for generation. Each video is associated with the model name that generated it.
The videos are sampled from our benchmark dataset text-2-video-human-preferences-pika2.2. Follow us to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/1k-ranked-videos-coherence.text-2-audio-human-preference-benchmark
Text to Audio Human Benchmark
In this dataset, ~32k human responses collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
The annotators were asked Which voice is more friendly? and Which voice sounds more natural? respectively.
Check out the Benchmark!
Translation-deepseek-llama-mixtral-v-deepl
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
This dataset contains ~51k responses from ~11k annotators and compares the translation capabilities of DeepSeek-R1(deepseek-r1-distill-llama-70b-specdec), Llama(llama-3.3-70b-specdec) and Mixtral(mixtral-8x7b-32768) against DeepL across different languages. The comparison involved 100 distinct questions in 4 languages, with each translation being rated by 51 native… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Translation-deepseek-llama-mixtral-v-deepl.happiness-per-country
Rapidata Happiness per Country
This dataset contains responses from 100 people in each country to the question:
"How happy are you with your life right now?"
Important Biases and Limitations
This dataset is intended as a fun exercise and has several important biases to
acknowledge:
Selection bias: All respondents answered on a mobile phone, which means
the sample skews toward people with higher income and socioeconomic status,
especially in poorer regions.
Coverage… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/happiness-per-country.Translation-gpt4o_mini-v-gpt4o-v-deepl
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
This dataset compares the translation capabilities of GPT-4o and GPT-4o-mini against DeepL across different languages. The comparison involved 100 distinct questions (found under raw_files) in 4 languages, with each translation being rated by 100 native speakers. Texts that were translated identically across platforms were excluded from the analysis.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Translation-gpt4o_mini-v-gpt4o-v-deepl.
