Rapidata/Imagen4_t2i_human_preference
Rapidata Imagen 4 Preference This T2I dataset contains over 195k human responses from over 70k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Imagen 4 (imagen-4.0-ultra-generate-exp-05-20) across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Imagen4_t2i_human_preference.
<style>
.vertical-container { display: flex; flex-direction: column; gap: 60px; }
.image-container img { max-height: 250px; / Set the desired height / margin:0; object-fit: contain; / Ensures the aspect ratio is maintained / width: auto; / Adjust width automatically based on height / box-sizing: content-box; }
.image-container { display: flex; / Aligns images side by side / justify-content: space-around; / Space them evenly / align-items: center; / Align them vertically / gap: .5rem }
.container { width: 90%; margin: 0 auto; }
.text-center { text-align: center; }
.score-amount { margin: 0; margin-top: 10px; }
.score-percentage {Score: font-size: 12px; font-weight: semi-bold; }
</style>
Rapidata Imagen 4 Preference
<a href="https://www.rapidata.ai"> <img src="https://cdn-uploads.huggingface.co/production/uploads/66f5624c42b853e73e0738eb/jfxR79bOztqaC6_yNNnGU.jpeg" width="400" alt="Dataset visualization"> </a>
This T2I dataset contains over 195k human responses from over 70k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Imagen 4 (imagen-4.0-ultra-generate-exp-05-20) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it ❤️
Overview
This T2I dataset contains over 195k human responses from over 70k individual annotators, collected in just ~1 Day. Evaluating Imagen 4 (imagen-4.0-ultra-generate-exp-05-20) across three categories: preference, coherence, and alignment.
The evaluation consists of 1v1 comparisons between Imagen 4 (imagen-4.0-ultra-generate-exp-05-20) and 15 other models: Hidream I1 full, Halfmoon-4-4-2025, OpenAI 4o-26-3-25, Ideogram V2, Recraft V2, Lumina-15-2-25, Frames-23-1-25, Imagen-3, Flux-1.1-pro, Flux-1-pro, DALL-E 3, Midjourney-5.2, Stable Diffusion 3, Aurora and Janus-7b.
Note: The number following the model name (e.g., Halfmoon-4-4-2025) represents the date (April 4, 2025) on which the images were generated to give an understanding of what model version was used.
Alignment
The alignment score quantifies how well an video matches its prompt. Users were asked: "Which image matches the description better?".
<div class="vertical-container"> <div class="container"> <div class="text-center"> <q>A blue cup and a green cell phone.</q> </div> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4 </h3> <div class="score-percentage">Score: 100%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/8JcDYOeviHCwVQo0vLMB-.jpeg" width=500> </div> <div> <h3 class="score-amount">Frames-23-1-25 </h3> <div class="score-percentage">Score: 0%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/WmHuNZY4M_NWUMSTSofNL.jpeg" width=500> </div> </div> </div>
<div class="container"> <div class="text-center"> <q>A person is walking with a guidebook and taking a tour of a historic site.</q> </div> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4</h3> <div class="score-percentage">Score: 0%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/RiyfAra-C5QcCT9lDiohg.jpeg" width=500> </div> <div> <h3 class="score-amount">Aurora</h3> <div class="score-percentage">Score: 100%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/SPbiu_SE8S4vRv6p0OZHo.jpeg" width=500> </div> </div> </div> </div>
Coherence
The coherence score measures whether the generated video is logically consistent and free from artifacts or visual glitches. Without seeing the original prompt, users were asked: "Which image has more glitches and is more likely to be AI generated?"
<div class="vertical-container"> <div class="container"> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4 </h3> <div class="score-percentage">Glitch Rating: 0%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/X2jimr0Icoqb4RlX6jG1j.jpeg" width=500> </div> <div> <h3 class="score-amount">Stable Diffusion 3 </h3> <div class="score-percentage">Glitch Rating: 100%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/UUllaDc48oxLXxeXbF_CD.jpeg" width=500> </div> </div> </div>
<div class="container"> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4 </h3> <div class="score-percentage">Glitch Rating: 100%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/1a5dQzX2XxTLIUcgimocy.jpeg" width=500> </div> <div> <h3 class="score-amount">Ideogram</h3> <div class="score-percentage">Glitch Rating: 0%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/pI5B7RxysOcUhVEzdqHFq.jpeg" width=500> </div> </div> </div> </div>
Preference
The preference score reflects how visually appealing participants found each image, independent of the prompt. Users were asked: "Which image do you prefer?"
<div class="vertical-container"> <div class="container"> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4</h3> <div class="score-percentage">Score: 81.83%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/b3ny4C5fRdfnoIlK3ocFi.jpeg" width=500> </div> <div> <h3 class="score-amount">Frames-23-1-25</h3> <div class="score-percentage">Score: 18.17%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/KDJSo6s6FMKpXiovP03cJ.jpeg" width=500> </div> </div> </div>
<div class="container"> <div class="image-container"> <div> <h3 class="score-amount">Imagen 4 </h3> <div class="score-percentage">Score: 27.42%</div> <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/-g7n1pP8QRU0Gcv-iKF4q.jpeg" width=500> </div> <div> <h3 class="score-amount">Frames-23-1-25 </h3> <div class="score-percentage">Score: 72.58%</div> <img style="border: 3px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/BMXzpF1loC3wVfWuyj5VX.jpeg" width=500> </div> </div> </div> </div>
About Rapidata
Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.
