datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fine-t2i
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv]
by Xu Ma, Yitian Zhang,
Qihua Dong, Yun Fu
Northeastern Univeristy
Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient).
🆕 What's New
[2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️
[2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.T2I-CoReBench-Images
T2I-CoReBench-Images
📖 Overview
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.T2IScoreScore
Dataset Card for Text-to-Image ScoreScore (T2IScoreScore or TS2)
This dataset exists as part of the T2IScoreScore metaevaluation for assessing the faithfulness and consistency of text-to-image model prompt-image evaluation metrics.
Necessary code for utilizing the resource is present at github.com/michaelsaxon/T2IScoreScore
Dataset Details
Dataset Description
This is a test set of 165 "target prompts" which each have between 5 and 76 generated images of… See the full description on the dataset page: https://huggingface.co/datasets/saxon/T2IScoreScore.GPT4O_Image_T2IT2I-ImageNet-NormalT2I-RiskyPrompt-ImageDataset
T2I-RiskyPrompt-Derived Images
Overview
This image dataset is derived from the project T2I-RiskyPrompt. T2I-RiskyPrompt provides a hierarchical risk taxonomy (6 primary categories and 14 subcategories) and a set of 6,432 human-validated risky prompts, where each prompt is annotated with hierarchical labels and detailed risk reasons. This image dataset contains 20,373 images generated from SD3 and FLUX using T2I-RiskyPrompt, where each image is annotated with both… See the full description on the dataset page: https://huggingface.co/datasets/datarr/T2I-RiskyPrompt-ImageDataset.recap-t2i-evaluation-sample-2026
Recaptioned T2I Supervision Evaluation Sample
This repository is the small reviewer-inspection companion to the full anonymous caption-metadata release. The full release is hosted separately at https://huggingface.co/datasets/Anonymous1477/recap-t2i-evaluation-metadata-2026; this repository stays under the large-dataset sample threshold and gives reviewers a direct way to inspect redacted caption metadata, join structure, and selected image-conditioned audit packages.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1477/recap-t2i-evaluation-sample-2026.BW_illustrations_t2i_512x512Runway_Frames_t2i_human_preferences
Rapidata Frames Preference
This T2I dataset contains roughly 400k human responses from over 82k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Frames across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Runway_Frames_t2i_human_preferences.atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.UnifiedReward-2.0-T2X-score-data
Dataset Summary
UnifiedReward-2.0-T2X-score-data is added for our UnifiedReward-2.0-qwen-[3b/7b/32b/72b] training.
This dataset enables UnifiedReward-2.0 introducing several new capabilities:
Pairwise scoring for image and video generation assessment on Alignment, Coherence, Style dimensions.
Pointwise scoring for image and video generation assessment on Alignment, Coherence/Physics, Style dimensions.
Welcome to try the latest version, and the inference code is available at… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-2.0-T2X-score-data.OpenAI-4o_t2i_human_preference
Rapidata OpenAI 4o Preference
This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenAI-4o_t2i_human_preference.xAI_Aurora_t2i_human_preferences
Rapidata Aurora Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Aurora across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/xAI_Aurora_t2i_human_preferences.Imagen4_t2i_human_preference
Rapidata Imagen 4 Preference
This T2I dataset contains over 195k human responses from over 70k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Imagen 4 (imagen-4.0-ultra-generate-exp-05-20) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Imagen4_t2i_human_preference.Flux-2-pro_t2i_human_preference
Rapidata Flux 2 Pro Preference
This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Flux 2 Pro (version from 25.11.25) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux-2-pro_t2i_human_preference.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.glami-1m-t2i-mteb
GLAMI-1M text-to-image retrieval
This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association.
The unchanged source archives are already hosted by the original authors in glami/glami-1m. GLAMI-1M-dataset--test-only.zip is pinned at… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-t2i-mteb.HunyuanImage-2.1_t2i_human_preference
Rapidata Hunyuan Image 2.1 Preference
This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Hunyuan Image 2.1 (version from 19.9.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/HunyuanImage-2.1_t2i_human_preference.OpenGVLab_Lumina_t2i_human_preference
Rapidata Lumina Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Lumina across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/OpenGVLab_Lumina_t2i_human_preference.T2I-2Mt2i-model-comparison
T2I Model Comparison — 600 Prompts × 17 Models
A large-scale open comparison of 17 text-to-image models across 600 identical prompts, with 10,200 generated images and an interactive explorer.
🚀 Try the live comparison tool →
Pick a category and a prompt to see all 17 models render the exact same description — or drop any two side by side with full generation metadata.
Actively maintained: recently expanded to 17 models and fully re-synced image resolution and metadata across… See the full description on the dataset page: https://huggingface.co/datasets/kruatech/t2i-model-comparison.T2I-CoReBench
Easier Painting Than Thinking: Can Text-to-Image Models
Set the Stage, but Not Direct the Play?
Ouxiang Li1*, Yuan Wang1, Xinting Hu†, Huijuan Huang2‡, Rui Chen2, Jiarong Ou2,
Xin Tao2†, Pengfei Wan2, Xiaojuan Qi3, Fuli Feng1
1University of Science and Technology of China, 2Kling Team, Kuaishou Technology, 3The University of Hong Kong
*Work done during internship… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench.Science-T2I-Fullset
Science-T2I Fullset
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: SciScore
Huggingface: Science-T2I-S&C Benchmark
Data
The Science-T2I Fullset comprises a comprehensive collection of data for scientific T2I generation, including both training and test sets with a unified data structure. The test sets are split into 'test-S' and 'test-C,' corresponding to the Science-T2I-S and Science-T2I-C benchmarks, respectively.
Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Fullset.t2i-compbenchHub version of the T2I-CompBench dataset. All credits and licensing belong to the creators of the dataset.
This version was obtained as described below.
First, the ".txt" files were obtained from this directory.
Code
import requests
import os
# Set the necessary parameters
owner = "Karine-Huang"
repo = "T2I-CompBench"
branch = "main"
directory = "examples/dataset"
local_directory = "."
# GitHub API URL to get contents of the directoryurl =… See the full description on the dataset page: https://huggingface.co/datasets/NinaKarine/t2i-compbench.SyntheticFacesHighQuality-T2I122K curated 1024x1024 face images from text2image models (Flux1, SDXL, Dalle3)
This dataset consists of 122,726 high quality 1024x1024 curated face images, and was created by creating random prompt strings that were sent to multiple "text to image" models (Flux1.pro, Flux1.dev, Flux1.schnell, SDXL, DALL-E 3) and dropping bad generations using a semi manual curation process.
The prompts used to generate the dataset are of various faces with different attributes and various conditions to make… See the full description on the dataset page: https://huggingface.co/datasets/bitmind/SyntheticFacesHighQuality-T2I.Seedream-3_t2i_human_preference
Rapidata Seedream 3 Preference
This T2I dataset contains over ~400'000 human responses from over ~30'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Seedream-3_t2i_human_preference.fine-t2i-identity-low-2048
Fine T2I Identity Low 2048 — 100 Pairs
This dataset contains 100 original crops and their 100 matching GPT Image 2 reconstructions at 2048 × 2048 resolution: 100 examples and 200 JPEG files.
crops/<id>.jpg: the original crop.
generated/<id>.jpg: its GPT Image 2 output, requested with quality low.
manifest.json: the retained IDs, source dimensions, crop coordinates, sampling provenance, and reconstruction prompt.
comparison.html: a gallery with zoom and before/after sliders for… See the full description on the dataset page: https://huggingface.co/datasets/owenzlz/fine-t2i-identity-low-2048.Imagen-4-ultra-24-7-25_t2i_human_preference
Rapidata Imagen 4 Ultra 24.7.25 Preference
This T2I dataset contains over ~400'000 human responses from over ~83'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Imagen 4 Ultra (version from 24.7.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Imagen-4-ultra-24-7-25_t2i_human_preference.t20i-cricket-dataset-2005-2014
