datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.text-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated images… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.text2image100kSampled_AIGCBench_text2image_ar_0.625
Description
This dataset is intended for the implementation of image-to-video generation evaluations in the paper of AdaptiveDiffusion, which is composed of the original text-image pairs collected from AIGCBench v1.0 and a text file listing the randomly selected samples.
Data Organization
The dataset is organized into the following files:
AIGCBench_t2i_aspect_ratio_625.zip: 2002 images named by the index and the text description, adjusted to an aspect ratio of 0.625.… See the full description on the dataset page: https://huggingface.co/datasets/HankYe/Sampled_AIGCBench_text2image_ar_0.625.ShareGPT-4o-Text2Imageatomic2023-small_text2imagetext-2-image-dpo-human-preferences
Text-2-Image DPO Human Preferences
A large-scale, quality-controlled human preference dataset for text-to-image generation. 80,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
Built on the Datapoint annotation platform — purpose-built infrastructure for collecting high-quality human preference data at scale.
Overview
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences.Text2image-ChinesePainting
Text2image-ChinesePainting Dataset
This repository provides 2192 pairs of Traditional Chinese Landscape Paintings and their corresponding descriptive texts. The paintings are sourced from the Chinese Landscape Painting Dataset, and the descriptive texts are generated using GPT-4o's API.
The dataset is intended for use in training and fine-tuning text-to-image models, focusing on generating Chinese landscape art from textual descriptions.
Dataset Overview
Number of… See the full description on the dataset page: https://huggingface.co/datasets/zqman/Text2image-ChinesePainting.text2image1mtext2image-fupotext2image10ktext2image_en_vi_captionstext-2-image-dpo-human-preferences-small
Text-2-Image DPO Human Preferences (Small)
A quality-controlled human preference dataset for text-to-image generation. 40,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the highest-annotator-quality subset. For the full 5,000-pair dataset, see datapointai/text-2-image-dpo-human-preferences.
Built on the Datapoint annotation platform — purpose-built… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-small.text2image-man
