datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FoR-T2I
Can Text-to-Image Models Draw from the Right Frame of Reference?
Left is not always image-left. If a person is facing the viewer, their left hand appears on the right side of the image. A text-to-image model can therefore render a plausible scene with the requested objects while still drawing the spatial relation from the wrong perspective.
FoR-T2I is the official benchmark release for Can Text-to-Image Models Draw from the Right Frame of Reference?. It… See the full description on the dataset page: https://huggingface.co/datasets/ernie-research/FoR-T2I.T2I-CoReBench
Easier Painting Than Thinking: Can Text-to-Image Models
Set the Stage, but Not Direct the Play?
Ouxiang Li1*, Yuan Wang1, Xinting Hu†, Huijuan Huang2‡, Rui Chen2, Jiarong Ou2,
Xin Tao2†, Pengfei Wan2, Xiaojuan Qi3, Fuli Feng1
1University of Science and Technology of China, 2Kling Team, Kuaishou Technology, 3The University of Hong Kong
*Work done during internship… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench.DIM-T2I
[ICLR 2026] Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
Ziyun Zeng, David Junhao Zhang, Wei Li,
and Mike Zheng Shou
📰 News
[2026-05-12] The DIM project page is available.
[2026-01-26] 🎉 DIM is accepted to ICLR 2026!
[2025-10-08] 🚀 Released the DIM-Edit dataset and the DIM-4.6B-T2I / DIM-4.6B-Edit models.
[2025-09-02] 📝 The DIM paper is released on arXiv.
🌟 Highlights
🧠… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/DIM-T2I.Arena-T2I-Hard
Arena-T2I-Hard
A 310-prompt stress benchmark for evaluating faithfulness (prompt-following) of
text-to-image models, drawn from real, hard arena user requests — long, multi-entity
prompts with attributes, spatial relations, counts, and stylistic constraints. Each
prompt ships pre-decomposed into a dependency-aware DAG of yes/no questions; when
scoring an image, failing a parent question zeroes out its descendants. The benchmark
stays discriminative where DPG-Bench and DSG… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Arena-T2I-Hard.T2I-Eval-BenchThis is the human-annotated benchmark dataset for paper Automatic Evaluation for Text-to-Image Generation: Fine-grained Framework,
Distilled Evaluation Model and Meta-Evaluation Benchmark
NOTE: Please check out our github repository for more detailed usage.
LongBench-T2I
LongBench-T2I
LongBench-T2I is a benchmark dataset introduced in the paper Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.It is a standalone dataset designed specifically for evaluating text-to-image (T2I) generation models under long and compositionally rich prompts.
📦 Dataset Summary
This dataset contains 500 samples, each composed of:
A long-form instruction (complex natural language prompt).
A… See the full description on the dataset page: https://huggingface.co/datasets/YCZhou/LongBench-T2I.ea-cot-t2i
EA-CoT-T2I
EA-CoT-T2I is the text-to-image (T2I) companion release of the Evaluation Agent chain-of-thought supervision data. It is published separately from the T2V-only EA-CoT-10K dataset.
The dataset contains history-conditioned next-step records distilled from multi-round T2I model-evaluation trajectories. It is designed for supervised fine-tuning of an evaluation planner that selects an evaluation tool, interprets the resulting observation, and eventually produces a… See the full description on the dataset page: https://huggingface.co/datasets/open-ea/ea-cot-t2i.t2ivT2I-benchmark
T2I Benchmark: Frontier Text-to-Image Models on Image Description Prompts
Accompanying dataset for Benchmarking Frontier Text-to-Image Models on the Image Description Prompts (Perle AI).
Four frontier text-to-image systems are compared on the 48 hardest prompts in the DataSeeds.AI Sample Dataset (DSD). Every prompt is a verbatim, human-written description of a real photograph — no prompt engineering. Every generated image is graded against a per-prompt weighted rubric by an… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/T2I-benchmark.MMDecodingTrust-T2I
Overview
This repo contains the text-to-image dataset of MMDT (Multimodal DecodingTrust). This research endeavor is designed to help researchers and practitioners better understand the capabilities, limitations, and potential risks involved in deploying the state-of-the-art Multimodal foundation models (MMFMs). This dataset focuses on the following six primary perspectives of trustworthiness, including safety, hallucination, fairness, privacy, adversarial robustness, and… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/MMDecodingTrust-T2I.deco-t2i-256-80k-baseline-nogate-20260924
