datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenAI-Bench
GenAI-Bench
Paper |
🤗 GenAI Arena |
Github
Introduction
GenAI-Bench is a benchmark designed to benchmark MLLMs’s ability in judging the quality of AI generative contents by comparing with human preferences collected through our 🤗 GenAI-Arnea. In other words, we are evaluting the capabilities of existing MLLMs as a multimodal reward model, and in this view, GenAI-Bench is a reward-bench for multimodal generative models.
We filter existing votes collecte visa NSFW filter… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/GenAI-Bench.GenAI-Bench
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
Baiqi Li1*, Zhiqiu Lin1,2*, Deepak Pathak1, Jiayao Li1, Yixin Fei1, Kewen Wu1, Tiffany Ling1, Xide Xia2†, Pengchuan Zhang2†, Graham Neubig1†, and Deva Ramanan1†.
1Carnegie Mellon University, 2Meta
Links:
📖Paper | | 🏠Home Page | | 🔍GenAI-Bench Dataset Viewer | 🏆Leaderboard|
🗂️GenAI-Bench-1600(ZIP format) | | 🗂️GenAI-Bench-Video(ZIP format) | |… See the full description on the dataset page: https://huggingface.co/datasets/BaiqiL/GenAI-Bench.cardio-mark
CardioMark Review Subset
This repository contains an anonymized review subset of the CardioMark benchmark introduced for automated vertebral heart score (VHS) estimation in canine thoracic radiographs.
The subset is provided to support reproducibility and data-quality inspection during peer review.
Dataset Overview
CardioMark is a large-scale benchmark for evaluating the complete VHS measurement pipeline, including:
cardiac landmark localization
geometric VHS estimation… See the full description on the dataset page: https://huggingface.co/datasets/gen-ai-researcher/cardio-mark.genai-bench-sdxl-latents
GenAI-Bench SDXL with Latent Trajectories
SDXL generations for the GenAI-Bench prompts, paired with
the per-step decoded-latent trajectory of each generation. This dataset was created to evaluate NoisyCLIP.
For every prompt, 10 images were generated (1600 prompts → 16,000 generations). Each generation provides:
the full-resolution final image,
the 50-step denoising trajectory (each step's latent decoded to a small preview image), and
a CLIP-FlanT5-XXL VQAScore measuring… See the full description on the dataset page: https://huggingface.co/datasets/asiimo/genai-bench-sdxl-latents.GenAI-Bench-1600
This page contains only the zip format data of GenAI-Bench-1600, which can be downloaded manually or using git clone.
We recommend using the parquet format. The code is as follows:
from datasets import load_dataset
dataset = load_dataset("BaiqiL/GenAI-Bench")
hleGenAI-Bench-picturesvideo-framesmultimodal-vqa-self-instruct-enriched
Multimodal VQA – Self-Instruct-enriched
Overview
This dataset is an enriched, cleaned, and metadata-enhanced version of zwq2018/Multi-modal-Self-instruct.It pairs images with natural language questions and answers, making it ideal for Vision-Language Model (VLM) training, benchmarking, and instruction tuning.
Dataset Summary
Total samples: 75,000+ (64,796 train, 11,193 test)
Modalities: Image + Text (Questions) + Text (Answers)
Task Types: Visual Question… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/multimodal-vqa-self-instruct-enriched.GenAI_ImageClassificationLargeRationalRewards-EvalData-GenAIBench-MMRB2-ERBenchTLDR: this is the RewardModel Evaluation dataset for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba
RationalRewards is a… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-EvalData-GenAIBench-MMRB2-ERBench.GenAI-BenchGenAI-Bench-527GenAI-Image-Ranking-800GenAI-RealEstate-TestSet
🏙️ GenAI Real Estate Test Set (Track B)
Dataset for the MenaML Winter School 2026 Challenge.
📊 Dataset Structure
This dataset contains 1,000 images split evenly between:
Authentic: Real estate photography from the Places365 dataset.
Manipulated: Synthetically generated deepfake artifacts (Inpainting, Diffusion Noise, GAN Grids).
🕵️ How to Use
This dataset is designed for testing forensic detection models.
genai-book-imagesneutral-sd-outputsgenai-manipulation-detection-interior
GenAI Manipulation Detection Dataset - Interior Design Images
📋 Dataset Description
This dataset contains 1000 paired images (real + manipulated) for training and evaluating GenAI manipulation detection models. Created for the MenaML Winter School 2026 Hackathon.
Dataset Summary
Total Images: 1000 pairs (2000 total images)
Image Size: 512x512
Format: JPEG
Source: Pinterest Interior Design Images (Kaggle)
License: MIT
🎯 Challenge Context
This… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/genai-manipulation-detection-interior.rewardzero-t2i-hpd-genaigenai_datasetsGenAI-Arena-Bench-3GenAI-Bench_image_edition_processedGenAI-Bench
GenAI-Bench
Paper |
🤗 GenAI Arena |
Github
Introduction
GenAI-Bench is a benchmark designed to benchmark MLLMs’s ability in judging the quality of AI generative contents by comparing with human preferences collected through our 🤗 GenAI-Arnea. In other words, we are evaluting the capabilities of existing MLLMs as a multimodal reward model, and in this view, GenAI-Bench is a reward-bench for multimodal generative models.
We filter existing votes collecte visa NSFW… See the full description on the dataset page: https://huggingface.co/datasets/futurehuhu/GenAI-Bench.GenAI-Bench_image_generation_processedgenaibenchflickr30k-enriched-qa
Flickr30k Enriched QA (CLIP-Aware)
This dataset enriches the original Flickr30k validation/test set with:
Automatically generated QA prompts
Answers derived from human captions
Image tags extracted using CLIP embeddings
Each entry now supports multimodal QA, grounding tasks, and caption-based reasoning.
Fields:
image: Raw image
caption: List of 5 human-written captions
qa_prompt: e.g. "What is happening in the image that contains: outdoor, playing, child?"
qa_answer: Chosen from… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/flickr30k-enriched-qa.imagesGenAI-Arena-human-evalGenAI_Emu_1kThis is the 1k images generated by Emu, which is supposed to be used as the benchmark in GenAi Benchmark Challenge Workshop @ CVPR2015
qwen-image-base-genaibenchmark
