datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.Obshazard-bench
ObsCrisis-Bench
A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks.
Dataset Description
ObsCrisis-Bench contains 4,202 VQA samples across 127 extreme weather events in 8 disaster categories, covering 61 countries. Each sample combines satellite multispectral imagery (AMSU-A, HIRS, MHS sensors) with optional weather station data, and requires models to perform risk assessment, type classification, timing… See the full description on the dataset page: https://huggingface.co/datasets/YYQ898/Obshazard-bench.Co-Spy-Bench
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI (CVPR 2025)
With the rapid advancement of generative AI, it is now possible to synthesize high-quality images in a few seconds. Despite the power of these technologies, they raise significant concerns regarding misuse.
To address this, various synthetic image detectors have been proposed. However, many of them struggle to generalize across diverse generation parameters and emerging generative models.
In… See the full description on the dataset page: https://huggingface.co/datasets/ruojiruoli/Co-Spy-Bench.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.DamageTriage-Bench
DamageTriage-Bench
DamageTriage-Bench is a footprint-conditioned benchmark for per-building damage
typing from single post-event aerial images. Its five classes distinguish roof
from structural damage and partial from total affected extent:
ID
Class
0
Undamaged
1
Partial Roof Damage
2
Total Roof Damage
3
Partial Structural Damage
4
Total Structural Collapse
Quick statistics
Item
Value
Tiles
7,472 (1024 × 1024 PNG)
Labeled… See the full description on the dataset page: https://huggingface.co/datasets/Ymx1025/DamageTriage-Bench.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.holisafe-bench
⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised.
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings)
🌐 Website | 📑 Paper
📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.CTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.GeoFidelity-Bench
GeoFidelity-Bench
GeoFidelity-Bench evaluates whether generated street-view images match a
requested location at the level of named street blocks. The release contains
109 named street blocks from 25 cities, 7,117 curated Mapillary reference
images, generated images from six open-weight text-to-image models, prompt
control metadata, and 109-target benchmark result summaries. The generated-image index covers
15,696 released JPEG files across six models, six prompt or control… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.PCF-Bench
PCF-Bench
A photonic-crystal-fiber (PCF) inverse-design benchmark for vision-language models. Each sample bundles geometric parameters, simulated mode-field images, and a four-level expert-style annotation suite supporting tasks across geometry perception, physics understanding, multimodal reasoning, inverse design, and code generation.
Note (anonymous review). This dataset card omits identifying information for double-blind review.
Note (review subset). The full source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/PCF-Bench/PCF-Bench.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,355,161 human responses, collected with the
Rapidata Python SDK, comparing how well 30 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.MWS-Antifraud-Bench
MWS Antifraud Bench (Validation)
Experimental document-authenticity task for general-purpose multimodal language
models. This is the public validation part of MWS Vision Bench anti-fraud v0.1.
The dataset is released for research and model comparison. It is not a
certification tool, a production fraud-detection system, or a universal
leaderboard that is expected to be resistant to deliberate optimization.
Data
The validation split contains 209 items:
44 ai_gen;… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/MWS-Antifraud-Bench.SFHQ-VirtualID-Bench
SFHQ-VirtualID-Bench Dataset Card
Summary
SFHQ-VirtualID-Bench is a synthetic, identity-conditioned face dataset of
750 synthetic identities, built as a benchmark for machine unlearning.
Each identity is a deletion unit: all 90 aligned 224×224 crops of one
identity share a single identity-level split (retain or forget), and 75
identities follow a sequential 15-step forgetting protocol. Every identity
contributes both per-image image_subset values (train + holdout)… See the full description on the dataset page: https://huggingface.co/datasets/FaizPalwala/SFHQ-VirtualID-Bench.locus-tag-bench
Locus-Tag Bench
Synthetic benchmark suite for fiducial-tag detection and camera calibration, rendered with render-tag.
Each config corresponds to one campaign (board family × resolution, or a lighting/sensor variant). Images are rendered in Blender Cycles with ground-truth geometry recovered directly from the scene — no detector-in-the-loop, no human labels.
Configs
Config
Purpose
Images
Board/Tag
Resolution
locus_v1_tag36h11_640x480
Detection (low-res)… See the full description on the dataset page: https://huggingface.co/datasets/NoeFontana/locus-tag-bench.ppe-benchmark-eval
PPE Benchmark Eval Set (v1)
A held-out, human-verified benchmark for evaluating vision-language models on
personal protective equipment (PPE) detection — specifically hardhat and
safety-vest presence — framed as a VQA-style classification task.
What this is
96 images, balanced 24/24/24/24 across the four hardhat × vest combinations
(yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of
the karabuk-university PPE dataset
on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.MFC-Bench
MFC-Bench: Multimodal Fact-Checking Benchmark
MFC-Bench is a comprehensive Multimodal Fact-Checking testbed designed to evaluate LVLMs in terms of identifying factual inconsistencies and counterfactual scenarios.
Dataset Description
From the paper: "MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models"
MFC-Bench encompasses a wide range of visual and textual queries, organized into three binary classification tasks:
1. Manipulation… See the full description on the dataset page: https://huggingface.co/datasets/MM-Hallu/MFC-Bench.HUGO-Bench
HUGO-Bench
Hierarchical Unsupervised Grouping of Organisms Benchmark
A comprehensive benchmark dataset for evaluating zero-shot clustering of wildlife camera trap images using Vision Transformer embeddings.
Overview
HUGO-Bench contains 139,111 expert-validated cropped images of 60 animal species (30 birds, 30 mammals), derived from 23 camera trap projects across LILA BC. The dataset enables benchmarking of Vision Transformer models for unsupervised species-level… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench.GeneLab_BPS_BenchmarkData
Dataset Card for Dataset GeneLab_BPS_BenchmarkData
Dataset Details
This dataset is a version of the Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset managed by NASA and hosted on an S3 Bucket here: https://registry.opendata.aws/bps_microscopy/
Fluorescence microscopy images of individual nuclei from mouse fibroblast cells, irradiated with Fe particles or X-rays with fluorescent foci indicating 53BP1 positivity, a marker of DNA damage.… See the full description on the dataset page: https://huggingface.co/datasets/kenobi/GeneLab_BPS_BenchmarkData.CC-Bench
CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection
CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection.
This repository currently contains:
4,282 images in total
2,157… See the full description on the dataset page: https://huggingface.co/datasets/annoymous-1/CC-Bench.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.PANDA-PLUS-Bench
PANDA-PLUS-Bench
A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models.
Dataset Description
PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations.
Dataset Summary
Patches: ~2,770 per augmentation condition
Resolution: 224×224 pixels at 20× magnification
Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3)
Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.MM-Bench-E-CommerceThis is the HuggingFace repository of the paper named MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding in WSDM 2026 (oral).
In this paper, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning.
We propose the first generative MLLM-based model named MOON for product representation learning.
Furthermore, we contruct and publish a large-scale real-world… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/MM-Bench-E-Commerce.Syncred-Bench
Syncred-Bench
SynCred-Bench is a benchmark designed to evaluate synthetic credibility: AI-generated images that appear trustworthy by imitating authoritative visual forms (e.g., fake notices, credentials, news layouts) and realistic circulation traces.
The benchmark contains 600 AI-generated misinformation images across six credible-form categories and seven circulation styles. It also introduces FP450, a real-image negative set for measuring false positives in detection… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/Syncred-Bench.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.chest-bench-example
ChestBench Example
DICOM-VLM Framework Reference Package v0.2.0
ChestBench Example is a four-case, DICOM-native reference package for developing and validating the data architecture of a medical vision-language model (VLM) pipeline.
It is intentionally small. Its purpose is to demonstrate how medical imaging data, annotations, text, knowledge, retrieval targets, QA, evidence requirements, perturbations, and audit metadata can be represented without confusing… See the full description on the dataset page: https://huggingface.co/datasets/NeeyuHuynh/chest-bench-example.Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.ai-detector-benchmark-test-data
🎯 AI Detector Benchmark Test Dataset
A comprehensive benchmark dataset for testing AI image detection models.
📊 Dataset Summary
Total Images: 700
AI-Generated: 250 images (from 5 different generators)
Real Images: 450 images (from 9 diverse datasets)
Perfect for:
✅ Testing AI detection models
✅ Creating leaderboards
✅ Comparing model performance
✅ Benchmarking new approaches
🤖 AI Generators Included
Generator
Images
Accuracy Baseline
FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.
