datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.clamp-benchmark
CLAMP: A Sim-to-Real Benchmark for Closed-Loop Kinematic Pose Estimation and Assembly Reasoning
Closed-Loop Assembly and Mechanism Perception
📦 Code: https://anonymous.4open.science/r/clamp-131F
Your browser does not support the video tag.
Overview of all 210 labeled real test scenes.
Overview of our benchmark. Top left: A robot arm with its kinematic graph overlaid—open-chain edges in cyan, closed-loop edges in orange. Removing a constraining link (red oval) breaks the… See the full description on the dataset page: https://huggingface.co/datasets/clamp-benchmark/clamp-benchmark.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.CTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.CoVAtt-Benchmark
CoVAtt-Benchmark
A large-scale benchmark for generated-image attribution: given an image
that was produced by some text-to-image model, decide which model
produced it, and decide whether it came from a model the system has never
seen before.
What this dataset is
This is the dataset used to train and evaluate CoVAtt
(Content-Based Verification for Attribution of AI-Generated Images,
BMVC 2026). CoVAtt is a Siamese network that takes a pair of images and
predicts… See the full description on the dataset page: https://huggingface.co/datasets/kenyag/CoVAtt-Benchmark.PRISM_Benchmark
PRISM: PhotoRealistic Image Synthesis and Manipulation
Dataset repository for the paper "The PRISM benchmark: PhotoRealistic Image Synthesis and Manipulation to detect generated images" — Bartolucci, Salti, Lisanti.
Real images
The corresponding real images can be downloaded separately from COCO:
Split
Source
Test set
COCO 2017 Val
Training set
COCO 2017 Train
Citation
@article{BARTOLUCCI2026104826,
title = {The PRISM… See the full description on the dataset page: https://huggingface.co/datasets/oppiliF/PRISM_Benchmark.ImageTime_Benchmark
ImagineTime Benchmark
This dataset repository contains the public benchmark assets for ImagineTime, released with the paper “Can Image Models Imagine Time?”
Paper: arXiv:2606.10620
ImagineTime evaluates whether image generation models can produce ordered 2x2 motion sheets with coherent entities, spatial relations, state transitions, interactions, and task constraints.
Contents
cases/
750 benchmark cases. Each case includes process specs, prompts… See the full description on the dataset page: https://huggingface.co/datasets/Xin-Rui/ImageTime_Benchmark.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,355,161 human responses, collected with the
Rapidata Python SDK, comparing how well 30 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.geofm-agriculture-benchmark
GeoFM Agriculture Benchmark
Sample data and fine-tuned weights accompanying our ACM SIGSPATIAL 2026 paper, released so
other researchers can run inference with SatMAE, Prithvi, and SpectralGPT on our
multi-temporal crop segmentation and change-detection tasks.
This is not the full training dataset — it's a set of representative chips per region/model
plus the fine-tuned checkpoints, enough to run and sanity-check inference end-to-end.
Contact the authors if you need the complete… See the full description on the dataset page: https://huggingface.co/datasets/sanmay4119/geofm-agriculture-benchmark.artist-style-benchmark
Artist Style Benchmark
A benchmark dataset of 34903 anime-style illustrations generated with
AnimaImagine, each using a
different Danbooru artist tag at fixed prompt/seed settings.
ASR Ranker Style Browser
This dataset also serves as the image host for
ASR Ranker, an interactive browser for comparing and
ranking artist styles.
Browser: https://ranker.kuronet.top/browser
Dataset: https://huggingface.co/datasets/Moeblack/artist-style-benchmark
选出最 hot 的画风吧。
总榜只有 trusted… See the full description on the dataset page: https://huggingface.co/datasets/Moeblack/artist-style-benchmark.ppe-benchmark-eval
PPE Benchmark Eval Set (v1)
A held-out, human-verified benchmark for evaluating vision-language models on
personal protective equipment (PPE) detection — specifically hardhat and
safety-vest presence — framed as a VQA-style classification task.
What this is
96 images, balanced 24/24/24/24 across the four hardhat × vest combinations
(yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of
the karabuk-university PPE dataset
on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.Reference-Update-Benchmark
COVER-Fish Reference-Update Benchmark
This benchmark studies how reference-corpus updates change a frozen recognition
system, instantiated on fine-grained fish identification. It publishes the
versioned control plane behind COVER-Fish: row-level manifests, gallery states,
taxonomy, tensor bindings, transition evidence, protocols and dependency locks.
The benchmark does not duplicate the 83 GB Full Payload Archive. Large source
archives and frozen tensors remain in the immutable… See the full description on the dataset page: https://huggingface.co/datasets/COVER-Fish/Reference-Update-Benchmark.scope-benchmark
SCOPE Benchmark
Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?
Contents
scope-benchmark/
scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.GeneLab_BPS_BenchmarkData
Dataset Card for Dataset GeneLab_BPS_BenchmarkData
Dataset Details
This dataset is a version of the Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset managed by NASA and hosted on an S3 Bucket here: https://registry.opendata.aws/bps_microscopy/
Fluorescence microscopy images of individual nuclei from mouse fibroblast cells, irradiated with Fe particles or X-rays with fluorescent foci indicating 53BP1 positivity, a marker of DNA damage.… See the full description on the dataset page: https://huggingface.co/datasets/kenobi/GeneLab_BPS_BenchmarkData.urban-perception-benchmark
Urban Perception Benchmark
Pretty name: Urban Perception Benchmark — Montreal 100Short name: UPB-MTL100License (data): CC BY-NC 4.0 (non-commercial)License (code): MITLanguages: French (source), English (normalized)Modalities: Images + structured annotationsSize: 100 images (50 synthetic, 50 real)Tasks: multi-label and single-choice annotation; evaluation of VLMs on urban perception
This repository hosts the dataset and annotation schema described in the paper:“Do Vision–Language… See the full description on the dataset page: https://huggingface.co/datasets/rsdmu/urban-perception-benchmark.diabetic-retinopathy-screening-benchmark-africa
DR-Africa-Benchmark — Screening-Prevalence-Corrected, Fairness-Instrumented DR Evaluation
An evaluation benchmark for diabetic-retinopathy grading under African
screening conditions. It does not introduce new labels; it introduces
evaluation validity — per-record importance weights that reweight a
referral-skewed image set to real Sub-Saharan-Africa population prevalence, plus
synthetic subgroup metadata for fairness reporting.
Version 1.0.0 · core dr_synth 1.0.0 · part of the… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-screening-benchmark-africa.willie-benchmark
WILLIE Wound Benchmark
Three public wound datasets unified into a single 5-class taxonomy with
fixed splits for classification, segmentation and localization.
The benchmark accompanying WILLIE, published at MLHC 2026.
Developed in the Qian Group, University of Houston.
Models: QianGroup/willie-weights
Code and notebooks: GitHub repository
Paper: MLHC 2026 (link to follow)
What this is
Three public wound datasets — FUSeg, AZH and Medetec — mapped onto one… See the full description on the dataset page: https://huggingface.co/datasets/QianGroup/willie-benchmark.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.ai-detector-benchmark-test-data
🎯 AI Detector Benchmark Test Dataset
A comprehensive benchmark dataset for testing AI image detection models.
📊 Dataset Summary
Total Images: 700
AI-Generated: 250 images (from 5 different generators)
Real Images: 450 images (from 9 diverse datasets)
Perfect for:
✅ Testing AI detection models
✅ Creating leaderboards
✅ Comparing model performance
✅ Benchmarking new approaches
🤖 AI Generators Included
Generator
Images
Accuracy Baseline
FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.willie-benchmark
WILLIE Wound Benchmark
Three public wound datasets unified into a single 5-class taxonomy with
fixed splits for classification, segmentation and localization.
The benchmark accompanying WILLIE, published at MLHC 2026.
Developed in the Qian Group, University of Houston.
Models: QianGroup/willie-weights
Code and notebooks: GitHub repository
Paper: MLHC 2026 (link to follow)
What this is
Three public wound datasets — FUSeg, AZH and Medetec — mapped onto one… See the full description on the dataset page: https://huggingface.co/datasets/CnGIndia/willie-benchmark.Face_Generation_Benchmark
Rapidata Human Face Generation Alignment
This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating 12 different image generation models on which one can generate faces more accurately.
The question that the annotators get asked is: "Which Image follows the description of the human better?"
To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.benchmark
EditJudge-Bench
EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as
automated judges for image-edit verification. Each row contains a source image,
an edited image, a factual edit instruction, counterfactual instructions, and
ground-truth scene parameters produced by a controlled Blender/Infinigen
generation pipeline.
This repository is an anonymous review release for a NeurIPS Evaluations and
Datasets submission.
Dataset Contents
1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.nra-benchmarks
🧬 NRA Benchmark Datasets
All benchmark datasets for Neural Ready Archive (NRA) — the Rust-native streaming format for ML training.
Train on gigabytes of real data without downloading a single byte. NRA replaces tar.gz and zip for the AI era.
📦 Available Datasets
File
Domain
Source
Files
Size
food-101.nra
🖼️ Vision
ethz/food101
101,000 images
4.7 GB
wikitext.nra
📝 Text
Salesforce/wikitext
23,767 text files
7.6 MB
pokemon.nra
🎨 Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/zevatov/nra-benchmarks.noble-ai-evidence-benchmark
NOBLE AI-Generated Evidence Detection Benchmark
A domain-specific benchmark for evaluating AI-generated image detection tools on law enforcement imagery (surveillance footage, bodycam, evidence-style photos). The benchmark spans multiple generator architectures and three image-quality levels designed to mimic the conditions in which real evidence reaches courtrooms.
Status: v1.1 release. All three generators (FLUX-schnell, Realistic Vision 5.1, SDXL) complete, paired-prompt design… See the full description on the dataset page: https://huggingface.co/datasets/ashleyscruse/noble-ai-evidence-benchmark.
