datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.clamp-benchmark
CLAMP: A Sim-to-Real Benchmark for Closed-Loop Kinematic Pose Estimation and Assembly Reasoning
Closed-Loop Assembly and Mechanism Perception
📦 Code: https://anonymous.4open.science/r/clamp-131F
Your browser does not support the video tag.
Overview of all 210 labeled real test scenes.
Overview of our benchmark. Top left: A robot arm with its kinematic graph overlaid—open-chain edges in cyan, closed-loop edges in orange. Removing a constraining link (red oval) breaks the… See the full description on the dataset page: https://huggingface.co/datasets/clamp-benchmark/clamp-benchmark.Co-Spy-Bench
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI (CVPR 2025)
With the rapid advancement of generative AI, it is now possible to synthesize high-quality images in a few seconds. Despite the power of these technologies, they raise significant concerns regarding misuse.
To address this, various synthetic image detectors have been proposed. However, many of them struggle to generalize across diverse generation parameters and emerging generative models.
In… See the full description on the dataset page: https://huggingface.co/datasets/ruojiruoli/Co-Spy-Bench.gently-perception-benchmark
Gently Perception Agent Benchmark
Light-sheet microscopy volumes of C. elegans embryo development, intended
for evaluating vision-based perception agents on embryo stage classification.
The dataset has two tiers:
Annotated benchmark set (embryo_1–embryo_8) — human ground-truth
stage transitions. Use this for evaluation.
Unannotated corpus (embryo_9–embryo_105) — 97 additional real embryo
timelapses with no human labels, provided for developing and stress-
testing perception… See the full description on the dataset page: https://huggingface.co/datasets/gently-project/gently-perception-benchmark.DamageTriage-Bench
DamageTriage-Bench
DamageTriage-Bench is a footprint-conditioned benchmark for per-building damage
typing from single post-event aerial images. Its five classes distinguish roof
from structural damage and partial from total affected extent:
ID
Class
0
Undamaged
1
Partial Roof Damage
2
Total Roof Damage
3
Partial Structural Damage
4
Total Structural Collapse
Quick statistics
Item
Value
Tiles
7,472 (1024 × 1024 PNG)
Labeled… See the full description on the dataset page: https://huggingface.co/datasets/Ymx1025/DamageTriage-Bench.MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.Obshazard-bench
ObsCrisis-Bench
A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks.
Dataset Description
ObsCrisis-Bench contains 4,202 VQA samples across 127 extreme weather events in 8 disaster categories, covering 61 countries. Each sample combines satellite multispectral imagery (AMSU-A, HIRS, MHS sensors) with optional weather station data, and requires models to perform risk assessment, type classification, timing… See the full description on the dataset page: https://huggingface.co/datasets/YYQ898/Obshazard-bench.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.Copernicus-Bench
Dataset Card for Copernicus-Bench
A hierarchical ML benchmark for Copernicus Sentinels, with 15 datasets spread into three task levels covering all major Sentinel missions (S1,2,3,5P).
(Officially named "Copernicus-Bench", initially named "SentinelBench")
Dataset Details
Level
Name
Modality
Task
# Images
Image Size
# Classes
Source
License
L1
Cloud-S2
S2 TOA
segmentation (cloud)
1699/567/551
512x512x13
4
CloudSEN12
CC 0 1.0
L1
Cloud-S3
S3 OLCI… See the full description on the dataset page: https://huggingface.co/datasets/wangyi111/Copernicus-Bench.CTTA-AD-Benchmarks
CTTA-AD Benchmarks
Dataset collection for CTTA-AD: Continual Test-Time Adaptation for Unified Few-Shot Visual Anomaly Detection (AAAI 2027 submission).
Datasets
Dataset
Domain
Categories
Train Normal
License
MVTec-AD
Industrial
15
209–391 per category
CC BY-NC-SA 4.0
VisA
Industrial
12
400–905 per category
CC BY-NC-SA 4.0
MVTec-LOCO
Logical
5
varies
CC BY-NC-SA 4.0
BrainMRI
Medical
1
7,500
Research only
LiverCT
Medical
1
1,542
Research only… See the full description on the dataset page: https://huggingface.co/datasets/Hammadhaideerr/CTTA-AD-Benchmarks.holisafe-bench
⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised.
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings)
🌐 Website | 📑 Paper
📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.SA-BENCH
SA-BENCH
SA-BENCH is the benchmark dataset released with “Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics.”
Accepted to CVPRW 2026.
GitHub | CVF Open Access | arXiv | Model
It evaluates the spatial aesthetics of interior images along four dimensions:
distortion
harmony
layout
lighting
SA-BENCH contains 17,768 annotated examples across four spatial-aesthetic dimensions, with image assets and human annotations for training and… See the full description on the dataset page: https://huggingface.co/datasets/gaoyuan-ai/SA-BENCH.CoVAtt-Benchmark
CoVAtt-Benchmark
A large-scale benchmark for generated-image attribution: given an image
that was produced by some text-to-image model, decide which model
produced it, and decide whether it came from a model the system has never
seen before.
What this dataset is
This is the dataset used to train and evaluate CoVAtt
(Content-Based Verification for Attribution of AI-Generated Images,
BMVC 2026). CoVAtt is a Siamese network that takes a pair of images and
predicts… See the full description on the dataset page: https://huggingface.co/datasets/kenyag/CoVAtt-Benchmark.PRISM_Benchmark
PRISM: PhotoRealistic Image Synthesis and Manipulation
Dataset repository for the paper "The PRISM benchmark: PhotoRealistic Image Synthesis and Manipulation to detect generated images" — Bartolucci, Salti, Lisanti.
Real images
The corresponding real images can be downloaded separately from COCO:
Split
Source
Test set
COCO 2017 Val
Training set
COCO 2017 Train
Citation
@article{BARTOLUCCI2026104826,
title = {The PRISM… See the full description on the dataset page: https://huggingface.co/datasets/oppiliF/PRISM_Benchmark.ImageTime_Benchmark
ImagineTime Benchmark
This dataset repository contains the public benchmark assets for ImagineTime, released with the paper “Can Image Models Imagine Time?”
Paper: arXiv:2606.10620
ImagineTime evaluates whether image generation models can produce ordered 2x2 motion sheets with coherent entities, spatial relations, state transitions, interactions, and task constraints.
Contents
cases/
750 benchmark cases. Each case includes process specs, prompts… See the full description on the dataset page: https://huggingface.co/datasets/Xin-Rui/ImageTime_Benchmark.PCF-Bench
PCF-Bench
A photonic-crystal-fiber (PCF) inverse-design benchmark for vision-language models. Each sample bundles geometric parameters, simulated mode-field images, and a four-level expert-style annotation suite supporting tasks across geometry perception, physics understanding, multimodal reasoning, inverse design, and code generation.
Note (anonymous review). This dataset card omits identifying information for double-blind review.
Note (review subset). The full source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/PCF-Bench/PCF-Bench.GeoFidelity-Bench
GeoFidelity-Bench
GeoFidelity-Bench evaluates whether generated street-view images match a
requested location at the level of named street blocks. The release contains
109 named street blocks from 25 cities, 7,117 curated Mapillary reference
images, generated images from six open-weight text-to-image models, prompt
control metadata, and 109-target benchmark result summaries. The generated-image index covers
15,696 released JPEG files across six models, six prompt or control… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,355,161 human responses, collected with the
Rapidata Python SDK, comparing how well 30 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.MWS-Antifraud-Bench
MWS Antifraud Bench (Validation)
Experimental document-authenticity task for general-purpose multimodal language
models. This is the public validation part of MWS Vision Bench anti-fraud v0.1.
The dataset is released for research and model comparison. It is not a
certification tool, a production fraud-detection system, or a universal
leaderboard that is expected to be resistant to deliberate optimization.
Data
The validation split contains 209 items:
44 ai_gen;… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/MWS-Antifraud-Bench.SFHQ-VirtualID-Bench
SFHQ-VirtualID-Bench Dataset Card
Summary
SFHQ-VirtualID-Bench is a synthetic, identity-conditioned face dataset of
750 synthetic identities, built as a benchmark for machine unlearning.
Each identity is a deletion unit: all 90 aligned 224×224 crops of one
identity share a single identity-level split (retain or forget), and 75
identities follow a sequential 15-step forgetting protocol. Every identity
contributes both per-image image_subset values (train + holdout)… See the full description on the dataset page: https://huggingface.co/datasets/FaizPalwala/SFHQ-VirtualID-Bench.geofm-agriculture-benchmark
GeoFM Agriculture Benchmark
Sample data and fine-tuned weights accompanying our ACM SIGSPATIAL 2026 paper, released so
other researchers can run inference with SatMAE, Prithvi, and SpectralGPT on our
multi-temporal crop segmentation and change-detection tasks.
This is not the full training dataset — it's a set of representative chips per region/model
plus the fine-tuned checkpoints, enough to run and sanity-check inference end-to-end.
Contact the authors if you need the complete… See the full description on the dataset page: https://huggingface.co/datasets/sanmay4119/geofm-agriculture-benchmark.auto-research-bench-data
auto-research-bench-data
同行评审语料 + 论文 PDF,覆盖五个机器学习会议 2022–2026 年。
用于研究「AI 生成的论文与人类论文有何差异」,特别是图表质量与实验管理两个维度。
数据来自 OpenReview,用官方 API 抓取。
内容
1. 评审元数据(18 个 jsonl,3.11 GB)
每行一篇投稿,字段如下:
字段
说明
forum
OpenReview 论文 ID,与 PDF 文件名一致,是关联两部分数据的 key
title / abstract / keywords
论文元信息
venue / venueid
录用层级。注意层级只在 venue 里(如 ICLR 2024 oral),venueid 对所有录用论文都是 .../Conference
reviews
全部 Official_Review,含评分、置信度、正文各字段
comments
作者 rebuttal 与其他… See the full description on the dataset page: https://huggingface.co/datasets/chengwanru/auto-research-bench-data.HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.locus-tag-bench
Locus-Tag Bench
Synthetic benchmark suite for fiducial-tag detection and camera calibration, rendered with render-tag.
Each config corresponds to one campaign (board family × resolution, or a lighting/sensor variant). Images are rendered in Blender Cycles with ground-truth geometry recovered directly from the scene — no detector-in-the-loop, no human labels.
Configs
Config
Purpose
Images
Board/Tag
Resolution
locus_v1_tag36h11_640x480
Detection (low-res)… See the full description on the dataset page: https://huggingface.co/datasets/NoeFontana/locus-tag-bench.artist-style-benchmark
Artist Style Benchmark
A benchmark dataset of 34903 anime-style illustrations generated with
AnimaImagine, each using a
different Danbooru artist tag at fixed prompt/seed settings.
ASR Ranker Style Browser
This dataset also serves as the image host for
ASR Ranker, an interactive browser for comparing and
ranking artist styles.
Browser: https://ranker.kuronet.top/browser
Dataset: https://huggingface.co/datasets/Moeblack/artist-style-benchmark
选出最 hot 的画风吧。
总榜只有 trusted… See the full description on the dataset page: https://huggingface.co/datasets/Moeblack/artist-style-benchmark.MFC-Bench
MFC-Bench: Multimodal Fact-Checking Benchmark
MFC-Bench is a comprehensive Multimodal Fact-Checking testbed designed to evaluate LVLMs in terms of identifying factual inconsistencies and counterfactual scenarios.
Dataset Description
From the paper: "MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models"
MFC-Bench encompasses a wide range of visual and textual queries, organized into three binary classification tasks:
1. Manipulation… See the full description on the dataset page: https://huggingface.co/datasets/MM-Hallu/MFC-Bench.HUGO-Bench
HUGO-Bench
Hierarchical Unsupervised Grouping of Organisms Benchmark
A comprehensive benchmark dataset for evaluating zero-shot clustering of wildlife camera trap images using Vision Transformer embeddings.
Overview
HUGO-Bench contains 139,111 expert-validated cropped images of 60 animal species (30 birds, 30 mammals), derived from 23 camera trap projects across LILA BC. The dataset enables benchmarking of Vision Transformer models for unsupervised species-level… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench.ppe-benchmark-eval
PPE Benchmark Eval Set (v1)
A held-out, human-verified benchmark for evaluating vision-language models on
personal protective equipment (PPE) detection — specifically hardhat and
safety-vest presence — framed as a VQA-style classification task.
What this is
96 images, balanced 24/24/24/24 across the four hardhat × vest combinations
(yes/yes, yes/no, no/yes, no/no). Sourced from a forked, filtered subset of
the karabuk-university PPE dataset
on Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/khadijah00/ppe-benchmark-eval.uijudge-bench
UIJudgeBench v0.4.0 (pre-release)
UIJudgeBench evaluates systems that judge web UI quality from frozen, versioned page
artifacts with machine-checkable ground truth. It covers accessibility, layout,
referring/computed-style questions, and a separate pairwise design-quality instrument.
This is the dataset release. The independently versioned Python harness is available on
PyPI, and the canonical source release is
GitHub v0.4.0. The PyPI
wheel intentionally does not bundle this… See the full description on the dataset page: https://huggingface.co/datasets/gojiberries/uijudge-bench.
