datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,355,161 human responses, collected with the
Rapidata Python SDK, comparing how well 30 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/FForty7/svg-benchmark.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.ai-detector-benchmark-test-data
🎯 AI Detector Benchmark Test Dataset
A comprehensive benchmark dataset for testing AI image detection models.
📊 Dataset Summary
Total Images: 700
AI-Generated: 250 images (from 5 different generators)
Real Images: 450 images (from 9 diverse datasets)
Perfect for:
✅ Testing AI detection models
✅ Creating leaderboards
✅ Comparing model performance
✅ Benchmarking new approaches
🤖 AI Generators Included
Generator
Images
Accuracy Baseline
FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.diabetic-retinopathy-screening-benchmark-africa
DR-Africa-Benchmark — Screening-Prevalence-Corrected, Fairness-Instrumented DR Evaluation
An evaluation benchmark for diabetic-retinopathy grading under African
screening conditions. It does not introduce new labels; it introduces
evaluation validity — per-record importance weights that reweight a
referral-skewed image set to real Sub-Saharan-Africa population prevalence, plus
synthetic subgroup metadata for fairness reporting.
Version 1.0.0 · core dr_synth 1.0.0 · part of the… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-screening-benchmark-africa.benchmark
EditJudge-Bench
EditJudge-Bench is a synthetic benchmark for auditing vision-language models used as
automated judges for image-edit verification. Each row contains a source image,
an edited image, a factual edit instruction, counterfactual instructions, and
ground-truth scene parameters produced by a controlled Blender/Infinigen
generation pipeline.
This repository is an anonymous review release for a NeurIPS Evaluations and
Datasets submission.
Dataset Contents
1… See the full description on the dataset page: https://huggingface.co/datasets/EDAnonSubmission/benchmark.Face_Generation_Benchmark
Rapidata Human Face Generation Alignment
This T2I dataset contains over ~22'000 human responses, collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating 12 different image generation models on which one can generate faces more accurately.
The question that the annotators get asked is: "Which Image follows the description of the human better?"
To evaluate your own models and create leaderboard check out our… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Face_Generation_Benchmark.doc-openvocab-benchmark
Open-Vocab Document & Figure Classification Benchmark
Given a document or figure image and an arbitrary set of text labels, which one is right? This is a
zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is
scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task
is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label
supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.tibetan-script-classification-benchmark
Tibetan Script Classification Benchmark
Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training.
All images are BDRC manuscript page scans, balanced by subclass.
Class
Images
Subclasses
Danyig
60
DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6
Druma
60
Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6
Gyuyig
60
Khyuyig: 31, Tsumachug: 15, Yigchung: 14
Pedri
60
Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.ai-detector-benchmark-test-data
🎯 AI Detector Benchmark Test Dataset
A comprehensive benchmark dataset for testing AI image detection models.
📊 Dataset Summary
Total Images: 700
AI-Generated: 250 images (from 5 different generators)
Real Images: 450 images (from 9 diverse datasets)
Perfect for:
✅ Testing AI detection models
✅ Creating leaderboards
✅ Comparing model performance
✅ Benchmarking new approaches
🤖 AI Generators Included
Generator
Images
Accuracy Baseline
FLUX… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/ai-detector-benchmark-test-data.sheep-facial-expression-benchmark
Sheep Facial Expression Benchmark
Prepared OpenFARM sheep facial-expression benchmark data from the public Mendeley Data record 10.17632/y5sm4smnfr.5.
Source
Source dataset: https://data.mendeley.com/datasets/y5sm4smnfr
Source DOI: 10.17632/y5sm4smnfr.5
Related paper DOI: 10.1016/j.compag.2020.105528
License: CC BY 4.0
Splits
{
"train": 172,
"test": 74,
"train_raw": 898,
"test_raw": 225
}
train and test are filtered/balanced views for benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/sheep-facial-expression-benchmark.multi-species-benchmark
multi-species benchmark
Photographs where 2+ species appear in the same frame. Designed to evaluate
multi-label species identification and steering capabilities of biological
vision-language models. Two sources, unified into one parquet schema.
Sources
inat21_multilabel (299 rows, 147 images)
In-distribution: drawn from iNat21
validation images that already carry an iNat-supplied primary label. We use
InternVL3-AWQ to surface
images that also… See the full description on the dataset page: https://huggingface.co/datasets/dcher95/multi-species-benchmark.
