datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
screenparse
ScreenParse: Large-Scale Dataset for Complete Screen Parsing
News
May 2026: ScreenParse v2 is released on main with more robust quality filtering, varied viewport resolutions, leaf-element annotations that reduce annotation noise, and 1,447,100 high-quality training screenshots. The first release is retained on the v1 branch.
Dataset Description
ScreenParse is a large-scale dataset for complete screen parsing, providing dense annotations of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/screenparse.document-haystack-10pages
Dataset Card for document-haystack-10pages
This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/document-haystack-10pages")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.noisy-medical-document-images-ocr
🏥 Noisy Medical Document Images (OCR)
1,000 noisy, synthetic medical document images with structured JSON ground truth — built for Document AI, LayoutLM fine-tuning, and clinical NLP research.
🧭 Overview
This dataset provides 1,000 high-resolution images of two healthcare document types, each degraded with realistic scanning artifacts to simulate real-world OCR conditions:
Category
Count
Description
🧾 Hospital Bills
500
Itemized statements with… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/noisy-medical-document-images-ocr.icdar2021-historical-document-dating
ICDAR 2021 Historical Document Classification — Task 2 (Dating)
13,810 manuscript page images labelled with the period in which they were produced.
Images come from e-codices, the virtual manuscript library
of Switzerland.
Split
Images
Date range
Median span
Dated to a single year
train
11,294
800–1899
45 years
1,409
test
2,516
800–1921
49 years
264
The label is an interval, not a year
Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.plant_doc_classification
Plant Doc Classification
A dataset for disease classification of various plants. The dataset contains 2,569 images across 28 classes:Images per class:
Apple Scab Leaf: 93
Apple leaf: 91
Apple rust leaf: 88
Bell_pepper leaf: 61
Bell_pepper leaf spot: 71
Blueberry leaf: 114
Cherry leaf: 57
Corn Gray leaf spot: 68
Corn leaf blight: 191
Corn rust leaf: 116
Peach leaf: 111
Potato leaf early blight: 116
Potato leaf late blight: 105
Raspberry leaf: 119
Soyabean leaf: 65
Squash Powdery… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/plant_doc_classification.doctype
Document Classification
It works per image input.
Classes
Class
Processing Pipeline
Training Datasets
document_printed
Docling pipeline
HuggingFaceFW/finepdfs, RVL-CDIP
document_handwritten
Docling pipeline with handwriting OCR
IAM Handwritten Forms, RVL-CDIP
photo
VLLM: "Describe this image"
COCO 2017
diagram
VLLM: "Describe this diagram"
AI2D, UML Diagrams
chart
VLLM: "Analyze this chart/graph"
PlotQA
screenshot
VLLM: "Analyze screenshot"
RICO… See the full description on the dataset page: https://huggingface.co/datasets/monkt/doctype.synthetic-australian-medical-documents-sample
Synthetic Australian Medical Documents - Sample
A 50-document free sample of a 5,000-document library of synthetic Australian medical PDFs. PHI-free. Modelled on Australian healthcare documentation. Pre-labelled with structured ground truth and pixel-precise bounding boxes. Released under CC-BY-NC 4.0 for evaluation and non-commercial research.
See Pricing & licensing below.
What's in this sample
Field
Value
Documents
50
Document types
29 (of 45 in full… See the full description on the dataset page: https://huggingface.co/datasets/RootCauseAnalytics/synthetic-australian-medical-documents-sample.rvl-cdip-document-classification
rvl-cdip-document-classification
This dataset is created from original aharley/rvl_cdip dataset using this notebook
Dataset Summary
This dataset consists of 8992 grayscale images in 16 classes, with 562 images per class.
There are 8000 training images(500 image per class) and 992 test images(62 images per class).
The images are sized so their largest dimension does not exceed 1000 pixels.
Document-Type-Detection
Document-Type-Detection
Dataset Summary
The Document-Type-Detection dataset is a large-scale image classification dataset consisting of scanned or photographed document images. Each image is categorized into one of nine document types. This dataset is ideal for training document classification models in finance, administration, OCR, and automation workflows.
Supported Tasks
Multiclass Document Classification
Classify an input document image into one of the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Document-Type-Detection.CMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.muninn-ww1-documents
Muninn WWI Documents (CEF Attestation Papers & War Diaries)
A tabular conversion of the document records in the Muninn Project's
World War I Linked Open Data archive (https://rdf.muninn-project.org/). The Muninn Project is a
research project that extracts structured data from digitized WWI-era archival documents.
The dataset covers 28,742 documents — 28,718 Canadian Expeditionary Force (CEF) attestation
papers
(enlistment forms of individual soldiers) and 24 Canadian unit war… See the full description on the dataset page: https://huggingface.co/datasets/biglam/muninn-ww1-documents.document-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.siglip-doc-understanding-classifier
SigLIP Doc Understanding — Unanswerable Question Detection Dataset
A mixed answerable / unanswerable benchmark dataset built from DocVQA and MP-DocVQA, used to
train and evaluate the siglip-doc-understanding-classifier
unanswerable-question detector.
Each row pairs a document image with a question. Half of the questions are the original,
answerable DocVQA/MP-DocVQA questions; the other half are corrupted versions of those same
questions — modified so the document image no longer… See the full description on the dataset page: https://huggingface.co/datasets/giacolees/siglip-doc-understanding-classifier.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.doc-openvocab-benchmark
Open-Vocab Document & Figure Classification Benchmark
Given a document or figure image and an arbitrary set of text labels, which one is right? This is a
zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is
scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task
is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label
supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.doc-content-clustering-740
Document Content-Clustering Benchmark (12 classes, 740 items)
A small, curated benchmark for clustering documents by their content topic
(not by their visual form/layout). Each item is a single document page provided
as an image plus two text views (a VLM description and OCR markdown), with a
ground-truth content class.
The set is intentionally "tangle-stripped": 27 borderline items whose content
sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.doclang
DocLang: Multilingual Document Type Classification
A multilingual document type classification dataset for identifying various document and visual content types. The dataset contains 13,200 images across 11 languages, with 1,200 images per language.
Dataset Structure
The dataset is organized by language code:
ar/ - Arabic (1,200 images)
bg/ - Bulgarian (1,200 images)
de/ - German (1,200 images)
en/ - English (1,200 images)
es/ - Spanish (1,200 images)
fr/ - French (1,200… See the full description on the dataset page: https://huggingface.co/datasets/monkt/doclang.docdet-scamai-crops
tzj04/docdet-scamai-crops
Training crops derived from the Scam-AI document-forgery datasets, for the
DocDet authentic-vs-AI-generated detector.
This is a derivative work. It is not an official Scam-AI release.
What a row is
Each forgery in the source data patches a single field into an otherwise
authentic scan - roughly 0.3% of the page. At a 224px whole-page input that
edit survives as a handful of pixels, and a random-resized crop can miss it
altogether. So… See the full description on the dataset page: https://huggingface.co/datasets/tzj04/docdet-scamai-crops.image-in-Words400_DOCCI_Test
