datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DocDownstream-2.0DocDownstream 2.0 is a collection of MP-DocVQA, DUDE, NewsVideoQA used in DocOwl2.
For MP-DocVQA and DUDE, the maximum number of pages of each sample is set to 20.
For NewsVideoQA, the maximum number of frames of each sample is set to 20.
VL-DocIR
Abstract
VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.RxnBench-Doc
RxnBench-Doc: A Benchmark for Multimodal Understanding of Chemistry Reaction Literature
News: This work has been accepted by Journal of Chemical Information and Modeling
📘 Benchmark Summary
RxnBench (FD-QA) is a document-level question answering (DocQA) benchmark comprising 540 multiple-select questions designed to assess PhD-level understanding of organic chemistry reactions in textual and multimodal contexts. All questions underwent multiple rounds of expert… See the full description on the dataset page: https://huggingface.co/datasets/UniParser/RxnBench-Doc.docflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.technical_documents
Technical Documents: Shafts Process (v1, PNG)
license: cc-by-4.0
pretty_name: Technical Documents: Shafts Process (v1, PNG)
language:
- en
task_categories:
- visual-question-answering
size_categories:
- 1K<n<10K
annotations_creators:
- machine-generated
source_datasets:
- original
~1000 image→text pairs (stepped shafts → machining process)
Fixed 10–11 step template; right-side Z=0; tailstock rule (L/D>3 or total length > 190 mm)Images converted from SVG to PNG… See the full description on the dataset page: https://huggingface.co/datasets/meatfly/technical_documents.MP-DocReason51KMP-DocReason51K is a Multi-image instruction tuning dataset on OCR-free Document Understanding used in DocOwl2.
Each answer comprises a concise answer, a reference image, and detailed explanation.
DocGenome12KDocMIDE
DocMIDE
Document images paired with question/answer pairs that require implicit reasoning over the document — deriving, cross-referencing, or computing a value — rather than plain OCR transcription. This is the held-out test split used by the DocMIDE Benchmark eval harness; the reasoning-annotated training splits (with raw field data and derivation traces, used for SFT/GRPO) live in the DocMIDE training repo.
Most document images are real-world documents sourced from the Unikie… See the full description on the dataset page: https://huggingface.co/datasets/VXRealLimited/DocMIDE.cdnVL-DocIR-RepresentativeSubset
Representative Subset Creation
This dataset represents a representative subset of VL-DocIR dataset (https://huggingface.co/datasets/anonymous-8421/VL-DocIR).
Subset creation method:
Random sampling of 5 queries for the combination of each data source and evidence structure type (we ensure that no query is taken more than once).
This results in 120 queries.
Collection of all documents that are referenced by query evidence.
Abstract
VL-DocIR is a page-level benchmark… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR-RepresentativeSubset.ml-design-doc-reviewer-data
ml-system-design/ml-design-doc-reviewer-data (v1.0.0)
Evaluation artifacts for the ML Design Doc Reviewer project.
Layout
Path
Description
manifest/sample_manifest.csv
Stratified 100-case sample manifest
manifest/error_topology.csv
Controlled error taxonomy for flawed docs
raw/
Raw markdown exports, metadata sidecars, OCR image blocks
raw/images/
Downloaded article images
normalized/
Canonical 14-section ML design documents
flawed/
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.MMLongBench_Doc_val_500_subsetdoctor-adventures
