datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DocumentVQAwikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.ocr-document-processing-eval
ocr_document_processing_eval
Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks.
Repo: himalaya-ai/ocr-document-processing-eval
Task: document_processing_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.icdar2021-historical-document-dating
ICDAR 2021 Historical Document Classification — Task 2 (Dating)
13,810 manuscript page images labelled with the period in which they were produced.
Images come from e-codices, the virtual manuscript library
of Switzerland.
Split
Images
Date range
Median span
Dated to a single year
train
11,294
800–1899
45 years
1,409
test
2,516
800–1921
49 years
264
The label is an interval, not a year
Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.document-visual-retrieval-test
Model Card: Document Visual Retrieval Test (internal)
Dataset Overview
This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.test-document-invoiceDocumentVQA
Neural Metrics · Asking documents questions, and grading the answers.
Document visual question answering: given a page image and a natural-language question, produce the answer. It measures whether a model genuinely read the layout or merely pattern-matched the text.
We use it for: evaluating question-answering over extracted documents - catching models that read text but misread structure.
Attribution
This is an unmodified fork of… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/DocumentVQA.khmer-document-synthetic-low-reswikimedia-commons-documents-ml_deprecated
Wikimedia Commons Document Retrieval
Wikimedia Commons Documents
This dataset is created for the evaluation of retrieval models. It contains images of (mostly historic) documents which should be identified based on their description. We extracted those descriptions from Wikimedia Commons. We have included the license type and a link (license_text) to the original Wikimedia Commons page for each extracted image.
The text_description column contains OCR text extracted from the images… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_deprecated.example-documentsTranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.sroie_document_understanding
Dataset Card for "sroie_document_understanding"
Dataset Description
This dataset is an enriched version of SROIE 2019 dataset with additional labels for line descriptions and line totals for OCR and layout understanding.
Dataset Structure
DatasetDict({
train: Dataset({
features: ['image', 'ocr'],
num_rows: 652
})
})
Data Fields
{
'image': PIL Image object,
'ocr': [
# text box 1
{
'box':… See the full description on the dataset page: https://huggingface.co/datasets/arvindrajan92/sroie_document_understanding.rvl-cdip-document-classification
rvl-cdip-document-classification
This dataset is created from original aharley/rvl_cdip dataset using this notebook
Dataset Summary
This dataset consists of 8992 grayscale images in 16 classes, with 562 images per class.
There are 8000 training images(500 image per class) and 992 test images(62 images per class).
The images are sized so their largest dimension does not exceed 1000 pixels.
document-parts
Dataset Card for document-parts
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
document-parts
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/document-parts.DocumentVQAMAVIS_documentsDocument-Type-Detection
Document-Type-Detection
Dataset Summary
The Document-Type-Detection dataset is a large-scale image classification dataset consisting of scanned or photographed document images. Each image is categorized into one of nine document types. This dataset is ideal for training document classification models in finance, administration, OCR, and automation workflows.
Supported Tasks
Multiclass Document Classification
Classify an input document image into one of the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Document-Type-Detection.kyc-document-extraction-vlmCMDS_Multimodal_Document
Dataset Card for Cyrillic Multimodel Document (CMDS)
This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.myanmar_complex_document_layouts
🇲🇲 Myanmar Complex Document Layouts
A large-scale, high-quality synthetic dataset containing 17,632 images of complex document layouts, dashboards, and infographics entirely in the Myanmar (Burmese) language.
This dataset is specifically designed to train and benchmark modern Computer Vision and multimodal LLMs on complex Myanmar typography, structured data, and diverse graphical layouts.
📊 Dataset Overview
Total Images: 17,632 high-resolution pages.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_complex_document_layouts.document-classification-benchmark
Document Classification Benchmark (open-vocab, zero-shot)
Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot,
open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked
into a head. Test split only; not for training. Every image is drawn from a permissively-licensed,
redistributable source.
Powers the
document-classification-leaderboard
and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.samyx-document-sercolpali_italian_documentskyc-document-extraction-vlmolmOCR-mix-0225-documents_cleaned
olmOCR-mix-0225-documents_cleaned
The olmOCR-mix-0225-documents__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
156,730
QA turns
569,306
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
127
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/olmOCR-mix-0225-documents_cleaned.sifta-document-forgery-datasetOMR-scanned-documentsA medical forms dataset containing scanned documents is a valuable resource for healthcare professionals, researchers, and institutions seeking to streamline and improve their administrative and patient care processes. This dataset comprises digitized versions of various medical forms, such as patient intake forms, consent forms, health assessment questionnaires, and more, which have been scanned for electronic storage and easy access.
These scanned medical forms preserve the layout and… See the full description on the dataset page: https://huggingface.co/datasets/saurabh1896/OMR-scanned-documents.DocumentIDEFICS_VQADocumentIDEFICS_QA
