datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docs-imagesdiffusers-images-docspixmo-docs
PixMo-Docs
We now recommend using CoSyn-400k and CoSyn-point over these
datasets. They are improved versions with more images categories and an improved generation pipeline.
PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
DocStruct4MVDocRetriever-Pretrain-DocStructDocStruct4MmPLUG/DocStruct4M reformated for VSFT with TRL's SFT Trainer.Referenced the format of HuggingFaceH4/llava-instruct-mix-vsft
I've merged the multi_grained_text_localization and struct_aware_parse datasets, removing problematic images.
However, I kept the images that trigger DecompressionBombWarning. In the multi_grained_text_localization dataset, 777 out of 1,000,000 images triggered this warning. For the struct_aware_parse dataset, 59 out of 3,036,351 images triggered the same warning.
I used… See the full description on the dataset page: https://huggingface.co/datasets/Ryoo72/DocStruct4M.eikon-docsrussian_docs_bypagepixmo-docs-corpusplain_docs_2_100kdai_docssynth_docs_pts_10kheb-noisy-docs-5kmellon_docspl-government-docs-mix-ocr-dataset
Polish municipal administrative documents OCR dataset
An OCR-oriented image dataset built from publicly available Polish municipal administrative documents.
The dataset contains pages collected from materials such as:
resolutions (uchwały)
ordinances / regulations
annexes
official notices
tabular administrative pages
stamped and signed office documents
other municipal administrative materials
All files in this repository are provided as images only.
What is inside… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset.docs_on_several_languages
Dataset Card for "docs_on_several_languages"
This dataset is a collection of different images in different languages.
The daset includes the following languages: Azerbaijani (az: 0), Belorussian (be: 1), Chinese (zh: 16), English (en: 2), Estonian (et: 3), Finnish (fn: 4), Georgian (gr: 5), Japanese (ja: 6), Korean (ko: 7), Kazakh (kk: 8), Latvian (lv: 10), Lithuanian (lt: 9), Mongolian (mn: 11), Norwegian (no: 12), Polish (pl: 13), Russian (ru: 14), Ukranian (uk: 15).
Each language… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyScorpi/docs_on_several_languages.pl-mixed-docs-ocr-dataset-100
Polish Mixed Documents OCR Dataset 100
A small public image-only OCR dataset containing 100 Polish-language document images sampled from a heterogeneous set of document types.
This dataset is designed as a lightweight evaluation sample for:
OCR models
VLM-based document understanding
testing OCR robustness on varied Polish document layouts
Contents
100 JPG images
Polish-language document-like images
mixed document types, including:
official forms
templates… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-mixed-docs-ocr-dataset-100.docs-forai-data
Docs-ForAI Data 🚀
This dataset contains the pre-computed LanceDB vector indices for the Docs-ForAI project.
It provides high-quality RAG (Retrieval-Augmented Generation) context, allowing AI Agents to access verified information from the most popular AI framework documentations without hallucinations.
📂 Included Documentations
This index currently covers:
LangGraph
CrewAI
PydanticAI
AutoGen
LlamaIndex
OpenAI Swarm
Phidata
Haystack
PraisonAI
LiteLLM
🛠… See the full description on the dataset page: https://huggingface.co/datasets/Iruziky/docs-forai-data.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.Thai_Insurance_Docs_OCRlux-typed-docs
Yale LUX dots.ocr layout/OCR outputs
Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest.
Rows: 626,586.
Built from the Yale LUX manifest processing database.
arndee-docsyeji-logic-docs
██╗ ██████╗ ██████╗ ██╗ ██████╗ ██████╗ ██████╗ ██████╗███████╗
██║ ██╔═══██╗██╔════╝ ██║██╔════╝ ██╔══██╗██╔═══██╗██╔════╝██╔════╝
██║ ██║ ██║██║ ███╗██║██║ ██║ ██║██║ ██║██║ ███████╗
██║ ██║ ██║██║ ██║██║██║ ██║ ██║██║ ██║██║ ╚════██║
███████╗╚██████╔╝╚██████╔╝██║╚██████╗ ██████╔╝╚██████╔╝╚██████╗███████║
╚══════╝ ╚═════╝ ╚═════╝ ╚═╝ ╚═════╝ ╚═════╝ ╚═════╝ ╚═════╝╚══════╝
⚡ FORTUNE-TELLING LOGIC ⚡… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-logic-docs.TR-Visual-Docsplain_docs_1legal-docs-images-labelsmlx_support_docsgenerate_pixmo_docs_v2
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
synthetic-math-docs-rigorous-20250624_125218
