datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pixmo-docs
PixMo-Docs
We now recommend using CoSyn-400k and CoSyn-point over these
datasets. They are improved versions with more images categories and an improved generation pipeline.
PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
VDocRetriever-Pretrain-DocStructDocStruct4MmPLUG/DocStruct4M reformated for VSFT with TRL's SFT Trainer.Referenced the format of HuggingFaceH4/llava-instruct-mix-vsft
I've merged the multi_grained_text_localization and struct_aware_parse datasets, removing problematic images.
However, I kept the images that trigger DecompressionBombWarning. In the multi_grained_text_localization dataset, 777 out of 1,000,000 images triggered this warning. For the struct_aware_parse dataset, 59 out of 3,036,351 images triggered the same warning.
I used… See the full description on the dataset page: https://huggingface.co/datasets/Ryoo72/DocStruct4M.pixmo-docs-corpussynth_docs_pts_10kplain_docs_2_100kheb-noisy-docs-5kdoc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.docs_on_several_languages
Dataset Card for "docs_on_several_languages"
This dataset is a collection of different images in different languages.
The daset includes the following languages: Azerbaijani (az: 0), Belorussian (be: 1), Chinese (zh: 16), English (en: 2), Estonian (et: 3), Finnish (fn: 4), Georgian (gr: 5), Japanese (ja: 6), Korean (ko: 7), Kazakh (kk: 8), Latvian (lv: 10), Lithuanian (lt: 9), Mongolian (mn: 11), Norwegian (no: 12), Polish (pl: 13), Russian (ru: 14), Ukranian (uk: 15).
Each language… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyScorpi/docs_on_several_languages.Thai_Insurance_Docs_OCRlegal-docs-images-labelsplain_docs_1generate_pixmo_docs_v2
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
mlx_support_docssynthetic-math-docs-rigorous-20250624_130247Title-Block-Engineering-Docssynthetic-math-docs-rigorous-20250624_123239synthetic-math-docs-rigorous-20250624_125218Thai_Insurance_Docs_OCR-calib-chandra2synthetic-math-docs-rigorous-20250624_124739synthetic-math-docs-rigorous-20250624_125646generate_pixmo_docs_v1
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
synthetic-math-docs-rigorous-20250621_140753docs_samplesynthetic-math-docs-rigorous-20250623_132049synthetic-docs-June24th-131907docs_pro_max_Jun_23heb-noisy-docs-1krussian-docs-vidore-valid
