CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes3.3k downloads2y agoHugging Face02cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes971 downloads11mo agoHugging Face03NTT-hil-insight /VDocRetriever-Pretrain-DocStructimagetext-generation100K<n<1M1 likes413 downloads1y agoHugging Face04Ryoo72 /DocStruct4MmPLUG/DocStruct4M reformated for VSFT with TRL's SFT Trainer.Referenced the format of HuggingFaceH4/llava-instruct-mix-vsft I've merged the multi_grained_text_localization and struct_aware_parse datasets, removing problematic images. However, I kept the images that trigger DecompressionBombWarning. In the multi_grained_text_localization dataset, 777 out of 1,000,000 images triggered this warning. For the struct_aware_parse dataset, 59 out of 3,036,351 images triggered the same warning. I used… See the full description on the dataset page: https://huggingface.co/datasets/Ryoo72/DocStruct4M.image1M<n<10M2 likes359 downloads2y agoHugging Face05Tevatron /pixmo-docs-corpusimage100K<n<1M0 likes123 downloads2y agoHugging Face06Akajackson /synth_docs_pts_10kimage10K<n<100K0 likes114 downloads2y agoHugging Face07aallail /plain_docs_2_100kimage100K<n<1M0 likes81 downloads1y agoHugging Face08asafd60 /heb-noisy-docs-5kimage1K<n<10K0 likes74 downloads2y agoHugging Face09nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes73 downloads1mo agoHugging Face10AlekseyScorpi /docs_on_several_languages Dataset Card for "docs_on_several_languages" This dataset is a collection of different images in different languages. The daset includes the following languages: Azerbaijani (az: 0), Belorussian (be: 1), Chinese (zh: 16), English (en: 2), Estonian (et: 3), Finnish (fn: 4), Georgian (gr: 5), Japanese (ja: 6), Korean (ko: 7), Kazakh (kk: 8), Latvian (lv: 10), Lithuanian (lt: 9), Mongolian (mn: 11), Norwegian (no: 12), Polish (pl: 13), Russian (ru: 14), Ukranian (uk: 15). Each language… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyScorpi/docs_on_several_languages.imagetext-classification1K<n<10K2 likes64 downloads1y agoHugging Face11fwgpiyawudk /Thai_Insurance_Docs_OCRgatedimagen<1K0 likes46 downloads9d agoHugging Face12ihsanbasheer /legal-docs-images-labelsimage1K<n<10K0 likes23 downloads1y agoHugging Face13aallail /plain_docs_1image10K<n<100K0 likes19 downloads2y agoHugging Face14nglebm19 /generate_pixmo_docs_v2 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes19 downloads2y agoHugging Face15besartshyti /mlx_support_docsimagen<1K0 likes18 downloads1y agoHugging Face16laxmacl /synthetic-math-docs-rigorous-20250624_130247imagen<1K0 likes15 downloads1y agoHugging Face17IHateStats /Title-Block-Engineering-Docsimagen<1K1 likes15 downloads1y agoHugging Face18laxmacl /synthetic-math-docs-rigorous-20250624_123239imagen<1K0 likes14 downloads1y agoHugging Face19laxmacl /synthetic-math-docs-rigorous-20250624_125218imagen<1K0 likes14 downloads1y agoHugging Face20fwgpiyawudk /Thai_Insurance_Docs_OCR-calib-chandra2gatedimagen<1K0 likes14 downloads1d agoHugging Face21laxmacl /synthetic-math-docs-rigorous-20250624_124739imagen<1K0 likes13 downloads1y agoHugging Face22laxmacl /synthetic-math-docs-rigorous-20250624_125646imagen<1K0 likes13 downloads1y agoHugging Face23nglebm19 /generate_pixmo_docs_v1 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes12 downloads2y agoHugging Face24Nayana-cognitivelab /synthetic-math-docs-rigorous-20250621_140753imagen<1K0 likes11 downloads1y agoHugging Face253sara /docs_sampleimage1K<n<10K0 likes10 downloads1y agoHugging Face26Nayana-cognitivelab /synthetic-math-docs-rigorous-20250623_132049imagen<1K0 likes10 downloads1y agoHugging Face27laxmacl /synthetic-docs-June24th-131907imagen<1K0 likes10 downloads1y agoHugging Face28dhruv107 /docs_pro_max_Jun_23image1K<n<10K0 likes8 downloads2y agoHugging Face29asafd60 /heb-noisy-docs-1kimagen<1K0 likes8 downloads2y agoHugging Face30zimble /russian-docs-vidore-validimage10K<n<100K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.