CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01diffusers /docs-imagesimagen<1K0 likes10k downloads6mo agoHugging Face02diffusers /diffusers-images-docsimagen<1K0 likes7.8k downloads2y agoHugging Face03allenai /pixmo-docs PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images categories and an improved generation pipeline. PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents. The data was created by using the Claude large language model to generate code that can be executed to render an image, and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.imagevisual-question-answering100K<n<1M35 likes3.9k downloads2y agoHugging Face04cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes860 downloads11mo agoHugging Face05mPLUG /DocStruct4Mimagen<1K13 likes697 downloads2y agoHugging Face06NTT-hil-insight /VDocRetriever-Pretrain-DocStructimagetext-generation100K<n<1M1 likes418 downloads1y agoHugging Face07Ryoo72 /DocStruct4MmPLUG/DocStruct4M reformated for VSFT with TRL's SFT Trainer.Referenced the format of HuggingFaceH4/llava-instruct-mix-vsft I've merged the multi_grained_text_localization and struct_aware_parse datasets, removing problematic images. However, I kept the images that trigger DecompressionBombWarning. In the multi_grained_text_localization dataset, 777 out of 1,000,000 images triggered this warning. For the struct_aware_parse dataset, 59 out of 3,036,351 images triggered the same warning. I used… See the full description on the dataset page: https://huggingface.co/datasets/Ryoo72/DocStruct4M.image1M<n<10M2 likes360 downloads2y agoHugging Face08DEVAIEXP /eikon-docsimagen<1K0 likes245 downloads2y agoHugging Face09zimble /russian_docs_bypageimagevisual-document-retrieval1K<n<10K0 likes125 downloads4mo agoHugging Face10Tevatron /pixmo-docs-corpusimage100K<n<1M0 likes124 downloads2y agoHugging Face11aallail /plain_docs_2_100kimage100K<n<1M0 likes91 downloads1y agoHugging Face12h2oai /dai_docsimagen<1K0 likes89 downloads3y agoHugging Face13Akajackson /synth_docs_pts_10kimage10K<n<100K0 likes83 downloads2y agoHugging Face14asafd60 /heb-noisy-docs-5kimage1K<n<10K0 likes74 downloads2y agoHugging Face15OzzyGT /mellon_docsimagen<1K0 likes74 downloads7mo agoHugging Face16Lukaszl /pl-government-docs-mix-ocr-dataset Polish municipal administrative documents OCR dataset An OCR-oriented image dataset built from publicly available Polish municipal administrative documents. The dataset contains pages collected from materials such as: resolutions (uchwały) ordinances / regulations annexes official notices tabular administrative pages stamped and signed office documents other municipal administrative materials All files in this repository are provided as images only. What is inside… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset.imageimage-to-textn<1K1 likes69 downloads6mo agoHugging Face17AlekseyScorpi /docs_on_several_languages Dataset Card for "docs_on_several_languages" This dataset is a collection of different images in different languages. The daset includes the following languages: Azerbaijani (az: 0), Belorussian (be: 1), Chinese (zh: 16), English (en: 2), Estonian (et: 3), Finnish (fn: 4), Georgian (gr: 5), Japanese (ja: 6), Korean (ko: 7), Kazakh (kk: 8), Latvian (lv: 10), Lithuanian (lt: 9), Mongolian (mn: 11), Norwegian (no: 12), Polish (pl: 13), Russian (ru: 14), Ukranian (uk: 15). Each language… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyScorpi/docs_on_several_languages.imagetext-classification1K<n<10K2 likes64 downloads1y agoHugging Face18Lukaszl /pl-mixed-docs-ocr-dataset-100 Polish Mixed Documents OCR Dataset 100 A small public image-only OCR dataset containing 100 Polish-language document images sampled from a heterogeneous set of document types. This dataset is designed as a lightweight evaluation sample for: OCR models VLM-based document understanding testing OCR robustness on varied Polish document layouts Contents 100 JPG images Polish-language document-like images mixed document types, including: official forms templates… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-mixed-docs-ocr-dataset-100.imageimage-to-textn<1K0 likes63 downloads6mo agoHugging Face19Iruziky /docs-forai-data Docs-ForAI Data 🚀 This dataset contains the pre-computed LanceDB vector indices for the Docs-ForAI project. It provides high-quality RAG (Retrieval-Augmented Generation) context, allowing AI Agents to access verified information from the most popular AI framework documentations without hallucinations. 📂 Included Documentations This index currently covers: LangGraph CrewAI PydanticAI AutoGen LlamaIndex OpenAI Swarm Phidata Haystack PraisonAI LiteLLM 🛠… See the full description on the dataset page: https://huggingface.co/datasets/Iruziky/docs-forai-data.imagen<1K0 likes52 downloads9mo agoHugging Face20nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes52 downloads1mo agoHugging Face21fwgpiyawudk /Thai_Insurance_Docs_OCRgatedimagen<1K0 likes42 downloads6d agoHugging Face22yale-cultural-heritage /lux-typed-docs Yale LUX dots.ocr layout/OCR outputs Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest. Rows: 626,586. Built from the Yale LUX manifest processing database. image100K<n<1M0 likes38 downloads3mo agoHugging Face23fwgpiyawudk /arndee-docsgatedimagen<1K0 likes38 downloads28d agoHugging Face24tellang /yeji-logic-docs ██╗ ██████╗ ██████╗ ██╗ ██████╗ ██████╗ ██████╗ ██████╗███████╗ ██║ ██╔═══██╗██╔════╝ ██║██╔════╝ ██╔══██╗██╔═══██╗██╔════╝██╔════╝ ██║ ██║ ██║██║ ███╗██║██║ ██║ ██║██║ ██║██║ ███████╗ ██║ ██║ ██║██║ ██║██║██║ ██║ ██║██║ ██║██║ ╚════██║ ███████╗╚██████╔╝╚██████╔╝██║╚██████╗ ██████╔╝╚██████╔╝╚██████╗███████║ ╚══════╝ ╚═════╝ ╚═════╝ ╚═╝ ╚═════╝ ╚═════╝ ╚═════╝ ╚═════╝╚══════╝ ⚡ FORTUNE-TELLING LOGIC ⚡… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-logic-docs.imagetext-generationn<1K0 likes27 downloads8mo agoHugging Face25ucsahin /TR-Visual-Docsimagen<1K0 likes23 downloads2y agoHugging Face26aallail /plain_docs_1image10K<n<100K0 likes18 downloads2y agoHugging Face27ihsanbasheer /legal-docs-images-labelsimage1K<n<10K0 likes18 downloads1y agoHugging Face28besartshyti /mlx_support_docsimagen<1K0 likes18 downloads1y agoHugging Face29nglebm19 /generate_pixmo_docs_v2 Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. imagen<1K0 likes16 downloads1y agoHugging Face30laxmacl /synthetic-math-docs-rigorous-20250624_125218imagen<1K0 likes14 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.