CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01docling-project /screenparse ScreenParse: Large-Scale Dataset for Complete Screen Parsing News May 2026: ScreenParse v2 is released on main with more robust quality filtering, varied viewport resolutions, leaf-element annotations that reduce annotation noise, and 1,447,100 high-quality training screenshots. The first release is retained on the v1 branch. Dataset Description ScreenParse is a large-scale dataset for complete screen parsing, providing dense annotations of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/screenparse.imageobject-detection1M<n<10M7 likes2.1k downloads4mo agoHugging Face02biglam /icdar2021-historical-document-dating ICDAR 2021 Historical Document Classification — Task 2 (Dating) 13,810 manuscript page images labelled with the period in which they were produced. Images come from e-codices, the virtual manuscript library of Switzerland. Split Images Date range Median span Dated to a single year train 11,294 800–1899 45 years 1,409 test 2,516 800–1921 49 years 264 The label is an interval, not a year Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.imageimage-classification10K<n<100K2 likes411 downloads2mo agoHugging Face03Project-AgML /plant_doc_classification Plant Doc Classification A dataset for disease classification of various plants. The dataset contains 2,569 images across 28 classes:Images per class: Apple Scab Leaf: 93 Apple leaf: 91 Apple rust leaf: 88 Bell_pepper leaf: 61 Bell_pepper leaf spot: 71 Blueberry leaf: 114 Cherry leaf: 57 Corn Gray leaf spot: 68 Corn leaf blight: 191 Corn rust leaf: 116 Peach leaf: 111 Potato leaf early blight: 116 Potato leaf late blight: 105 Raspberry leaf: 119 Soyabean leaf: 65 Squash Powdery… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/plant_doc_classification.imageimage-classification1K<n<10K1 likes244 downloads3mo agoHugging Face04hf-tuner /rvl-cdip-document-classification rvl-cdip-document-classification This dataset is created from original aharley/rvl_cdip dataset using this notebook Dataset Summary This dataset consists of 8992 grayscale images in 16 classes, with 562 images per class. There are 8000 training images(500 image per class) and 992 test images(62 images per class). The images are sized so their largest dimension does not exceed 1000 pixels. imageimage-classification1K<n<10K0 likes121 downloads1y agoHugging Face05prithivMLmods /Document-Type-Detection Document-Type-Detection Dataset Summary The Document-Type-Detection dataset is a large-scale image classification dataset consisting of scanned or photographed document images. Each image is categorized into one of nine document types. This dataset is ideal for training document classification models in finance, administration, OCR, and automation workflows. Supported Tasks Multiclass Document Classification Classify an input document image into one of the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Document-Type-Detection.imageimage-classification10K<n<100K1 likes80 downloads1y agoHugging Face06nutrientdocs /doc-split-benchmark Doc-Split Benchmark The evaluation slice for page-stream segmentation — the exact set behind the leaderboard and the cloud-VLM comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible. This is the benchmark, not the training corpus (which stays private). 🏆 Leaderboard: doc-split-leaderboard 🎯 Demo: doc-split-demo 🟢 Model: doc-split-mini-e5 (open weights) 🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.imageimage-classificationn<1K0 likes80 downloads2mo agoHugging Face07thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes70 downloads4mo agoHugging Face08sitloboi2012 /CMDS_Multimodal_Document Dataset Card for Cyrillic Multimodel Document (CMDS) This is the dataset consists of 3789 pairs of images and text across 31 categories downloaded from the Bulgarian ministry of finance Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards Uses this dataset for downstream task like Document Classification, Image Classification or Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/sitloboi2012/CMDS_Multimodal_Document.imageimage-classification1K<n<10K0 likes66 downloads3y agoHugging Face09nutrientdocs /document-classification-benchmark Document Classification Benchmark (open-vocab, zero-shot) Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot, open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked into a head. Test split only; not for training. Every image is drawn from a permissively-licensed, redistributable source. Powers the document-classification-leaderboard and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.imagezero-shot-image-classification1K<n<10K0 likes63 downloads2mo agoHugging Face10giacolees /siglip-doc-understanding-classifier SigLIP Doc Understanding — Unanswerable Question Detection Dataset A mixed answerable / unanswerable benchmark dataset built from DocVQA and MP-DocVQA, used to train and evaluate the siglip-doc-understanding-classifier unanswerable-question detector. Each row pairs a document image with a question. Half of the questions are the original, answerable DocVQA/MP-DocVQA questions; the other half are corrupted versions of those same questions — modified so the document image no longer… See the full description on the dataset page: https://huggingface.co/datasets/giacolees/siglip-doc-understanding-classifier.textvisual-question-answering1K<n<10K0 likes54 downloads3mo agoHugging Face11biglam /muninn-ww1-documents Muninn WWI Documents (CEF Attestation Papers & War Diaries) A tabular conversion of the document records in the Muninn Project's World War I Linked Open Data archive (https://rdf.muninn-project.org/). The Muninn Project is a research project that extracts structured data from digitized WWI-era archival documents. The dataset covers 28,742 documents — 28,718 Canadian Expeditionary Force (CEF) attestation papers (enlistment forms of individual soldiers) and 24 Canadian unit war… See the full description on the dataset page: https://huggingface.co/datasets/biglam/muninn-ww1-documents.tabularimage-classification100K<n<1M0 likes44 downloads2mo agoHugging Face12nutrientdocs /doc-openvocab-benchmark Open-Vocab Document & Figure Classification Benchmark Given a document or figure image and an arbitrary set of text labels, which one is right? This is a zero-shot, open-vocabulary image-classification benchmark for the document-AI setting: every image is scored against a broad ~48-label candidate vocabulary (document types + figure/zone types), and the task is to pick the correct label. The labels are supplied at inference — which is precisely what a fixed-label supervised… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-openvocab-benchmark.imagezero-shot-image-classification1K<n<10K1 likes37 downloads2mo agoHugging Face13langminer /doc-content-clustering-740 Document Content-Clustering Benchmark (12 classes, 740 items) A small, curated benchmark for clustering documents by their content topic (not by their visual form/layout). Each item is a single document page provided as an image plus two text views (a VLM description and OCR markdown), with a ground-truth content class. The set is intentionally "tangle-stripped": 27 borderline items whose content sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.imageimage-classificationn<1K0 likes25 downloads2mo agoHugging Face14tzj04 /docdet-scamai-cropsgated tzj04/docdet-scamai-crops Training crops derived from the Scam-AI document-forgery datasets, for the DocDet authentic-vs-AI-generated detector. This is a derivative work. It is not an official Scam-AI release. What a row is Each forgery in the source data patches a single field into an otherwise authentic scan - roughly 0.3% of the page. At a 224px whole-page input that edit survives as a handful of pixels, and a random-resized crop can miss it altogether. So… See the full description on the dataset page: https://huggingface.co/datasets/tzj04/docdet-scamai-crops.tabularimage-classification10K<n<100K1 likes21 downloads26d agoHugging Face15Mouwiya /image-in-Words400_DOCCI_Testimageimage-to-textn<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.