CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01docling-project /PubTabNet_OTSL Dataset Card for PubTabNet_OTSL Dataset Summary This dataset is a conversion of the original PubTabNet into the OTSL format presented in our paper "Optimized Table Tokenization for Table Structure Recognition". The dataset includes the original annotations amongst new additions. Dataset Structure cells: origunal dataset cell groundtruth (content). otsl: new reduced table structure token format html: original dataset groundtruth HTML (structure). html_restored:… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/PubTabNet_OTSL.imageobject-detection100K<n<1M6 likes4.8k downloads3y agoHugging Face02docling-project /DocLayNet-v1.2 Dataset Card for DocLayNet v1.2 Dataset Summary This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column. DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.image10K<n<100K21 likes4.7k downloads2y agoHugging Face03HuggingFaceM4 /DoclingMatix DoclingMatix DoclingMatix is a large-scale, multimodal dataset designed for training vision-language models in the domain of document intelligence. It was created specifically for training the SmolDocling model, an ultra-compact model for end-to-end document conversion. The dataset is constructed by augmenting Hugging Face's Docmatix. Each sample in Docmatix, which consists of a document image and a few questions and answers about it, has been transformed. The text field is now… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/DoclingMatix.imagevisual-question-answering1M<n<10M56 likes4.2k downloads1y agoHugging Face04docling-project /DocLayNet-v1.1 Dataset Card for DocLayNet v1.1 Dataset Summary DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank: Human Annotation: DocLayNet is hand-annotated by well-trained experts, providing a gold-standard in layout segmentation through human recognition and interpretation of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.1.imageobject-detection10K<n<100K27 likes3.5k downloads3y agoHugging Face05docling-project /PubTables-1M_OTSL Dataset Card for PubTables-1M_OTSL Dataset Summary This dataset enables the evaluation of both object detection models and image-to-text methods. PubTables-1M is introduced in the publication "PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents" by Smock et al. The conversion into HF (Hugging Face) and the addition of the OTSL (Optimized Table Structure Language) format is presented in our paper "Optimized Table Tokenization for Table… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/PubTables-1M_OTSL.imageobject-detection1M<n<10M8 likes2.9k downloads3y agoHugging Face06docling-project /SynthCodeNet SynthCodeNet SynthCodeNet is a multimodal dataset created for training the SmolDocling model. It consists of over 9.3 million synthetically generated image-text pairs, covering code snippets from 56 different programming languages. Text data was sourced from permissively licensed sources, while images were synthetically generated at 120 DPI using LaTeX and Pygments to ensure visual diversity. Dataset Statistics Total samples: 9,334,257 Training set: 8,400… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/SynthCodeNet.imageimage-text-to-text1M<n<10M15 likes2.8k downloads1y agoHugging Face07docling-project /screenparse ScreenParse: Large-Scale Dataset for Complete Screen Parsing News May 2026: ScreenParse v2 is released on main with more robust quality filtering, varied viewport resolutions, leaf-element annotations that reduce annotation noise, and 1,447,100 high-quality training screenshots. The first release is retained on the v1 branch. Dataset Description ScreenParse is a large-scale dataset for complete screen parsing, providing dense annotations of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/screenparse.imageobject-detection1M<n<10M7 likes2.1k downloads4mo agoHugging Face08docling-project /SynthTabNet_OTSL Dataset Card for SynthTabNet_OTSL Dataset Summary This dataset is a conversion of the original SynthTabNet into the OTSL format presented in our paper "Optimized Table Tokenization for Table Structure Recognition". The dataset includes the original annotations amongst new additions. SynthTabNet is organized into 4 parts of 150k tables (600k in total). Each part contains tables with different appearances in regard to their size, structure, style and content. All parts are… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/SynthTabNet_OTSL.imageobject-detection100K<n<1M10 likes1.5k downloads3y agoHugging Face09docling-project /PubTables-1M_OTSL-v1.1 Dataset Card for PubTables-1M_OTSL Dataset Summary This dataset contains tables enriched with information about headers, it is filtered version of original PubTables-1M, with less samples. This dataset enables the evaluation of both object detection models and image-to-text methods. PubTables-1M is introduced in the publication "PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents" by Smock et al. The conversion into HF (Hugging Face) and the… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/PubTables-1M_OTSL-v1.1.imageobject-detection100K<n<1M4 likes1.4k downloads2y agoHugging Face10docling-project /SynthChartNet SynthChartNet SynthChartNet is a multimodal dataset designed for training the SmolDocling model on chart-based document understanding tasks. It consists of 1,981,157 synthetically generated samples, where each image depicts a chart (e.g., line chart, bar chart, pie chart, stacked bar chart), and the associated ground truth is given in OTSL format. Charts were rendered at 120 DPI using a diverse set of visualization libraries: Matplotlib, Seaborn, and Pyecharts, enabling… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/SynthChartNet.imageimage-text-to-text1M<n<10M16 likes1.4k downloads1y agoHugging Face11docling-project /FinTabNet_OTSL Dataset Card for FinTabNet_OTSL Dataset Summary This dataset is a conversion of the original FinTabNet into the OTSL format presented in our paper "Optimized Table Tokenization for Table Structure Recognition". The dataset includes the original annotations amongst new additions. Dataset Structure cells: origunal dataset cell groundtruth (content). otsl: new reduced table structure token format html: original dataset groundtruth HTML (structure). html_restored:… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/FinTabNet_OTSL.imageobject-detection100K<n<1M8 likes1.3k downloads3y agoHugging Face12docling-project /MolGrapher-Synthetic-300K MolGrapher-Synthetic-300K MolGrapher-Synthetic-300K is the synthetic dataset introduced in MolGrapher: Graph-based Visual Recognition of Chemical Structures. Our dataset is created using molecule SMILES retrieved from the database PubChem. Training images are then generated from SMILES using the molecule drawing library RDKit. The synthetic training set is augmented at multiple levels: Molecule level: Molecules are randomly transformed by: (1) displaying explicit hydrogens, (2)… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MolGrapher-Synthetic-300K.image100K<n<1M14 likes962 downloads2y agoHugging Face13docling-project /SynthFormulaNet SynthFormulaNet SynthFormulaNet is a multimodal dataset designed for training the SmolDocling model. It contains over 6.4 million pairs of synthetically rendered images depicting mathematical formulas and their corresponding LaTeX representations. The LaTeX data was collected from permissively licensed sources, and the images were generated using LaTeX at 120 DPI with diverse rendering styles, fonts, and layout configurations to maximize visual variability. This dataset also… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/SynthFormulaNet.imageimage-text-to-text1M<n<10M19 likes724 downloads1y agoHugging Face14docling-project /docling-dpbench Docling-DPBench license: cdla-permissive-2.0 textn<1K1 likes292 downloads2y agoHugging Face15docling-project /HF-CC-v0-0001-images-filteredimage100K<n<1M0 likes279 downloads10mo agoHugging Face16docling-project /doclaynet-pt-enriched-formulaimage100K<n<1M2 likes275 downloads11mo agoHugging Face17docling-project /USPTO-30K USPTO-30K USPTO-30K is the benchmark dataset introduced in MolGrapher: Graph-based Visual Recognition of Chemical Structures. Existing benchmarks for Optical Chemical Structure Recognition have some limitations. Being created using only a few documents, they contain batches of very similar molecules. For example in a patent, a molecule could typically be displayed together with all the substituent of one particular substructure, resulting in large batches of almost identical… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/USPTO-30K.image10K<n<100K11 likes263 downloads3y agoHugging Face18docling-project /MarkushGrapher-2-Datasets MarkushGrapher 2 Datasets Datasets for training and evaluating MarkushGrapher 2, a model for converting patent Markush structure images into CXSMILES representations. Dataset Subsets Subset Train Test Description OCR uspto-mol-m-54k-new 54,785 200 USPTO-MOL-M Markush samples ChemicalOCR predictions uspto-markush — 74 USPTO Markush structures benchmark Ground Truth OCR m2s — 103 Mol2Smiles (M2S) benchmark Ground Truth OCR IP5-markush — 878 IP5 Markush… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MarkushGrapher-2-Datasets.image10K<n<100K3 likes224 downloads6mo agoHugging Face19Xenova /docling-documentstext1M<n<10M1 likes82 downloads1y agoHugging Face20Elliot-Data /DoclingMatix_cleanedgated DoclingMatix_cleaned The DoclingMatix__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 588,763 QA turns 6,394,614 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 503 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/DoclingMatix_cleaned.imagevisual-question-answering100K<n<1M0 likes70 downloads16d agoHugging Face21pranavvmurthy26 /DoclingMatix_5K [!NOTE]This dataset is a fork of HuggingFaceM4/DoclingMatix DoclingMatix DoclingMatix is a large-scale, multimodal dataset designed for training vision-language models in the domain of document intelligence. It was created specifically for training the SmolDocling model, an ultra-compact model for end-to-end document conversion. The dataset is constructed by augmenting Hugging Face's Docmatix. Each sample in Docmatix, which consists of a document image and a few questions and… See the full description on the dataset page: https://huggingface.co/datasets/pranavvmurthy26/DoclingMatix_5K.imagevisual-question-answering1K<n<10K1 likes69 downloads11mo agoHugging Face22pranavvmurthy26 /DoclingMatix_500 [!NOTE]This dataset is a fork of HuggingFaceM4/DoclingMatix DoclingMatix DoclingMatix is a large-scale, multimodal dataset designed for training vision-language models in the domain of document intelligence. It was created specifically for training the SmolDocling model, an ultra-compact model for end-to-end document conversion. The dataset is constructed by augmenting Hugging Face's Docmatix. Each sample in Docmatix, which consists of a document image and a few questions and… See the full description on the dataset page: https://huggingface.co/datasets/pranavvmurthy26/DoclingMatix_500.imagevisual-question-answeringn<1K0 likes39 downloads11mo agoHugging Face23asoria /docling_sample_pdftextn<1K0 likes38 downloads2y agoHugging Face24elliot-mllm /DoclingMatix_cleanedgated DoclingMatix_cleaned The DoclingMatix__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 588,763 QA turns 6,394,614 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 503 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/DoclingMatix_cleaned.imagevisual-question-answering100K<n<1M0 likes22 downloads25d agoHugging Face25davidberenstein1957 /daily-papers-docling Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/daily-papers-docling.image100K<n<1M2 likes17 downloads2y agoHugging Face26asoria /pdf-papers-doclingtextn<1K0 likes17 downloads2y agoHugging Face27baobabtech /test-eval-docs-docling-plain PDF Document Processing with Docling This dataset contains structured markdown extraction from PDFs in baobabtech/test-eval-documents using Docling with hierarchical parsing. Processing Details Source Dataset: baobabtech/test-eval-documents Number of PDFs: 20 Processing Time: 8.4 minutes Processing Date: 2025-12-02 15:40 UTC Configuration PDF Column: pdf_bytes Dataset Split: train Dataset Structure The dataset contains all original columns plus:… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/test-eval-docs-docling-plain.documentn<1K0 likes17 downloads10mo agoHugging Face28asoria /pdf-dataset-doclingtextn<1K0 likes15 downloads2y agoHugging Face29hynky /docling-issues-first-50textn<1K0 likes12 downloads5mo agoHugging Face30asoria /pdf_bench-doclingtextn<1K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.