datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DocLayNet-v1.2
Dataset Card for DocLayNet v1.2
Dataset Summary
This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column.
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.DocLayNet-v1.1
Dataset Card for DocLayNet v1.1
Dataset Summary
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
Human Annotation: DocLayNet is hand-annotated by well-trained experts, providing a gold-standard in layout segmentation through human recognition and interpretation of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.1.doclaynet_processed
Dataset Card for "doclaynet_processed"
Clean version of DocLayNet ready for finetuning.
DocLayNet-baseAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.doclaynet_benchDocLayNet-smallAccurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very effective at layout detection and segmentation. While these datasets are of adequate size to train such models, they severely lack in layout variability since they are sourced from scientific article repositories such as PubMed and arXiv only. Consequently, the accuracy of the layout segmentation drops significantly when these models are applied on more challenging and diverse layouts. In this paper, we present \textit{DocLayNet}, a new, publicly available, document-layout annotation dataset in COCO format. It contains 80863 manually annotated pages from diverse data sources to represent a wide variability in layouts. For each PDF page, the layout annotations provide labelled bounding-boxes with a choice of 11 distinct classes. DocLayNet also provides a subset of double- and triple-annotated pages to determine the inter-annotator agreement. In multiple experiments, we provide smallline accuracy scores (in mAP) for a set of popular object detection models. We also demonstrate that these models fall approximately 10\% behind the inter-annotator agreement. Furthermore, we provide evidence that DocLayNet is of sufficient size. Lastly, we compare models trained on PubLayNet, DocBank and DocLayNet, showing that layout predictions of the DocLayNet-trained models are more robust and thus the preferred choice for general-purpose document-layout analysis.doclaynet-pt-enriched-formulaDocLayNet_rankDoclaynet-Full
NOtice: The category Ids are not mapped btw 0-10 doclaynet classes, rather they are 2-12 Use the following Classes map.
{'caption': 2,
'footnote': 3,
'formula': 4,
'list_item': 5,
'page_footer': 6,
'page_header': 7,
'picture': 8,
'section_header': 9,
'table': 10,
'text': 11,
'title': 12}
dataset_info:
config_name: all
features:
name: image, dtype: image
name: category_ids, sequence: int32
name: image_id, dtype: int32
name: boxes, sequence:… See the full description on the dataset page: https://huggingface.co/datasets/thewalnutaisg/Doclaynet-Full.DocLayNet-Instruct-v1-preprocessedsmall-DocLayNet-v1.1DocLayNet-tiny
Dataset Card for "DocLayNet-tiny"
Tiny set for unit tests based on https://huggingface.co/datasets/pierreguillou/DocLayNet-small.
Total ~0.1% of DocLayNet.
DocLayNet-Instruct-v1DocLayNet-Smalldoclaynet_train_cleaned
doclaynet_train_cleaned
The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
54,199
QA turns
145,370
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
43
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/doclaynet_train_cleaned.doclaynet-smallarocrbench_doclaynetPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
doclaynet_mathDocLayNet_ValidationJust the validation part of the DocLayNet dataset.
The full dataset can be found at https://developer.ibm.com/exchanges/data/all/doclaynet/
DocLayNet-base-lawdoclaynet-full
Category IDs
1 - Caption
2 - Footnote
3 - Formula
4 - List Item
5 - Page Footer
6 - Page Header
7 - Picture
8 - Section Header
9 - Table
10 - Text
11 - Title
doclaynet10classesdoclaynet_train_cleaned
doclaynet_train_cleaned
The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
54,199
QA turns
145,370
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
43
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/doclaynet_train_cleaned.doclaynet-for-yoloUsed by https://github.com/ppaanngggg/yolo-doclaynet
Doclaynet3doclaynet-yoloarocrbench_doclaynetv2doclaynet3classestable-extracted-yolo-data-doclaynet-zipDoclayNet10
