datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DocLayNet-v1.2
Dataset Card for DocLayNet v1.2
Dataset Summary
This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column.
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
Human Annotation: DocLayNet is… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.2.DocLayNet-v1.1
Dataset Card for DocLayNet v1.1
Dataset Summary
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
Human Annotation: DocLayNet is hand-annotated by well-trained experts, providing a gold-standard in layout segmentation through human recognition and interpretation of… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/DocLayNet-v1.1.doclaynet_processed
Dataset Card for "doclaynet_processed"
Clean version of DocLayNet ready for finetuning.
doclaynet_benchdoclaynet-pt-enriched-formulaDoclaynet-Full
NOtice: The category Ids are not mapped btw 0-10 doclaynet classes, rather they are 2-12 Use the following Classes map.
{'caption': 2,
'footnote': 3,
'formula': 4,
'list_item': 5,
'page_footer': 6,
'page_header': 7,
'picture': 8,
'section_header': 9,
'table': 10,
'text': 11,
'title': 12}
dataset_info:
config_name: all
features:
name: image, dtype: image
name: category_ids, sequence: int32
name: image_id, dtype: int32
name: boxes, sequence:… See the full description on the dataset page: https://huggingface.co/datasets/thewalnutaisg/Doclaynet-Full.DocLayNet-Instruct-v1-preprocessedsmall-DocLayNet-v1.1DocLayNet-tiny
Dataset Card for "DocLayNet-tiny"
Tiny set for unit tests based on https://huggingface.co/datasets/pierreguillou/DocLayNet-small.
Total ~0.1% of DocLayNet.
DocLayNet-SmallDocLayNet-Instruct-v1arocrbench_doclaynetPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
doclaynet_train_cleaned
doclaynet_train_cleaned
The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
54,199
QA turns
145,370
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
43
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/doclaynet_train_cleaned.doclaynet-smalldoclaynet_mathdoclaynet-full
Category IDs
1 - Caption
2 - Footnote
3 - Formula
4 - List Item
5 - Page Footer
6 - Page Header
7 - Picture
8 - Section Header
9 - Table
10 - Text
11 - Title
DocLayNet-base-lawdoclaynet10classesdoclaynet_train_cleaned
doclaynet_train_cleaned
The doclaynet_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
54,199
QA turns
145,370
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
43
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/doclaynet_train_cleaned.arocrbench_doclaynetv2Doclaynet3doclaynet-yolodoclaynet3classesDoclayNet10doclaynet-yolodoclaynet-bambara-bomudoclaynet-bambara-bomu-200
