datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doclaynet-document-level
DocLayNet Document-Level Reconstruction and 8K Expansion
This dataset is a normalized, one-row-per-document view over the page-level
DocLayNet v1.1
dataset. Pages are grouped using DocLayNet's source metadata and ordered by their original
page number.
Dataset summary
2,944 logical documents
80,863 observed pages
896 complete document groups
2,048 partial document groups
Train: 2,355 documents / 60,810 pages
Validation: 294 documents / 7,964 pages
Test: 295… See the full description on the dataset page: https://huggingface.co/datasets/operant-ai/doclaynet-document-level.DoCLayNet-large-wt-image
Dataset Card for DocLayNet large without image
About this card (02/14/2024)
Property and license
All information from this page but the content of this paragraph "About this card (02/14/2025)" has been copied/pasted from Dataset Card for DocLayNet.
DocLayNet is a dataset created by Deep Search (IBM Research) published under license CDLA-Permissive-1.0.
I do not claim any rights to the data taken from this dataset and published on this page.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agomberto/DoCLayNet-large-wt-image.
