CoolFace
Datasetpublic

SBB/page_extraction_dataset

Title Page Extraction Dataset Description In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created.… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.

sourceHugging Facecc-by-4.0updated 8mo agoView on Hugging Face
0likes22downloads
3 commits on main
cde2d458mo ago

Upload data

Jrglmn
85963dd8mo ago

create datasheet

Jrglmn
1fe54038mo ago

initial commit

Jrglmn