SBB/page_extraction_dataset
Title Page Extraction Dataset Description In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created.… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.
022
page_extraction.tar.gzdownload
