CoolFace
Datasetpublic

biglam/bnl_ground_truth_newspapers_before_1878

Dataset description 33.000 transcribed text lines from historical newspapers (before 1878) along with the cropped images of the original scans Text line based OCR 19.000 text lines in Antiqua 14.000 text lines in Fraktur Transcribed using double-keying (99.95% accuracy) Public Domain, CC0 (See copyright notice) Best for training an OCR engine The newspapers used are: Le Gratis luxembourgeois (1857-1858) Luxemburger Volks-Freund (1869-1876) L'Arlequin (1848-1848) Courrier du… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bnl_ground_truth_newspapers_before_1878.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
2likes72downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
biglam/bnl_ground_truth_newspapers_before_1878 · CoolFace