CoolFace
Datasetpublic

nevmenandr/russian-old-orthography-ocr

Basic Description Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography. The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917). Data structure For each text there is a file with its image and the text corresponding to this image. The names of these files are… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.

sourceHugging Facemitupdated 2y agoView on Hugging Face
7likes769downloads

nevmenandr/russian-old-orthography-ocr · main · files are served by the source, never re-hosted here