CoolFace
Datasetpublic

tetrak/armenian-ocr-crops

Tetrak Armenian OCR crops Training data for tetrak_hy, the Armenian text recogniser we are building as an EasyOCR custom model in tetrak-hy-trainer for Tetrak, an OCR pipeline for community archives. The dataset has three configurations: corpus — 1,190 proofread pages of the Armenian Soviet Encyclopedia, as plain text with full Wikisource provenance. crops — the v0 synthetic pre-training set: 181,800 rendered word crops with transcriptions. crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.

sourceHugging Facecc-by-sa-4.0updated 27d agoView on Hugging Face
1likes124downloads

tetrak/armenian-ocr-crops · main · files are served by the source, never re-hosted here