CoolFace
Datasetpublic

harsha-desaraju/telugu-line-ocr-bench

Telugu Wikisource OCR — human-verified line crops 1044 single-line crops from Telugu Wikisource page scans, each with a transcription checked against the image by a human. Grayscale, height 64px, width a multiple of 8 — the form the encoder consumes. Columns column meaning image the line crop text gold transcription, human-verified n_graphemes akshara count of text (regex.\X) has_english text contains a Latin-script letter. Digits/punctuation do… See the full description on the dataset page: https://huggingface.co/datasets/harsha-desaraju/telugu-line-ocr-bench.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes32downloads
3 commits on main
2f558212mo ago

Upload README.md with huggingface_hub

harsha-desaraju
e68036e2mo ago

Upload dataset

harsha-desaraju
bc6d2ae2mo ago

initial commit

harsha-desaraju