CoolFace
Datasetpublic

harsha-desaraju/telugu-line-ocr-bench

Telugu Wikisource OCR — human-verified line crops 1044 single-line crops from Telugu Wikisource page scans, each with a transcription checked against the image by a human. Grayscale, height 64px, width a multiple of 8 — the form the encoder consumes. Columns column meaning image the line crop text gold transcription, human-verified n_graphemes akshara count of text (regex.\X) has_english text contains a Latin-script letter. Digits/punctuation do… See the full description on the dataset page: https://huggingface.co/datasets/harsha-desaraju/telugu-line-ocr-bench.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes32downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face