AlbertoChestnut/telugu-ocr
Telugu OCR Dataset A corpus of aligned scanned page images and human-transcribed Telugu text, sourced from Telugu Wikisource. Built for OCR model training and evaluation. Stats Total page pairs ~25,565 Books 221 Total size ~11 GB License CC BY-SA 4.0 Dataset Structure dataset/ <book_title>/ page_0001.jpg ← scan image page_0001.txt ← transcribed Telugu text (UTF-8) page_0004.jpg page_0004.txt ...… See the full description on the dataset page: https://huggingface.co/datasets/AlbertoChestnut/telugu-ocr.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face