datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telugu-synthetic-line-imagesTelugu-book-text-imagesTelugu-text-imagetelugu-wikisource-text-imagesOCR-Bench1000-Telugu
OCR-Bench1000-Telugu
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Telugu OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
telugu_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Telugu.telugu-pdf-line-image-texttelugu-line-ocr-bench
Telugu Wikisource OCR — human-verified line crops
1044 single-line crops from Telugu Wikisource page scans, each with a
transcription checked against the image by a human. Grayscale, height 64px,
width a multiple of 8 — the form the encoder consumes.
Columns
column
meaning
image
the line crop
text
gold transcription, human-verified
n_graphemes
akshara count of text (regex.\X)
has_english
text contains a Latin-script letter. Digits/punctuation do… See the full description on the dataset page: https://huggingface.co/datasets/harsha-desaraju/telugu-line-ocr-bench.telugu-ocr-datasetchemistry-multimodal-exams-telugutelugu-line-text-imageMedQA_ODD_Telugu_testtelugu-book-line-images-sampletelugu-book-line-imagesMedQA_tran_COT_Telugu_train2MedQA_tran_COT_Telugu_test2telugu-ocr
Telugu OCR Dataset
A corpus of aligned scanned page images and human-transcribed Telugu text, sourced from Telugu Wikisource. Built for OCR model training and evaluation.
Stats
Total page pairs
~25,565
Books
221
Total size
~11 GB
License
CC BY-SA 4.0
Dataset Structure
dataset/
<book_title>/
page_0001.jpg ← scan image
page_0001.txt ← transcribed Telugu text (UTF-8)
page_0004.jpg
page_0004.txt
...… See the full description on the dataset page: https://huggingface.co/datasets/AlbertoChestnut/telugu-ocr.
