CoolFace
Datasetpublic

meharuhanzz/OCR-Bench1000-Kannada

OCR-Bench1000-Kannada 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Kannada OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category kannada_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Kannada.

sourceHugging Facemitupdated 10d agoView on Hugging Face
0likes163downloads
Dataset Card

OCR-Bench1000-Kannada

1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Kannada OCR training corpus. This is a benchmark/sample release, not the full training set.

Data fields

FieldDescription
file_namerelative path to the image (images/...)
textground-truth transcription
categorykannada_only / english_only / mixed / numeric_and_symbols
length_bucketshort / medium / long, by character count

Category distribution (this sample)

CategoryCount
kannada_only972
english_only18
numeric_and_symbols10

Length-bucket distribution (this sample)

BucketCount
medium445
long293
short262

Note: this language currently has no mixed (native+Latin in the same line) samples.

A note on data quality and how to help

This corpus was reviewed by someone who cannot personally read or verify Kannada.

If you notice anything wrong — a mistranscription, an odd or non-natural sentence, wrong script/font rendering, or anything else that looks off — please open a thread in this repo's Community tab (Discussions) and mention the file_name and what's wrong. Real corrections from people who can actually read Kannada are exactly what's needed to improve this further, and are genuinely welcome.

License

MIT — see LICENSE.