CoolFace
Datasetpublic

meharuhanzz/OCR-Bench1000-Hindi

OCR-Bench1000-Hindi 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Hindi OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category hindi_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character count… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Hindi.

sourceHugging Facemitupdated 12d agoView on Hugging Face
0likes247downloads
5 commits on main
050b6a712d ago

Remove generation-methodology details from README

meharuhanzz
3bb251f12d ago

Update data-quality note wording

meharuhanzz
d5e8efc12d ago

Reorder metadata.jsonl fields to match viewer column order

meharuhanzz
c43be2712d ago

Add tofu-fixed OCR-Bench1000 hindi samples

meharuhanzz
b7f213512d ago

initial commit

meharuhanzz