meharuhanzz/OCR-Bench1000-Tamil
OCR-Bench1000-Tamil 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Tamil OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category tamil_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character count… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Tamil.
This repository belongs to meharuhanzz on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
