datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-deva-ocr-eval
indic_deva_eval
Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting.
Repo: himalaya-ai/indic-deva-ocr-eval
Task: indic_devanagari_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.indic-mozhi-ocrindic-mozhi-ocr
Mozhi (Printed Word Images) - Indic OCR Dataset
This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for
"Towards Deployable OCR Models for Indic Languages". The data is organized by language and split
(train/val/test) and is intended for upload to Hugging Face.
Source
Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php
Paper: Towards Deployable OCR Models for Indic Languages
Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.Indic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.
