indic_ocr
IndicOCR
Large-scale multilingual OCR and document dataset across 23 Pan-Indic languages and 12 writing systems.
1. Overview
IndicOCR (IndicPixel) is a large-scale multilingual Optical Character Recognition (OCR) and document dataset covering the South Asian linguistic landscape. The dataset provides dense document coverage across 23 official and literary languages representing 12 distinct writing systems.
Dataset Specifications:
Scale & Scope: Over 12… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/IndicOCR.indic-deva-ocr-eval
indic_deva_eval
Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting.
Repo: himalaya-ai/indic-deva-ocr-eval
Task: indic_devanagari_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.indic-mozhi-ocrindic-mozhi-ocr
Mozhi (Printed Word Images) - Indic OCR Dataset
This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for
"Towards Deployable OCR Models for Indic Languages". The data is organized by language and split
(train/val/test) and is intended for upload to Hugging Face.
Source
Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php
Paper: Towards Deployable OCR Models for Indic Languages
Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.Indic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.
