datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-ocr-bench
Sarvam Indic OCR Bench
Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material.
The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.Devanagari-OCR-ICL-Benchmark
Devanagari Post-OCR Correction Benchmark
A benchmark for evaluating post-OCR correction systems on Hindi and Marathi
text rendered in Devanagari script, accompanying the paper
"Evaluating In-Context Learning and Retrieval Strategies for Devanagari
Post-OCR Correction" (Bhandari and Harit, 2026).
Summary
20,000 evaluation pairs (10k Hindi, 10k Marathi) of (OCR-corrupted, ground-truth) sentences
~6,600 shot-bank pairs (~3.3k per language) for in-context example… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark.OCRBenchV2-DocParsing-UpdatedGTOCRBenchV2-DocParsing-UpdatedGT is an improved ground truth for the document parsing subset of OCRBench V2, created and verified by Tensorlake.
This version is used in Tensorlake’s OCR and document understanding benchmark comparisons.
The dataset is intended for evaluation and research purposes only.For the original benchmark and other subsets, please refer to OCRBench V2
.
