Faizaniqbal/IndicOCR
A large-scale, high-fidelity synthetic Document AI & OCR benchmark spanning 23 Pan-Indic languages and 12 distinct writing systems. 1. Overview IndicOCR (IndicPixel) is a large-scale, multilingual Optical Character Recognition (OCR) and Document AI benchmark purpose-built for the South Asian linguistic ecosystem. Spanning all 22 Eighth Schedule Constitutional Languages of India plus Bhojpuri, the dataset covers 12 distinct writing systems including Devanagari… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/IndicOCR.
This repository belongs to Faizaniqbal on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
