CoolFace
15 results

indic_ocr

Faizaniqbal /IndicOCRgated Large-scale multilingual OCR and document dataset across 23 Pan-Indic languages and 12 writing systems. 1. Overview IndicOCR (IndicPixel) is a large-scale multilingual Optical Character Recognition (OCR) and document dataset covering the South Asian linguistic landscape. The dataset provides dense document coverage across 23 official and literary languages representing 12 distinct writing systems. Dataset Specifications: Scale & Scope: Over 12… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/IndicOCR.imageimage-to-text10M<n<100M0 likes1.2k downloads4h agoHugging Facehimalaya-ai /indic-deva-ocr-eval indic_deva_eval Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting. Repo: himalaya-ai/indic-deva-ocr-eval Task: indic_devanagari_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.imageimage-to-text1K<n<10K1 likes660 downloads4mo agoHugging FaceAtharvImmverse /indic-mozhi-ocrimage1M<n<10M0 likes413 downloads6mo agoHugging Facedarknight054 /indic-mozhi-ocr Mozhi (Printed Word Images) - Indic OCR Dataset This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for "Towards Deployable OCR Models for Indic Languages". The data is organized by language and split (train/val/test) and is intended for upload to Hugging Face. Source Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php Paper: Towards Deployable OCR Models for Indic Languages Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.image1M<n<10M2 likes299 downloads8mo agoHugging FaceAbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes123 downloads3mo agoHugging Facesarvamai /indic-ocr-bench Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material. The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.imageimage-to-text1K<n<10K9 likes109 downloads5d agoHugging Face