CoolFace
15 results

indic-ocr

himalaya-ai /indic-deva-ocr-eval indic_deva_eval Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting. Repo: himalaya-ai/indic-deva-ocr-eval Task: indic_devanagari_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.imageimage-to-text1K<n<10K1 likes658 downloads4mo agoHugging FaceFaizaniqbal /IndicOCRgated Large-scale multilingual OCR and document dataset across 23 Pan-Indic languages and 12 writing systems. 1. Overview IndicOCR (IndicPixel) is a large-scale multilingual Optical Character Recognition (OCR) and document dataset covering the South Asian linguistic landscape. The dataset provides dense document coverage across 23 official and literary languages representing 12 distinct writing systems. Dataset Specifications: Scale & Scope: Over 12… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/IndicOCR.image-to-text10M<n<100M0 likes519 downloads0m agoHugging FaceAtharvImmverse /indic-mozhi-ocrimage1M<n<10M0 likes389 downloads6mo agoHugging Facedarknight054 /indic-mozhi-ocr Mozhi (Printed Word Images) - Indic OCR Dataset This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for "Towards Deployable OCR Models for Indic Languages". The data is organized by language and split (train/val/test) and is intended for upload to Hugging Face. Source Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php Paper: Towards Deployable OCR Models for Indic Languages Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.image1M<n<10M2 likes318 downloads8mo agoHugging FaceAbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes112 downloads3mo agoHugging Facesarvamai /indic-ocr-bench Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks, newspapers, magazines, and other published material. The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.imageimage-to-text1K<n<10K8 likes78 downloads5d agoHugging Face