CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MR3z4 /persian-handwriting-ocr Persian Handwriting OCR Dataset Dataset Summary A standardized dataset of Persian (Farsi) handwritten pages with word-level bounding-box annotations and transcriptions. The dataset is page-level: each sample is a full page scan; annotations are one row per word bbox on that page. This is the most flexible form -- users can train page-level OCR, word detection (DBNet/PaddleOCR), or derive word/line crops as needed. Pages: 1115 scanned pages (canonical IDs… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-handwriting-ocr.imageimage-to-text100K<n<1M0 likes220 downloads1mo agoHugging Face02openpecha /OCR_Uchan Configuration: default Split: train Total Rows: 1,800,000 printMethod Type: categorical Data Type: object Unique Values: 1 Value Distribution: Value Count Percentage PrintMethod_Modern 1,800,000 100.00% script Type: categorical Data Type: object Unique Values: 3 Value Distribution: Value Count Percentage ScriptTibt 1,476,272 82.02% ScriptDbuCan 314,945 17.50% ScriptHani 8,783 0.49%… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR_Uchan.image1M<n<10M0 likes191 downloads11mo agoHugging Face03Arko007 /assistive-ocr-data-acquisition Assistive OCR Benchmark Data Multilingual OCR benchmark for visually impaired assistance — Indian medicine labels, packaged goods, and signage in Bengali, Hindi, and English. Dataset Summary Property Value Total images 7,004 rows in manifest Image sources images/hf_medicines/, images/openfoodfacts/, images/synthetic/ Domains medicine_packaging (6,588), packaged_goods (386), signage (30) Languages bn+en (5,337), hi+en (868), en (799) Splits dev… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/assistive-ocr-data-acquisition.imageimage-to-text1K<n<10K0 likes156 downloads2mo agoHugging Face04zmlm2001 /OCR_Uchan Configuration: default Split: train Total Rows: 1,800,000 printMethod Type: categorical Data Type: object Unique Values: 1 Value Distribution: Value Count Percentage PrintMethod_Modern 1,800,000 100.00% script Type: categorical Data Type: object Unique Values: 3 Value Distribution: Value Count Percentage ScriptTibt 1,476,272 82.02% ScriptDbuCan 314,945 17.50% ScriptHani 8,783 0.49%… See the full description on the dataset page: https://huggingface.co/datasets/zmlm2001/OCR_Uchan.image1M<n<10M0 likes149 downloads9mo agoHugging Face05FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes125 downloads1y agoHugging Face06getomni-ai /ocr-leaderboard OmniAI OCR Leaderboard A comprehensive leaderboard comparing OCR and data extraction performance across traditional OCR providers and multimodal LLMs, such as gpt-4o and gemini-2.0. The dataset includes full results from testing 9 providers on 1,000 pages each. Benchmark Results (Feb 2025) | Source Code image1K<n<10K8 likes101 downloads2y agoHugging Face07ordaktaktak /Persian-OCR-230k. ├── Images/ ├── train.csv (184k) └── test.csv (46k) imageimage-to-text100K<n<1M2 likes100 downloads2y agoHugging Face08sunbv56 /vibook-OCR_VQAimage1K<n<10K0 likes5 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.