datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-handwriting-ocr
Persian Handwriting OCR Dataset
Dataset Summary
A standardized dataset of Persian (Farsi) handwritten pages with word-level
bounding-box annotations and transcriptions. The dataset is page-level:
each sample is a full page scan; annotations are one row per word bbox on
that page. This is the most flexible form -- users can train page-level OCR,
word detection (DBNet/PaddleOCR), or derive word/line crops as needed.
Pages: 1115 scanned pages (canonical IDs… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-handwriting-ocr.OCR_Uchan
Configuration: default
Split: train
Total Rows: 1,800,000
printMethod
Type: categorical
Data Type: object
Unique Values: 1
Value Distribution:
Value
Count
Percentage
PrintMethod_Modern
1,800,000
100.00%
script
Type: categorical
Data Type: object
Unique Values: 3
Value Distribution:
Value
Count
Percentage
ScriptTibt
1,476,272
82.02%
ScriptDbuCan
314,945
17.50%
ScriptHani
8,783
0.49%… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR_Uchan.assistive-ocr-data-acquisition
Assistive OCR Benchmark Data
Multilingual OCR benchmark for visually impaired assistance — Indian medicine labels, packaged goods, and signage in Bengali, Hindi, and English.
Dataset Summary
Property
Value
Total images
7,004 rows in manifest
Image sources
images/hf_medicines/, images/openfoodfacts/, images/synthetic/
Domains
medicine_packaging (6,588), packaged_goods (386), signage (30)
Languages
bn+en (5,337), hi+en (868), en (799)
Splits
dev… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/assistive-ocr-data-acquisition.OCR_Uchan
Configuration: default
Split: train
Total Rows: 1,800,000
printMethod
Type: categorical
Data Type: object
Unique Values: 1
Value Distribution:
Value
Count
Percentage
PrintMethod_Modern
1,800,000
100.00%
script
Type: categorical
Data Type: object
Unique Values: 3
Value Distribution:
Value
Count
Percentage
ScriptTibt
1,476,272
82.02%
ScriptDbuCan
314,945
17.50%
ScriptHani
8,783
0.49%… See the full description on the dataset page: https://huggingface.co/datasets/zmlm2001/OCR_Uchan.Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.ocr-leaderboard
OmniAI OCR Leaderboard
A comprehensive leaderboard comparing OCR and data extraction performance across traditional OCR providers and multimodal LLMs, such as gpt-4o and gemini-2.0. The dataset includes full results from testing 9 providers on 1,000 pages each.
Benchmark Results (Feb 2025) | Source Code
Persian-OCR-230k.
├── Images/
├── train.csv (184k)
└── test.csv (46k)
vibook-OCR_VQA
