datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-handwriting-ocr
Persian Handwriting OCR Dataset
Dataset Summary
A standardized dataset of Persian (Farsi) handwritten pages with word-level
bounding-box annotations and transcriptions. The dataset is page-level:
each sample is a full page scan; annotations are one row per word bbox on
that page. This is the most flexible form -- users can train page-level OCR,
word detection (DBNet/PaddleOCR), or derive word/line crops as needed.
Pages: 1115 scanned pages (canonical IDs… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-handwriting-ocr.ocr-leaderboard
OmniAI OCR Leaderboard
A comprehensive leaderboard comparing OCR and data extraction performance across traditional OCR providers and multimodal LLMs, such as gpt-4o and gemini-2.0. The dataset includes full results from testing 9 providers on 1,000 pages each.
Benchmark Results (Feb 2025) | Source Code
ocr
Chagatai OCR
Chagatai text in Arabic script, with Cyrillic transcriptions and Kazakh translations.
The dataset is available as Dataset_OCR.xlsx and UTF-8 Dataset_OCR.csv. Both contain 875 non-empty data rows.
Column
Contents
Original
Text in Arabic script
Transcript
Cyrillic transcription
Translation (?)
Kazakh translation; the original column label is retained
Pages
Page markers, provided on selected rows
The Excel workbook is provided unchanged. The CSV… See the full description on the dataset page: https://huggingface.co/datasets/chagatai-project/ocr.ocr_textsexpert-ocr_resultsexpert-ocr-with-text_results
