datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-ocr-community-dataset-argilla
Persian OCR community dataset - Argilla view
Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects.
persian-license-plate-v1
Dataset is downloaded from here which was provided at Amirkabir University of Technology.
The dataset is labeled by the authors.
Experimental results show that the fine-tuned model works well in Persian License Plate.
Usage
You can download the dataset easily using HF datasets package in Python:
!pip install datasets
from datasets import load_dataset
dataset = load_dataset("hezarai/persian-license-plate-v1", split="train") # Other splits: validation, test
print(dataset[0])
persian-ocr-community-datasetpersian-handwriting-pages-3.69m
Persian Handwriting Pages 3.69M
3,690,000 deterministic, densely composed Persian handwriting pages.
This expansion uses new random seeds and is complementary to
Reza2kn/persian-handwriting-pages-369k,
not a repetition of its rendered pages.
The public viewer intentionally exposes exactly two columns: image and label.
Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification.
Source handwriting
Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.persian-printed-ocr-3.5m
Persian Printed OCR 3.5M
A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public
datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733
rejected rows are excluded. The viewer exposes exactly image and label.
Sources
AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0)
hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.persian-handwritten-digits
Persian Handwritten Digits (Farsi)
80,000 grayscale images of handwritten Persian (Farsi) digits — ۰۱۲۳۴۵۶۷۸۹ —
organized as an ImageFolder dataset with 10 classes (0–9), 8,000 images per class.
Each image is a 28×28 grayscale PNG of a single digit.
Classes
Class
Count
0 (۰)
8,000
1 (۱)
8,000
2 (۲)
8,000
3 (۳)
8,000
4 (۴)
8,000
5 (۵)
8,000
6 (۶)
8,000
7 (۷)
8,000
8 (۸)
8,000
9 (۹)
8,000
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Mehdinmz/persian-handwritten-digits.persian-handwriting-pages-369k
Persian Handwriting Pages 369K
Full-page Persian handwriting compositions on scanned paper backgrounds.
Each row deliberately has only two fields:
image: the composed full-page image
label: its complete line-separated Persian transcription, ordered from top to bottom
The pages are composed from labeled real handwriting crops with page-level ink normalization,
controlled RTL layout variation, collision prevention, and exact transcription provenance.
The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.persian-ocr-bench-submitted10-bbox-crops
Persian OCR benchmark — selected submitted bbox crops
This dataset contains the non-empty OCR bboxes from the ten explicitly selected
submitted pages in persian_ocr_bench_bbox_review.
Each row is one PNG crop. gold_text is the current editable OCR content from
the live Argilla bbox field (content_text). Geometry is stored both as source
page pixels and as percentages of the source page. The original record ID,
external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.bookroom-persian-book-covers-and-titlesPersian_Arabic_TextLine_Image_Ocr_Mediumpersian-ocr-gemini37-wins-bina-misses-viewer
Gemini 3.7 exact / Bina miss OCR crops
53 non-empty bbox crops from the PersianVLM submitted-10 benchmark where
google/gemini-3.7-flash was normalized-exact and Bina Koochik 0.1 was not.
This is a minimal Hugging Face ImageFolder dataset for reliable viewer support.
It contains exactly two columns: image and ocr. The ocr value is Gemini's
actual output for the corresponding crop.
persian-ocr-gemini37-wins-bina-misses-v2
Gemini 3.7 exact / Bina miss OCR crops
53 non-empty bbox crops from the PersianVLM submitted-10 benchmark where
google/gemini-3.7-flash produced a normalized-exact transcription and
Bina Koochik 0.1 did not.
This release uses the standard Hugging Face ImageFolder layout. The viewer's
first column is image, followed by ocr containing Gemini's actual output.
Gold and Bina outputs are retained for comparison.
persian-ocr-bench
Persian OCR Evaluation Dataset
This benchmark contains paired Persian document images and UTF-8 transcription
targets for OCR evaluation. Each JSONL row references one image under
bench_data/images/ and contains its transcription in the text field.
Schema
image: image path relative to the repository
id: stable image identifier
page: page number, currently 1
type: evaluation item type, currently transcription
text: reference transcription
language: fa
checked:… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-ocr-bench.persian-ocr-garshasp-70c
Persian OCR Garshasp: Large-Scale Synthetic Persian OCR Dataset
Persian-OCR-Garshasp is a large-scale Persian (Farsi) OCR dataset containing about 2.6M image–text pairs.
It is designed for training and evaluating Optical Character Recognition (OCR) and scene text recognition models
for Persian text in the wild.
Each sample consists of:
an RGB image with height 48 px and width 640 px
up to 70 Persian (Farsi) characters as the ground-truth text label
a style field describing the… See the full description on the dataset page: https://huggingface.co/datasets/AliShafiee2003/persian-ocr-garshasp-70c.persian-handwriting-ocr
Persian Handwriting OCR Dataset
Dataset Summary
A standardized dataset of Persian (Farsi) handwritten pages with word-level
bounding-box annotations and transcriptions. The dataset is page-level:
each sample is a full page scan; annotations are one row per word bbox on
that page. This is the most flexible form -- users can train page-level OCR,
word detection (DBNet/PaddleOCR), or derive word/line crops as needed.
Pages: 1115 scanned pages (canonical IDs… See the full description on the dataset page: https://huggingface.co/datasets/MR3z4/persian-handwriting-ocr.Persian_Pixel
Persian Pixel
Persian Pixel is a synthetic optical character recognition (OCR) dataset for Persian / Farsi (fa), in which Unicode text is rendered to images and paired with its exact transcription. It is built for OCR recognition, image-to-text modeling, fine-tuning, and evaluation workflows that need clean, controllable image/label pairs at scale.
Because the text is rendered programmatically, every image ships with a perfectly aligned ground-truth label — making the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Persian_Pixel.persian-ocr-benchmark
Persian OCR Evaluation Dataset
This benchmark contains paired Persian document images and UTF-8 transcription
targets for OCR evaluation. Each JSONL row references one image under
bench_data/images/ and contains its transcription in the text field.
Schema
image: image path relative to the repository
id: stable image identifier
page: page number, currently 1
type: evaluation item type, currently transcription
text: reference transcription
language: fa
checked:… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-ocr-benchmark.persian-ocr-gemini37-wins-bina-misses
Gemini 3.7 vs Bina OCR comparison crops
53 non-empty bbox crops where Gemini 3.7 Flash was normalized-exact and Bina
Koochik 0.1 was not. The first columns are image, gemini_ocr, bina_ocr,
and gold_text for direct visual and OCR comparison.
Persian_Arabic_TextLine_Image_Ocr_Small
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/mohajesmaeili/Persian_Arabic_TextLine_Image_Ocr_Small.persian-ocr-double-benchmark
Persian OCR Double Benchmark
A frozen, leakage-controlled Persian OCR benchmark with two equally weighted splits:
printed: 6,669 rows carved from Reza2kn/persian-printed-ocr-3.5m at ba02f36c3d496838d8fad9aff352b77763af1ea4.
handwriting: 6,669 rows carved from Reza2kn/persian-handwriting-pages-3.69m at b114f0a36a6a2e397bc93517dd084431bfca2329.
The exact rows were uploaded here before being removed from their source training repositories.
Each row retains its original repository… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-double-benchmark.Persian-OCR-230k.
├── Images/
├── train.csv (184k)
└── test.csv (46k)
persian-word-ocrpersian_images_doc_tags_all_v12Persian-Image-Captioning
Dataset Card for "Persian-Image-Captioning"
More Information needed
persian_images_doc_tags_all_v11persian-ocr-gemini37-wins-bina-misses-parquet
Gemini 3.7 exact / Bina miss OCR crops
53 bbox crops. Columns: image and ocr.
crrn_persian_license_v3Persian_English_Roadsign_OCR_Dataset_Relabeledpersian-tweets-2024
Dataset Description
This dataset contains high-engagement Persian language tweets collected from Twitter/X during 2024. The dataset includes comprehensive tweet metadata and user information, making it valuable for various NLP tasks, social media analysis, and Persian language processing research.
Dataset Details
Size: 900 tweets
Language: Persian (Farsi)
Time Period: 2024
Collection Criteria:
Language: Persian
Minimum Likes: 1,000+
Date Range: January 1, 2024… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-tweets-2024.persian_images_doc_tags_all_v15
