datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr_arabic_books
Arabic OCR Books Dataset (ocr_arabic_books)
This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions.
📚 Master Book Inventory
Total running pages in repository: 181,427
#
Book Name (English)
Book Name (Arabic)
Subset / Config Name
Page Count
Image Index Range
1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.synth_shamela_ocr_arabic_books
Synthetic Arabic Books Dataset
Structured book pages rendered dynamically with style, font, and degradation variations.
Persian_Arabic_TextLine_Image_Ocr_Mediumarabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.Historical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.arabic-ocrarocrbench_arabicocrPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
arabic_ocr_synth_2Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.Persian_Arabic_TextLine_Image_Ocr_Small
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/mohajesmaeili/Persian_Arabic_TextLine_Image_Ocr_Small.arabic-ocr-imagesarabic-ocr-datasetv3arabic-ocr
ocr-data
Alshams, The largest Arabic OCR dataset at the word level.
This dataset is specifically designed for fine-grained word-level OCR tasks, providing precise word-level bounding box annotations for each image.
Each word is annotated with pixel-accurate localization, enabling tasks such as text detection, text recognition, and end-to-end OCR.
This repository currently provides an Arabic OCR dataset.
📦 Available Datasets
The technical specifications of each… See the full description on the dataset page: https://huggingface.co/datasets/craneset/arabic-ocr.arabic_ocr_merged_datasetarabic-names-synthetic-ocrarabic-handwritten-ocr-eval
Arabic OCR Evaluation Dataset
This dataset contains Arabic paragraph images paired with ground truth text.
Dataset Structure
images/
train.csv
validation.csv
test.csv
Task
OCR evaluation for Arabic printed text.
Metrics
Models are evaluated using:
CER (Character Error Rate)
WER (Word Error Rate)
Size
~600 images
Use case
Evaluation of Arabic OCR models.
qari-arabic-ocr-10karabic-ocr-books
Scanned Book Page Images
This dataset contains page images rendered from 9 PDF file(s).
Dataset structure
One image per PDF page.
One subfolder per source PDF.
data/metadata.jsonl contains technical provenance for each page.
Image settings
DPI: 300
Maximum width: 1600
Grayscale: True
Contrast factor: 1.4
JPEG quality: 95
White-margin cropping: True
Intended use
OCR, document understanding, knowledge extraction, fine-tuning,
and… See the full description on the dataset page: https://huggingface.co/datasets/mustaphaelkady/arabic-ocr-books.arabic-ocr-markdown-dataset
Arabic Document OCR Markdown Dataset
Dataset Description
This dataset contains 1,256 pairs of document images and their corresponding Markdown representations, specifically designed for Arabic document OCR tasks. The dataset is intended for training and evaluating models that convert document images into structured Markdown text (image-to-markdown OCR).
Features
The dataset consists of two main features:
image: Document images in various formats
markdown:… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/arabic-ocr-markdown-dataset.arabic_ocr_synth_3arabic-ocr-datasetarabic_ocr_datasetarabic_arabicocrarabic_ocrisiarabic-ocr-dataset-7arabic-ocr-dataset-2Qari-OCR-0.2.2-Arabic-2B_Qari-0.2-eval-tashkilarabic-ocr-dataset-5arabic-manuscript-ocrarabic-ocr-datasetv2
