datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.arabic-ocrarocrbench_arabicocrPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
arabic_ocr_synth_2arabic-ocr-imagesarabic-ocr-datasetv3arabic-ocr
ocr-data
Alshams, The largest Arabic OCR dataset at the word level.
This dataset is specifically designed for fine-grained word-level OCR tasks, providing precise word-level bounding box annotations for each image.
Each word is annotated with pixel-accurate localization, enabling tasks such as text detection, text recognition, and end-to-end OCR.
This repository currently provides an Arabic OCR dataset.
📦 Available Datasets
The technical specifications of each… See the full description on the dataset page: https://huggingface.co/datasets/craneset/arabic-ocr.arabic_ocr_merged_datasetarabic-ocr-books
Scanned Book Page Images
This dataset contains page images rendered from 9 PDF file(s).
Dataset structure
One image per PDF page.
One subfolder per source PDF.
data/metadata.jsonl contains technical provenance for each page.
Image settings
DPI: 300
Maximum width: 1600
Grayscale: True
Contrast factor: 1.4
JPEG quality: 95
White-margin cropping: True
Intended use
OCR, document understanding, knowledge extraction, fine-tuning,
and… See the full description on the dataset page: https://huggingface.co/datasets/mustaphaelkady/arabic-ocr-books.arabic-ocr-markdown-dataset
Arabic Document OCR Markdown Dataset
Dataset Description
This dataset contains 1,256 pairs of document images and their corresponding Markdown representations, specifically designed for Arabic document OCR tasks. The dataset is intended for training and evaluating models that convert document images into structured Markdown text (image-to-markdown OCR).
Features
The dataset consists of two main features:
image: Document images in various formats
markdown:… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/arabic-ocr-markdown-dataset.arabic_ocr_synth_3arabic-ocr-datasetarabic_ocr_datasetarabic_arabicocrarabic_ocrisiarabic-ocr-dataset-7arabic-ocr-dataset-2arabic-ocr-dataset-5arabic-ocr-datasetv2arabic-ocrr-datasetarabic-ocr-dataset-4arabic-ocr-doctags-datasetarabic-ocr-dataset-1arabic-ocr-images
