ArabicOCR
arabic-ocr-labelsarabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.arabic-ocrarocrbench_arabicocrPlease see paper & code for more information:
https://github.com/mbzuai-oryx/KITAB-Bench
https://arxiv.org/abs/2502.14949
arabic_ocr_synth_2arabic-ocr-images
