NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset
QARI Markdown Mixed Dataset ๐ Dataset Summary The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding. This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition. This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7kโฆ See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.
QARI Markdown Mixed Dataset
<div align="center">
</div>
<div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/628f7a71dd993507cfcbe587/TxzHjVy6NsdmcghXqVH.png" alt="QARI Logo" width="400"> </div>
๐ Dataset Summary
The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding. This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition.
This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k validation, 3.7k test) with corresponding ground truth text in HTML/Markdown format, featuring:
- ๐ค Full diacritical marks (tashkeel) support
- ๐ Mixed font sizes within documents (headers, body text, annotations)
- ๐จ 12 distinct Arabic fonts ranging from common Naskh to ornate calligraphic styles
- ๐ Realistic document layouts with structural HTML tags
- ๐ผ๏ธ Multiple text sources including Basma2423 and YoussefAnwar Arabic news
๐ฏ Intended Use
This dataset is specifically designed for:
- Training OCR models that need to understand document structure
- Fine-tuning vision-language models for Arabic text recognition
- Developing systems that preserve formatting and layout information
- Research in Arabic document analysis and understanding
๐ Dataset Statistics
๐ง Data Generation Pipeline
<div align="center">
</div>
๐ Model Performance
When used to train QARI v0.3, this dataset enables:
Key Advantages:
- ๐ Superior layout understanding compared to plain text models
- ๐ท๏ธ HTML tag preservation for structured document conversion
- โก Resource efficient - 5x less training time than larger datasets
- ๐ฏ Specialized performance for document structure tasks
Citation
@article{wasfy2025qari,
title={QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation},
author={Wasfy, Ahmed and Nacar, Omer and Elkhateb, Abdelakreem and Reda, Mahmoud and Elshehy, Omar and Ammar, Adel and Boulila, Wadii},
journal={arXiv preprint arXiv:2506.02295},
year={2025}
}