zirak-ai/PashtoOCR
PsOCR - Pashto OCR Dataset ๐ Zirak.ai | ๐ค HuggingFace | GitHub | Kaggle | ๐ Paper PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR Introduction PsOCR is aโฆ See the full description on the dataset page: https://huggingface.co/datasets/zirak-ai/PashtoOCR.
PsOCR - Pashto OCR Dataset
<p align="center"> <img src="logo.png" width="200"/> <p> <p align="center"> ๐ <a href="https://zirak.ai/downloads/"><b>Zirak.ai</b></a>    |   ๐ค <a href="https://huggingface.co/datasets/zirak-ai/Pashto-OCR"><b>HuggingFace</b></a>    |    <a href="https://github.com/zirak-ai/PashtoOCR"><b>GitHub</b></a>    |    <a href="https://www.kaggle.com/datasets/drijaz/PashtoOCR"><b>Kaggle</b></a>    |   ๐ <a href="https://doi.org/10.1016/j.asej.2026.104024"><b>Paper</b></a> </p>
[PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language](https://doi.org/10.1016/j.asej.2026.104024)
The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR
Introduction
- PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
- This is the first publicly available comprehensive Pashto OCR dataset consisting of One Million synthetic images annotated at word, line, and document-level granularity, covering extensive variations including 1000 unique font families, diverse colors, image sizes, and text layouts.
- PsOCR includes the first publicly available OCR benchmark comprising 10,000 images, facilitating systematic evaluation and comparison of OCR systems for the low-resource Pashto.
- We conducted a pioneering evaluation and comparison of state-of-the-art LMMs on Pashto OCR, providing crucial insights into their zero-shot capabilities, strengths, and limitations for low-resource languages written in Perso-Arabic scripts.
- <span style="color:red"><b>โน๏ธ On this repo, only the test set (benchmark), containing 10K images is publicly available. To access the complete dataset of 1 Million images, please **visit this link.**</b></span>
<p align="center"> <img src="fig1.jpg" width="50%"/><br> Performance Comparison of various LMMs on PsOCR Benchmark <p>
Granularity
The annotation information is provided at three levels of granularity: page-level, line-level, and token-level
<p align="center"> <img src="fig2.jpg" width="100%"/> <p>
Font Variation
PsOCR features 1000 unique font families, a few of them are shown here. <p align="center"> <img src="fig4.jpg" width="100%"/> <p>
Citation
If you found our work useful, please feel free to cite it:
@article{Haq2026PsOCR,
title = {PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-Resource Pashto Language},
journal = {Ain Shams Engineering Journal},
volume = {17},
number = {3},
pages = {104024},
year = {2026},
issn = {2090-4479},
doi = {10.1016/j.asej.2026.104024},
url = {https://www.sciencedirect.com/science/article/pii/S2090447926000511},
author = {Ijazul Haq and Yingjie Zhang and Muhammad Saqib}
}Contact
Website: https://zirak.ai/ Email Address: contact@zirak.ai, mail@ijaz.me
