CoolFace
Datasetpublic

zirak-ai/PashtoOCR

PsOCR - Pashto OCR Dataset ๐ŸŒ Zirak.ai    |   ๐Ÿค— HuggingFace    |    GitHub    |    Kaggle    |   ๐Ÿ“‘ Paper PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR Introduction PsOCR is aโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/zirak-ai/PashtoOCR.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
6likes228downloads
Dataset Card

PsOCR - Pashto OCR Dataset

<p align="center"> <img src="logo.png" width="200"/> <p> <p align="center"> ๐ŸŒ <a href="https://zirak.ai/downloads/"><b>Zirak.ai</b></a> &nbsp&nbsp | &nbsp&nbsp๐Ÿค— <a href="https://huggingface.co/datasets/zirak-ai/Pashto-OCR"><b>HuggingFace</b></a> &nbsp&nbsp | &nbsp&nbsp <a href="https://github.com/zirak-ai/PashtoOCR"><b>GitHub</b></a> &nbsp&nbsp | &nbsp&nbsp <a href="https://www.kaggle.com/datasets/drijaz/PashtoOCR"><b>Kaggle</b></a> &nbsp&nbsp | &nbsp&nbsp๐Ÿ“‘ <a href="https://doi.org/10.1016/j.asej.2026.104024"><b>Paper</b></a> </p>

[PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language](https://doi.org/10.1016/j.asej.2026.104024)

The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR

Introduction

  • โ€”PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
  • โ€”This is the first publicly available comprehensive Pashto OCR dataset consisting of One Million synthetic images annotated at word, line, and document-level granularity, covering extensive variations including 1000 unique font families, diverse colors, image sizes, and text layouts.
  • โ€”PsOCR includes the first publicly available OCR benchmark comprising 10,000 images, facilitating systematic evaluation and comparison of OCR systems for the low-resource Pashto.
  • โ€”We conducted a pioneering evaluation and comparison of state-of-the-art LMMs on Pashto OCR, providing crucial insights into their zero-shot capabilities, strengths, and limitations for low-resource languages written in Perso-Arabic scripts.
  • โ€”<span style="color:red"><b>โ„น๏ธ On this repo, only the test set (benchmark), containing 10K images is publicly available. To access the complete dataset of 1 Million images, please **visit this link.**</b></span>

<p align="center"> <img src="fig1.jpg" width="50%"/><br> Performance Comparison of various LMMs on PsOCR Benchmark <p>

Granularity

The annotation information is provided at three levels of granularity: page-level, line-level, and token-level

<p align="center"> <img src="fig2.jpg" width="100%"/> <p>

Font Variation

PsOCR features 1000 unique font families, a few of them are shown here. <p align="center"> <img src="fig4.jpg" width="100%"/> <p>

Citation

If you found our work useful, please feel free to cite it:

bibtex
@article{Haq2026PsOCR,
  title   = {PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-Resource Pashto Language},
  journal = {Ain Shams Engineering Journal},
  volume  = {17},
  number  = {3},
  pages   = {104024},
  year    = {2026},
  issn    = {2090-4479},
  doi     = {10.1016/j.asej.2026.104024},
  url     = {https://www.sciencedirect.com/science/article/pii/S2090447926000511},
  author  = {Ijazul Haq and Yingjie Zhang and Muhammad Saqib}
}

Contact

Website: https://zirak.ai/ Email Address: contact@zirak.ai, mail@ijaz.me