datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PashtoOCR
PsOCR - Pashto OCR Dataset
🌐 Zirak.ai
| 🤗 HuggingFace
| GitHub
| Kaggle
| 📑 Paper
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
The dataset is also available at: https://www.kaggle.com/datasets/drijaz/PashtoOCR
Introduction
PsOCR is a… See the full description on the dataset page: https://huggingface.co/datasets/zirak-ai/PashtoOCR.zamai-pashto-vision
ZamAI Pashto Vision
Languages: psLicense: cc-by-4.0Task categories: image-to-text, image-classificationSize categories: 1K<n<10K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for image-to-text, image-classification tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-vision")
print(dataset)
Configs
pashto_captions: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-vision.pashto_multimodal_v1
Pashto Multimodal v1: A Decade of Pashtun Digital Activism
Dataset Summary
This is a curated, multilingual, and multimodal dataset of ~5,900 high-quality image-text pairs, extracted from the personal X (Twitter) archive of @Pashto_lab. This archive represents over a decade of active digital advocacy (since 2012) by a Pashtun AI researcher and activist based in Japan.
The dataset is a rich, firsthand chronicle of contemporary Pashtun socio-political discourse… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto_multimodal_v1.
