datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zamai-pashto-documents
ZamAI-Pashto Documents
This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language.
Project Structure
data/: Contains scanned documents, extracted text, translations, and summaries.
annotations/: OCR bounding boxes, handwriting labels, and domain tags.
scripts/: OCR processing, text cleaning, and translation alignment scripts.
configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.Pashto-Textbooks-PDFs-Corpus
Pashto Textbooks and PDFs Corpus
Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.zamai-pashto-video
ZamAI Pashto Video
ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation.
Dataset Summary
The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows.
Languages
Pashto
Dari
English
Modalities
Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.Pashto-Poetry
Pashto Poetry Dataset
Dataset Description
This dataset contains Pashto poetry from various poets. Each entry includes a line of poetry and the corresponding poet's name.
Dataset Structure
text: A line of Pashto poetry (string).
poet: The name of the poet (string).
Poets Included
The dataset includes works by poets such as Rahman Baba, Ghani Khan, Hamza Baba, and others. The complete dataset details are as follows:
No.
Poet
Couplet Count… See the full description on the dataset page: https://huggingface.co/datasets/AliMuhammad73/Pashto-Poetry.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.zamai-pashto-vision
ZamAI-Pashto Vision
This repository contains the dataset and scripts for the ZamAI-Pashto vision project, focusing exclusively on Pashto-related visual data, including scene captions, signs, and cultural objects.
Project Structure
images/: Raw and cleaned images, including signs, boards, and documents with Pashto text.
annotations/: Pashto captions, object bounding boxes, and scene labels.
metadata/: Cultural and location-specific tags.
scripts/: Image… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-vision.zamai-pashto-emotion
ZamAI Pashto Emotion
ZamAI Pashto Emotion is a dataset scaffold for emotional and cultural context modeling across Pashto-centered text, dialogue, and speech data.
Dataset Summary
The repository is organized for emotional language understanding, culturally grounded annotation, formality tracking, and speaker-aware metadata.
Languages
Pashto
Dari
English
Modalities
Text
Dialogue
Audio
Project Structure… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-emotion.pashto-mental-health-counseling-3k
🧠 Pashto Mental Health Counseling 3K
This dataset is a specialized collection of 3,000 conversational pairs focused on mental health counseling, translated and culturally adapted into Pashto. It is designed to train LLMs to provide empathetic, supportive, and culturally relevant responses in a therapeutic context.
🌟 Overview
Mental health resources in Pashto are scarce. This dataset aims to bridge that gap by providing high-quality counseling dialogues. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-counseling-3k.zamai-pashto-documents
Pashto
Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-documents")
print(dataset)
Citation
@misc{zamai_pashto_data,
title = {{Pashto}},
author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.pashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.pashto-opus-5k-reasoning-max
🚀 Pashto OPUS 5K Reasoning Max
This dataset is a high-quality collection of 5,000 reasoning-focused pairs, derived from the OPUS corpus and enhanced for Pashto Language Models. It is specifically curated to push the boundaries of "Chain-of-Thought" (CoT) and logical deduction in the Pashto language.
🌟 Overview
While standard OPUS data is often used for simple translation, Pashto-OPUS-5K-Reasoning-Max takes it a step further by focusing on complex instructions and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-opus-5k-reasoning-max.pashto-mental-health-support
Pashto Mental Health Support Dataset
Overview
Pashto-Mental-Health-Support is a specialized conversational dataset consisting of 172 high-quality samples focused on mental health awareness, emotional support, and psychological well-being. This dataset is a localized and translated version of the heliosbrahma/mental_health_chatbot_dataset.
This repository marks a strategic expansion of the iPashto.ai ecosystem, moving from logical reasoning into the domain of Emotional… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-support.pashto-otter-cot
Pashto-Otter-CoT Dataset
Overview
Pashto-Otter-CoT is a first-of-its-kind dataset specifically designed to bring Chain-of-Thought (CoT) Reasoning capabilities to Pashto language models. This dataset is a translated and curated version of a subset of the brendan-gho/gemma4b_paraphrased_otter_cot.
This project is part of the iPashto.ai initiative, led by Nassim الله (nassimjp), aimed at creating high-quality linguistic resources for the Pashto language.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-otter-cot.pashto-dragon-1k-cot
Pashto-Dragon-1K-CoT Dataset
Overview
Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot.
This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dragon-1k-cot.pashto-qwen-1k-cot
Pashto-Qwen-1K-CoT Dataset
Overview
Pashto-Qwen-1K-CoT is a high-quality reasoning dataset consisting of 1,024 samples, specifically curated to enhance the Chain-of-Thought (CoT) capabilities of Pashto language models. This dataset is a translated version of a subset from brendan-gho/qwen3b_paraphrased_cat_cot.
By focusing on "Reasoning" rather than just "Information," this dataset helps models like Baran and Roshan develop logical thinking paths in the Pashto language.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-qwen-1k-cot.
