datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ZamAI-Pashto-Dataset-Cleaned
ZamAI Pashto Dataset Cleaned
Languages: psLicense: apache-2.0Task categories: text-classification, text-generation, question-answeringSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-classification, text-generation, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Dataset-Cleaned")
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Dataset-Cleaned.afghanistan-post-2021-pashto-dataset
Afghanistan Post-2021 Pashto Dataset
Dataset Description
This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for:
Training and fine-tuning Pashto large language models (LLMs)
Question-answering tasks
Research on Afghanistan's post-2021 developments
Low-resource language AI development
The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.pashto-instruct-dataset
Pashto Instruct Dataset
This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions.
Dataset Structure
Each sample in the dataset contains the following fields:
id: Unique identifier for the sample.
messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.Driving-License-Pashto-QA
🚗 Driving License Pashto QA Dataset (د موټر چلولو جواز - پښتو ډاټاسیټ)
This dataset contains translated Pashto Questions and Answers related to Driving License exams and road traffic rules. It was originally sourced/translated from Persian driving theory test questions and formatted for fine-tuning Large Language Models (LLMs) and training Chat completions models.
دا ډاټاسیټ د موټر چلولو د لایسنس/جواز او ترافیکي مقرراتو پښتو پوښتنې او ځوابونه لري، چې له فارسي منبع څخه په معیاري… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Driving-License-Pashto-QA.pashto-algebra
د پښتو الجبرا پروژه 🤖🇦🇫📚🧠
د افغان نجونو لپاره ډالۍ 💝
"تعلیم یو حق دی، نه مرسته." ✨
هغو زړورو افغان نجونو ته چې له ښوونځي او کتابونو څخه محرومې دي — دا پروژه ستاسو لپاره ده! 🇦🇫❤️
📖 د پروژې په اړه
Pashto Algebra Dataset په پښتو ژبه کې لومړی او تر ټولو لوی ګام په ګام ریاضي ډیټاسیټ دی.
دا پروژه د هغو افغان ماشومانو لپاره جوړه شوې چې په ځانګړې توګه نجونې چې په افغانستان کې له ښوونځي تګ څخه منع دي او هلکان چې په لرو پرتو سیمو کې اوسي.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-algebra.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.ZamAI-Pashto-Mega-Dataset
ZamAI Pashto Mega Dataset
Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.Pashto-100k-Pairs
Qehwa AI - Pashto 100K Fine-Tuning Dataset
Overview
Qehwa AI presents a large-scale Pashto instruction tuning dataset containing 100,000+ high-quality instruction-response pairs designed for supervised fine-tuning, conversational AI, and downstream NLP tasks.
This dataset was created to advance AI research for the Pashto language, a significantly underrepresented low-resource language spoken by millions worldwide. The dataset covers more than 20 diverse domains and is… See the full description on the dataset page: https://huggingface.co/datasets/junaid008/Pashto-100k-Pairs.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.Medic_Chat-Pashto
📦 Dataset Summary
ژبه: Pashto
ډول: Chat‑style SFT (Supervised Fine‑Tuning)
موضوع: Traditional Chinese Medicine (TCM)
ریکارډونه: شاوخوا 10.8k
فورمټ: JSONL — messages: [{role, content}, ...]
لایسنس: CC‑BY‑NC‑4.0
کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs
🧬 Data Structure
هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده:
{
"messages": [
{"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.ZamAI-Pashto-High-Qualituly-Dataset
ZamAI Pashto High Quality Dataset
Languages: psLicense: mitTask categories: text-generation, question-answering, translationSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, question-answering, translation tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-High-Qualituly-Dataset")
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-High-Qualituly-Dataset.pashto-legal-qa-chat
Pashto Legal QA Chat Dataset
⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.)
Dataset Overview
The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.pashto-math
pashto-math
Dataset Summary
pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي.
Dataset Structure
هره نمونه د ChatML-style SFT په بڼه ده:
{
"id": "000005",
"messages": [
{ "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.Pashto-Medical-o1-Reasoning-SFT-Dataset
Pashto Medical o1 Reasoning SFT Dataset
This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models.
Dataset Structure
The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response.
Data Fields
Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.pashto-sociology
# Dataset Card for Pashto Sociology Dataset
## Dataset Description
- **Homepage:** [N/A]
- **Repository:** [Nassimjp/pashto-sociology](https://huggingface.co/datasets/nassimjp/pashto-sociology)
- **Paper:** [N/A]
- **Leaderboard:** [N/A]
- **Point of Contact:** [N/A]
### Dataset Summary
This dataset contains a collection of 100 sociological dialogue samples in the Pashto language. It is designed to facilitate research and development of conversational AI, natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-sociology.perfect-pashto-reasoning-sft
Perfect Pashto Reasoning SFT Dataset
د پښتو ژبې لپاره تر ټولو پاک او لوړ کیفیت لرونکی Reasoning Dataset
📌 الوتنه (Overview)
دا ډېټاسیټ د Magpie-Pro-300K-Filtered dataset پر بنسټ جوړ شوی دی چې د Pashto LLM او AI ټولنې لپاره په بشپړ ډول نوي سره انجنیر شوی او پروسس شوی دی.
ټول ډاټا په اتومي او لاین په لاین ډول ژباړل شوې او په لوړ کیفیت سره reformatted شوې ترڅو د alignment-handbook سره مستقیم مطابقت ولري. هدف یې د پښتو ژبې نوي نسل ماډلونو (لکه Rawanاو Ghanam لړۍ)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/perfect-pashto-reasoning-sft.pashto-quotes-dataset
Pashto Quotes Dataset with Chain-of-Thought Reasoning
Dataset Description
This dataset contains 990 Persian (Farsi/Dari) quotes from various philosophers, writers, and thinkers, each accompanied by:
A Pashto translation of the quote
5 step-by-step reasoning steps (Chain-of-Thought) in Pashto explaining the quote's meaning
A concise conclusion in Pashto summarizing the key insight
Each entry is designed to help language models learn reasoning, translation, and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-quotes-dataset.ZamAI-Pashto-MegaDataset-v1
ZamAI Pashto Mega Dataset v1
Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-MegaDataset-v1")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-MegaDataset-v1.pashto-fallacy-dataset
Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ)
The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts.
The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.Pashto-Reasoning-RescueBench
Pashto-Reasoning-RescueBench
A High-Quality Pashto Chain-of-Thought Dataset for Rescue, Survival & Emergency Preparedness
Dataset Description
Pashto-Reasoning-RescueBench is a Pashto-language dataset created for training language models with strong reasoning in survival, bushcraft, and emergency situations.
Origin & Creation Process
Base Dataset: Derived from mattwesney/CoT_Reasoning_Bushcraft_Survival
The original questions were translated/adapted into… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Reasoning-RescueBench.Pashto-Clean-100k-Pairs.QA
Pashto‑Clean‑100k‑Pairs.QA
A curated collection of 100,000 Pashto question–answer pairs, cleaned and normalized for general‑purpose Pashto NLP training.This dataset focuses on broad coverage, topic diversity, and clean formatting, without synthetic reasoning or long‑context generation.
Dataset Summary
Pashto-Clean-100k-Pairs.QA contains short, direct QA pairs across 70+ everyday topics:
Daily life
Community
Education
Work
Nature
Safety
Culture… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Clean-100k-Pairs.QA.Pashto-Ethical-Bench_Base
Pashto Ethical Benchmark Base (Pashto-Ethical-Bench_Base)
Overview
Pashto-Ethical-Bench_Base is a high-quality, carefully curated dataset containing 3,604 instruction-response pairs in Pashto (پښتو).
The dataset focuses on criminal law, evidence rules, investigation procedures, forensic science, presumption of innocence, and ethical/legal reasoning. It is designed to:
Improve safety and alignment of Pashto-language LLMs
Evaluate cultural and legal understanding
Support… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Ethical-Bench_Base.knowledge_qa_in_pashto
د پوهې QA ډیټا سیټ
دا ډیټا سیټ د مطالعې ډیټا سیټ دی چې د مختلفو موضوعاتو څخه د پوښتنې ځواب مثالونه لري.
په اړه
دا پروژه اوس مهال د زده کړې او پراختیا مرحله کې ده. زه د ډیټا سیټ چمتو کولو پرمهال د مصنوعي استخباراتو او ډیټا سیټ جوړولو تمرین کوم.
په ډیټا سیټ کې ځینې پوښتنې د ChatGPT په کارولو سره رامینځته شوي، ځینې یې د Qwen په کارولو سره، او ځینې یې زما لخوا چمتو شوي.
پوښتنې او ځوابونه مختلف موضوعات پوښي. د مثال په توګه:
عمومي پوهه
ریاضی
ساینس
کیمیا
کمپیوټر ساینس… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/knowledge_qa_in_pashto.pashto-mental-health-counseling-3k
🧠 Pashto Mental Health Counseling 3K
This dataset is a specialized collection of 3,000 conversational pairs focused on mental health counseling, translated and culturally adapted into Pashto. It is designed to train LLMs to provide empathetic, supportive, and culturally relevant responses in a therapeutic context.
🌟 Overview
Mental health resources in Pashto are scarce. This dataset aims to bridge that gap by providing high-quality counseling dialogues. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-counseling-3k.pashto-legal-reasoning
Pashto Legal Reasoning Dataset (iPashto.ai)
دا ډېټاسیټ د پښتو ژبې لپاره یو له خورا پرمختللو حقوقي سرچینو څخه دی چې په ځانګړي ډول د Reasoning (استدلالي) ماډلونو لکه Baran او Ghanam لپاره ډیزاین شوی. دا ډېټا یوازې ساده پوښتنې او ځوابونه نه دي، بلکې هر ریکارډ د یوې حقوقي مسلې په اړه څو اړخیز استدلال وړاندې کوي.
📋 عمومي ځانګړتیاوې
ساحه (Domain): نړیوال حقوق، مدني قانون، جنایي قانون، او سوداګریز قوانین (WTO).
جوړښت (Structure): هره پوښتنه درې ډوله استدلال لري:
حقوقي استدلال… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-reasoning.pashto-health-chat
🩺 Pashto-Health-Chat Dataset
Pashto‑Health‑Chat یو منظم، پاک، او د Pashto روغتیايي پوښتنو–ځوابونو چټ ډیټاسیټ دی چې دPashto طبي مرستندویانو، کلینیکي Reasoning، او SFT روزنې لپاره جوړ شوی.
دا ډیټاسیټ د LLaMA‑Chat, Qwen‑Chat, Ministral‑Instruct او نورو چټ ماډلونو سرهپه بشپړ ډول سازګار دی.
📦 Dataset Format
هره نمونه د چټ په بڼه ده:
{
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]
}
هیڅ system prompt نشته… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-health-chat.pashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.
