datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.afghanistan-post-2021-pashto-dataset
Afghanistan Post-2021 Pashto Dataset
Dataset Description
This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for:
Training and fine-tuning Pashto large language models (LLMs)
Question-answering tasks
Research on Afghanistan's post-2021 developments
Low-resource language AI development
The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.pashto-instruct-dataset
Pashto Instruct Dataset
This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions.
Dataset Structure
Each sample in the dataset contains the following fields:
id: Unique identifier for the sample.
messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.Da-Ploshi-OpenThoughts_Cache
Da-Ploshi OpenThoughts Cache
Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases.
The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs.
Dataset Structure
Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.pashto-reasoning-chat-dataset
Pashto Reasoning Chat Dataset
A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto.
📊 Dataset Structure
Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities:
system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.Pashto-Brain-Extraction-Dataset
🧠 Pashto Brain Extraction Dataset
A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer.
Keep the brain 🧠 — throw away the mouth 🗣️
🎯 Purpose
A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer.
Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.pashto-wikipedia-sft
Pashto Wikipedia SFT Dataset
This dataset is a cleaned, structured, and context-anchored variant of the Pashto Wikipedia corpus, formatted explicitly for Supervised Fine-Tuning (SFT) and Continual Pre-Training (CPT).
By restructuring raw encyclopedic data into explicit source-grounding prompts, this dataset trains Large Language Models (LLMs) to couple their factual generation directly with reference variables (URLs and Titles), mitigating hallucination tendencies in… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-wikipedia-sft.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.da-tolanpoheni-pokhtany
Da Tolanpoheni Pokhtany (د ټولنیزې پوښتنې)
Overview
da-tolanpoheni-pokhtany is a specialized dataset curated for Pashto language processing, natural language understanding, and related machine learning applications.
Structure
Data Fields
text / standard fields containing the primary dataset contents.
Data Instances
An example instance from the dataset:
{
"text": "Sample text entry..."
}
Usage
You can… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/da-tolanpoheni-pokhtany.Medic_Chat-Pashto
📦 Dataset Summary
ژبه: Pashto
ډول: Chat‑style SFT (Supervised Fine‑Tuning)
موضوع: Traditional Chinese Medicine (TCM)
ریکارډونه: شاوخوا 10.8k
فورمټ: JSONL — messages: [{role, content}, ...]
لایسنس: CC‑BY‑NC‑4.0
کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs
🧬 Data Structure
هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده:
{
"messages": [
{"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.Sindhi-Reasoning-Chat-Dataset
# 📘 **Sindhi‑Reasoning‑Chat‑Dataset**
A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages.
---
## 🧠 **Dataset Summary**
Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.questions
📝 Overview
Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی.
🎯 Purpose
دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی:
Pashto instruction-tuning
Pashto question-answering
Pashto reasoning
Pashto dialogue modeling
Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.pashto-legal-qa-chat
Pashto Legal QA Chat Dataset
⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.)
Dataset Overview
The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.pashto-math
pashto-math
Dataset Summary
pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي.
Dataset Structure
هره نمونه د ChatML-style SFT په بڼه ده:
{
"id": "000005",
"messages": [
{ "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.Pashto-Medical-o1-Reasoning-SFT-Dataset
Pashto Medical o1 Reasoning SFT Dataset
This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models.
Dataset Structure
The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response.
Data Fields
Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.pashto-sociology
# Dataset Card for Pashto Sociology Dataset
## Dataset Description
- **Homepage:** [N/A]
- **Repository:** [Nassimjp/pashto-sociology](https://huggingface.co/datasets/nassimjp/pashto-sociology)
- **Paper:** [N/A]
- **Leaderboard:** [N/A]
- **Point of Contact:** [N/A]
### Dataset Summary
This dataset contains a collection of 100 sociological dialogue samples in the Pashto language. It is designed to facilitate research and development of conversational AI, natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-sociology.perfect-pashto-reasoning-sft
Perfect Pashto Reasoning SFT Dataset
د پښتو ژبې لپاره تر ټولو پاک او لوړ کیفیت لرونکی Reasoning Dataset
📌 الوتنه (Overview)
دا ډېټاسیټ د Magpie-Pro-300K-Filtered dataset پر بنسټ جوړ شوی دی چې د Pashto LLM او AI ټولنې لپاره په بشپړ ډول نوي سره انجنیر شوی او پروسس شوی دی.
ټول ډاټا په اتومي او لاین په لاین ډول ژباړل شوې او په لوړ کیفیت سره reformatted شوې ترڅو د alignment-handbook سره مستقیم مطابقت ولري. هدف یې د پښتو ژبې نوي نسل ماډلونو (لکه Rawanاو Ghanam لړۍ)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/perfect-pashto-reasoning-sft.pashto-dingding-poet
Pashto DingDing Poet 🔔
A Pashto poetry error‑analysis and evaluation dataset
🧩 Overview
Pashto DingDing Poet is a curated collection of low‑quality, unnatural, machine‑translated, or semantically incoherent Pashto poetry outputs.It is designed specifically for LLM evaluation, error analysis, and quality‑control research in Pashto NLP.
This dataset helps models transition from:
Pashto-looking text → Natural, meaningful Pashto.
🎯 Purpose
This… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dingding-poet.multilingual-proverb-reasoning
📖 Multilingual Proverb Reasoning Dataset
This dataset contains 990 unique proverbs primarily in Japanese, with accompanying phonetic readings (Romaji), literal translations, and multilingual reasoning "Chain of Thought" (thought) fields. It is designed for training and evaluating LLMs on cultural nuance, metaphorical reasoning, and multilingual explanation tasks.
📊 Dataset Summary
The dataset was curated and cleaned to remove 182 duplicates, resulting in a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/multilingual-proverb-reasoning.pashto-fallacy-dataset
Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ)
The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts.
The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.test_bank_ALPACA_pashto
test_bank_ALPACA_pashto
Overview
test_bank_ALPACA_pashto د Pashto ژبې لپاره یو پاک، معیاري، ALPACA‑style dataset دی چې د instruction‑tuning موخو لپاره جوړ شوی. دا مجموعه د COLREGS Test Bank څخه Pashto ژباړل شوي او بیا ALPACA فورمټ ته restore شوي مثالونه لري.
Dataset درې برخې لري:
instruction — دنده یا پوښتنه
input — اضافي معلومات یا انتخابونه (که موجود وي)
output — د ماډل ځواب
دا جوړښت د Pashto LLMs لپاره د reasoning، rewriting، summarization، او structured… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/test_bank_ALPACA_pashto.
