datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.ipashto-glossary-ddup
🦅 iPashto.ai Culturally Aligned Master Glossary (ddup)
Welcome to the official iPashto.ai Master Glossary dataset repo (nassimjp/ipashto-glossary-ddup). This is a highly specialized, clean, and deduplicated dictionary network covering 39 strategic domains with 11,442 unique entries.
As verified in our deployment log (Screenshot from 2026-07-02 22-33-31.png), this dataset maps complex domain terminology and foreign personal names (English, Chinese, and Japanese) directly into… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/ipashto-glossary-ddup.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.nasjonalt-vitenarkiv
Nasjonalt vitenarkiv
Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national
repository where Norwegian research institutions publish their output: master's and PhD theses,
journal articles, and technical and research reports. Subjects span the disciplines - marine
science, forestry, archaeology, education, public health, engineering - and most documents are
recent.
Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.RefusEU
Multilingual Red-Teaming Preference + Evaluation
This dataset contains multilingual red-teaming prompts and preference pairs for safety research, plus a separate multilingual evaluation prompt set.
Dataset Structure
The repository is organized into multiple dataset configs:
multilingual
train: full preference training split
test: full preference test split
lang_pl, lang_en, lang_cs, lang_sk, lang_sl, lang_lt, lang_lv, lang_de, lang_it, lang_fr, lang_es, lang_pt
train:… See the full description on the dataset page: https://huggingface.co/datasets/NASK-PIB/RefusEU.afghanistan-post-2021-pashto-dataset
Afghanistan Post-2021 Pashto Dataset
Dataset Description
This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for:
Training and fine-tuning Pashto large language models (LLMs)
Question-answering tasks
Research on Afghanistan's post-2021 developments
Low-resource language AI development
The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.bloomee-sft-nasasmd-grounded-5m
🌸 bloomee-sft-nasasmd-grounded-5m
Supervised fine-tuning corpus of grounded, tool-calling conversations about flowering phenology
Every answer traced back to the chunks it was drawn from, and scored against them
🏆 Part of the Bloomee platform — NASA Space Apps Challenge 2025
Built by Team Ganespace for Bandung, Indonesia
🌍 About
1,899 conversations that teach a small model to behave like the Bloomee
agent: call the right NDVI tool with the right… See the full description on the dataset page: https://huggingface.co/datasets/bloomee-app/bloomee-sft-nasasmd-grounded-5m.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.pashto-instruct-dataset
Pashto Instruct Dataset
This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions.
Dataset Structure
Each sample in the dataset contains the following fields:
id: Unique identifier for the sample.
messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.Da-Ploshi-OpenThoughts_Cache
Da-Ploshi OpenThoughts Cache
Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases.
The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs.
Dataset Structure
Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.pashto-algebra
د پښتو الجبرا پروژه 🤖🇦🇫📚🧠
د افغان نجونو لپاره ډالۍ 💝
"تعلیم یو حق دی، نه مرسته." ✨
هغو زړورو افغان نجونو ته چې له ښوونځي او کتابونو څخه محرومې دي — دا پروژه ستاسو لپاره ده! 🇦🇫❤️
📖 د پروژې په اړه
Pashto Algebra Dataset په پښتو ژبه کې لومړی او تر ټولو لوی ګام په ګام ریاضي ډیټاسیټ دی.
دا پروژه د هغو افغان ماشومانو لپاره جوړه شوې چې په ځانګړې توګه نجونې چې په افغانستان کې له ښوونځي تګ څخه منع دي او هلکان چې په لرو پرتو سیمو کې اوسي.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-algebra.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.pashto-reasoning-chat-dataset
Pashto Reasoning Chat Dataset
A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto.
📊 Dataset Structure
Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities:
system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.Pashto-Brain-Extraction-Dataset
🧠 Pashto Brain Extraction Dataset
A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer.
Keep the brain 🧠 — throw away the mouth 🗣️
🎯 Purpose
A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer.
Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.pashto-wikipedia-sft
Pashto Wikipedia SFT Dataset
This dataset is a cleaned, structured, and context-anchored variant of the Pashto Wikipedia corpus, formatted explicitly for Supervised Fine-Tuning (SFT) and Continual Pre-Training (CPT).
By restructuring raw encyclopedic data into explicit source-grounding prompts, this dataset trains Large Language Models (LLMs) to couple their factual generation directly with reference variables (URLs and Titles), mitigating hallucination tendencies in… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-wikipedia-sft.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.da-tolanpoheni-pokhtany
Da Tolanpoheni Pokhtany (د ټولنیزې پوښتنې)
Overview
da-tolanpoheni-pokhtany is a specialized dataset curated for Pashto language processing, natural language understanding, and related machine learning applications.
Structure
Data Fields
text / standard fields containing the primary dataset contents.
Data Instances
An example instance from the dataset:
{
"text": "Sample text entry..."
}
Usage
You can… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/da-tolanpoheni-pokhtany.Medic_Chat-Pashto
📦 Dataset Summary
ژبه: Pashto
ډول: Chat‑style SFT (Supervised Fine‑Tuning)
موضوع: Traditional Chinese Medicine (TCM)
ریکارډونه: شاوخوا 10.8k
فورمټ: JSONL — messages: [{role, content}, ...]
لایسنس: CC‑BY‑NC‑4.0
کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs
🧬 Data Structure
هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده:
{
"messages": [
{"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.Sindhi-Reasoning-Chat-Dataset
# 📘 **Sindhi‑Reasoning‑Chat‑Dataset**
A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages.
---
## 🧠 **Dataset Summary**
Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.questions
📝 Overview
Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی.
🎯 Purpose
دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی:
Pashto instruction-tuning
Pashto question-answering
Pashto reasoning
Pashto dialogue modeling
Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.pashto-legal-qa-chat
Pashto Legal QA Chat Dataset
⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.)
Dataset Overview
The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.pashto-math
pashto-math
Dataset Summary
pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي.
Dataset Structure
هره نمونه د ChatML-style SFT په بڼه ده:
{
"id": "000005",
"messages": [
{ "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.
