datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.quran-tafsir
QuranLab — Multilingual Quran Tafsir Dataset
A ready-to-use collection of Quran commentaries and annotated translations,
aligned to the canonical 6,236 ayahs.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours.
The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.quran
QuranLab — Verse-Aligned Multilingual Quran Corpus
A unified, verse-aligned multilingual Quran corpus spanning 79
languages and 185 translations. Every recension and translation is a
separate config (subset), all row-aligned on the canonical 6,236-ayah
verse_key (Hafs ʿan ʿAsim reading, 114 surahs).
The corpus also contains 111 tafsir configs: verse-grain classical
and openly licensed Arabic works, plus the native-passage and verse-expanded
views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.afghanistan-post-2021-pashto-dataset
Afghanistan Post-2021 Pashto Dataset
Dataset Description
This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for:
Training and fine-tuning Pashto large language models (LLMs)
Question-answering tasks
Research on Afghanistan's post-2021 developments
Low-resource language AI development
The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.pashto-instruct-dataset
Pashto Instruct Dataset
This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions.
Dataset Structure
Each sample in the dataset contains the following fields:
id: Unique identifier for the sample.
messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.Driving-License-Pashto-QA
🚗 Driving License Pashto QA Dataset (د موټر چلولو جواز - پښتو ډاټاسیټ)
This dataset contains translated Pashto Questions and Answers related to Driving License exams and road traffic rules. It was originally sourced/translated from Persian driving theory test questions and formatted for fine-tuning Large Language Models (LLMs) and training Chat completions models.
دا ډاټاسیټ د موټر چلولو د لایسنس/جواز او ترافیکي مقرراتو پښتو پوښتنې او ځوابونه لري، چې له فارسي منبع څخه په معیاري… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Driving-License-Pashto-QA.nasa-smd-qa-benchmark
NASA-QA Benchmark
NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.pashto-algebra
د پښتو الجبرا پروژه 🤖🇦🇫📚🧠
د افغان نجونو لپاره ډالۍ 💝
"تعلیم یو حق دی، نه مرسته." ✨
هغو زړورو افغان نجونو ته چې له ښوونځي او کتابونو څخه محرومې دي — دا پروژه ستاسو لپاره ده! 🇦🇫❤️
📖 د پروژې په اړه
Pashto Algebra Dataset په پښتو ژبه کې لومړی او تر ټولو لوی ګام په ګام ریاضي ډیټاسیټ دی.
دا پروژه د هغو افغان ماشومانو لپاره جوړه شوې چې په ځانګړې توګه نجونې چې په افغانستان کې له ښوونځي تګ څخه منع دي او هلکان چې په لرو پرتو سیمو کې اوسي.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-algebra.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.neteval-examNetEval is a NetOps evaluation suite for foundation models, consisting of 5269 multi-choice questions. Please check our paper for more details about NetEval.
We hope NetEval could help developers track the progress and analyze the NetOps ability of their models.
Citation
Please cite our paper if you use our dataset.
@misc{miao2023empirical,
title={An Empirical Study of NetOps Capability of Pre-Trained Large Language Models},
author={Yukai Miao and Yu Bai and Li Chen and… See the full description on the dataset page: https://huggingface.co/datasets/NASP/neteval-exam.SearchBench
Dataset Card for SearchBench
Dataset Summary
SearchBench is a benchmark designed to evaluate Language Models' (LLMs) ability to solve state-based problems that require combinatorial search and backtracking. SearchBench problems require a systematic exploration of action paths and backtracking to feasible states, which poses a significant challenge for LLMs to solve end-to-end, due to their autoregressive next-token prediction architecture.
The dataset is composed of five… See the full description on the dataset page: https://huggingface.co/datasets/NasimBrz/SearchBench.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.Medic_Chat-Pashto
📦 Dataset Summary
ژبه: Pashto
ډول: Chat‑style SFT (Supervised Fine‑Tuning)
موضوع: Traditional Chinese Medicine (TCM)
ریکارډونه: شاوخوا 10.8k
فورمټ: JSONL — messages: [{role, content}, ...]
لایسنس: CC‑BY‑NC‑4.0
کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs
🧬 Data Structure
هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده:
{
"messages": [
{"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.questions
📝 Overview
Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی.
🎯 Purpose
دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی:
Pashto instruction-tuning
Pashto question-answering
Pashto reasoning
Pashto dialogue modeling
Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.pashto-legal-qa-chat
Pashto Legal QA Chat Dataset
⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.)
Dataset Overview
The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.pashto-math
pashto-math
Dataset Summary
pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي.
Dataset Structure
هره نمونه د ChatML-style SFT په بڼه ده:
{
"id": "000005",
"messages": [
{ "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.Pashto-Medical-o1-Reasoning-SFT-Dataset
Pashto Medical o1 Reasoning SFT Dataset
This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models.
Dataset Structure
The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response.
Data Fields
Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.pashto-sociology
# Dataset Card for Pashto Sociology Dataset
## Dataset Description
- **Homepage:** [N/A]
- **Repository:** [Nassimjp/pashto-sociology](https://huggingface.co/datasets/nassimjp/pashto-sociology)
- **Paper:** [N/A]
- **Leaderboard:** [N/A]
- **Point of Contact:** [N/A]
### Dataset Summary
This dataset contains a collection of 100 sociological dialogue samples in the Pashto language. It is designed to facilitate research and development of conversational AI, natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-sociology.perfect-pashto-reasoning-sft
Perfect Pashto Reasoning SFT Dataset
د پښتو ژبې لپاره تر ټولو پاک او لوړ کیفیت لرونکی Reasoning Dataset
📌 الوتنه (Overview)
دا ډېټاسیټ د Magpie-Pro-300K-Filtered dataset پر بنسټ جوړ شوی دی چې د Pashto LLM او AI ټولنې لپاره په بشپړ ډول نوي سره انجنیر شوی او پروسس شوی دی.
ټول ډاټا په اتومي او لاین په لاین ډول ژباړل شوې او په لوړ کیفیت سره reformatted شوې ترڅو د alignment-handbook سره مستقیم مطابقت ولري. هدف یې د پښتو ژبې نوي نسل ماډلونو (لکه Rawanاو Ghanam لړۍ)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/perfect-pashto-reasoning-sft.pashto-quotes-dataset
Pashto Quotes Dataset with Chain-of-Thought Reasoning
Dataset Description
This dataset contains 990 Persian (Farsi/Dari) quotes from various philosophers, writers, and thinkers, each accompanied by:
A Pashto translation of the quote
5 step-by-step reasoning steps (Chain-of-Thought) in Pashto explaining the quote's meaning
A concise conclusion in Pashto summarizing the key insight
Each entry is designed to help language models learn reasoning, translation, and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-quotes-dataset.pashto-fallacy-dataset
Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ)
The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts.
The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.tolanpohena
📘 Tolanpohena — Pashto Social Sciences (ټولنپوهنه) Dataset
A curated Pashto dataset focused on social sciences, sociology, community studies, and human behavior. Designed for Pashto LLM training, SFT, and educational applications.
📑 Overview
Tolanpohena is a high‑quality Pashto dataset containing questions, explanations, definitions, and conceptual discussions related to:
ټولنه (Society)
ټولنیز جوړښت (Social Structure)
کلتور (Culture)
ارزښتونه (Values)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/tolanpohena.Pashto-Reasoning-RescueBench
Pashto-Reasoning-RescueBench
A High-Quality Pashto Chain-of-Thought Dataset for Rescue, Survival & Emergency Preparedness
Dataset Description
Pashto-Reasoning-RescueBench is a Pashto-language dataset created for training language models with strong reasoning in survival, bushcraft, and emergency situations.
Origin & Creation Process
Base Dataset: Derived from mattwesney/CoT_Reasoning_Bushcraft_Survival
The original questions were translated/adapted into… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Reasoning-RescueBench.Pashto-Clean-100k-Pairs.QA
Pashto‑Clean‑100k‑Pairs.QA
A curated collection of 100,000 Pashto question–answer pairs, cleaned and normalized for general‑purpose Pashto NLP training.This dataset focuses on broad coverage, topic diversity, and clean formatting, without synthetic reasoning or long‑context generation.
Dataset Summary
Pashto-Clean-100k-Pairs.QA contains short, direct QA pairs across 70+ everyday topics:
Daily life
Community
Education
Work
Nature
Safety
Culture… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Clean-100k-Pairs.QA.Pashto-Ethical-Bench_Base
Pashto Ethical Benchmark Base (Pashto-Ethical-Bench_Base)
Overview
Pashto-Ethical-Bench_Base is a high-quality, carefully curated dataset containing 3,604 instruction-response pairs in Pashto (پښتو).
The dataset focuses on criminal law, evidence rules, investigation procedures, forensic science, presumption of innocence, and ethical/legal reasoning. It is designed to:
Improve safety and alignment of Pashto-language LLMs
Evaluate cultural and legal understanding
Support… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Ethical-Bench_Base.
