datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.jits-legal-dataset
JITS Legal Dataset
A production-ready, deterministic pipeline for processing Indian legal judgments into structured, high-quality legal datasets — with comprehensive extraction, self-citation exclusion, and multi-act statutory section detection.
Overview
Disclaimer: This dataset is independently created for research and engineering use. It is not an official government or judicial release and does not constitute legal advice.
The JITS Legal Dataset currently… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/jits-legal-dataset.DarijaDZ-DialectID
DarijaDZ Dialect Identification
DarijaDZ-DialectID is a labeled dataset for classifying Algerian
online text into one of six dialect/language classes: darija, msa,
arabize, french, english, code_switch. It is part of
DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija.
Dataset Description
Motivation
Algeria's online text is a mix of several dialects and scripts --
Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.xmc-lfamazontitles-131kNASA-EO-Bench
NASA-EO-Bench
A large-scale benchmark for geoscience dataset retrieval, derived from citation relationships in peer-reviewed NASA publications.
Paper: Bringing Agentic Search to Earth Observation Data Discovery — CIKM '26, 10.1145/3799682.3841109
Overview
Finding the right NASA Earth observation dataset for a given research need is hard even for domain experts. NASA-EO-Bench operationalises this task as an information retrieval problem: given a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/HamiltonMYu/NASA-EO-Bench.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.afghanistan-post-2021-pashto-dataset
Afghanistan Post-2021 Pashto Dataset
Dataset Description
This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for:
Training and fine-tuning Pashto large language models (LLMs)
Question-answering tasks
Research on Afghanistan's post-2021 developments
Low-resource language AI development
The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.pashto-instruct-training-datasetPashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.pashto-instruct-dataset
Pashto Instruct Dataset
This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions.
Dataset Structure
Each sample in the dataset contains the following fields:
id: Unique identifier for the sample.
messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.Driving-License-Pashto-QA
🚗 Driving License Pashto QA Dataset (د موټر چلولو جواز - پښتو ډاټاسیټ)
This dataset contains translated Pashto Questions and Answers related to Driving License exams and road traffic rules. It was originally sourced/translated from Persian driving theory test questions and formatted for fine-tuning Large Language Models (LLMs) and training Chat completions models.
دا ډاټاسیټ د موټر چلولو د لایسنس/جواز او ترافیکي مقرراتو پښتو پوښتنې او ځوابونه لري، چې له فارسي منبع څخه په معیاري… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Driving-License-Pashto-QA.History-of-India-in-Pashto
History of India in Pashto 📜
A clean, structured, and high‑quality Pashto dataset containing historical questions covering the entire span of Indian history — from the Indus Valley Civilization to the Delhi Sultanate, Bhakti movements, Maratha Empire, colonial era, independence, and global revolutions.
This dataset is designed for:
Pashto Question‑Answering (QA)
Pashto Reasoning Benchmarks
Historical knowledge modeling
Chat-style LLM training
Cultural and academic research… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/History-of-India-in-Pashto.Da-Ploshi-OpenThoughts_Cache
Da-Ploshi OpenThoughts Cache
Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases.
The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs.
Dataset Structure
Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.nasa-smd-qa-benchmark
NASA-QA Benchmark
NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.gpt-oss-pashto-conversation-3xEO-via-NLP
Dataset Summary
Toward Open Earth Science as Fast and Accessible as Natural Language
This dataset was curated to accompany the EO-via-NLP code and the following paper:
Ellis, M., Gurung, I., Ramasubramanian, M., & Ramachandran, R. (2025).Toward Open Earth Science as Fast and Accessible as Natural Language.arXiv:2505.15690
Supported Tasks
This dataset was primarily designed for:
Named Entity Recognition (NER) in earth science contexts.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/EO-via-NLP.pashto-reasoning-chat-dataset
Pashto Reasoning Chat Dataset
A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto.
📊 Dataset Structure
Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities:
system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.Pashto-Brain-Extraction-Dataset
🧠 Pashto Brain Extraction Dataset
A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer.
Keep the brain 🧠 — throw away the mouth 🗣️
🎯 Purpose
A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer.
Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.Pashto-Instruct
Pashto-Instruct
Pashto-Instruct is a curated, high-quality instruction-tuning dataset designed specifically for the Pashto language. This repository is part of the iPashto.ai initiative, which aims to bridge the resource gap for the Pashto language in modern Large Language Models (LLMs).
Dataset Overview
This dataset provides structured instruction-response pairs tailored for Supervised Fine-Tuning (SFT), alignment, and conversational capabilities in Pashto. It… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Instruct.convergent-wisdom-pashto
Dataset Card for Convergent Wisdom (Pashto)
This dataset explores the convergence of wisdom traditions, bridging Eastern philosophies (including the Bhagavad Gita) and Western philosophical thought, specifically tailored for the Pashto language.
Uses
This dataset is designed for training and fine-tuning language models to understand and generate philosophical discourse in Pashto, fostering cross-cultural dialogue and semantic analysis of timeless wisdom.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/convergent-wisdom-pashto.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.pashto-wikipedia-sft
Pashto Wikipedia SFT Dataset
This dataset is a cleaned, structured, and context-anchored variant of the Pashto Wikipedia corpus, formatted explicitly for Supervised Fine-Tuning (SFT) and Continual Pre-Training (CPT).
By restructuring raw encyclopedic data into explicit source-grounding prompts, this dataset trains Large Language Models (LLMs) to couple their factual generation directly with reference variables (URLs and Titles), mitigating hallucination tendencies in… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-wikipedia-sft.SearchBench
Dataset Card for SearchBench
Dataset Summary
SearchBench is a benchmark designed to evaluate Language Models' (LLMs) ability to solve state-based problems that require combinatorial search and backtracking. SearchBench problems require a systematic exploration of action paths and backtracking to feasible states, which poses a significant challenge for LLMs to solve end-to-end, due to their autoregressive next-token prediction architecture.
The dataset is composed of five… See the full description on the dataset page: https://huggingface.co/datasets/NasimBrz/SearchBench.
