CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes115 downloads8mo agoHugging Face02nassimjp /quran-tafsir QuranLab — Multilingual Quran Tafsir Dataset A ready-to-use collection of Quran commentaries and annotated translations, aligned to the canonical 6,236 ayahs. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were written — not to claim them as ours. The text here reaches you through the work of QuranEnc.com, Tafsir Center for Quranic Studies, Quranic Universal Library… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran-tafsir.tabulartext-generation100K<n<1M0 likes103 downloads6d agoHugging Face03nassimjp /quran QuranLab — Verse-Aligned Multilingual Quran Corpus A unified, verse-aligned multilingual Quran corpus spanning 79 languages and 185 translations. Every recension and translation is a separate config (subset), all row-aligned on the canonical 6,236-ayah verse_key (Hafs ʿan ʿAsim reading, 114 surahs). The corpus also contains 111 tafsir configs: verse-grain classical and openly licensed Arabic works, plus the native-passage and verse-expanded views of Diyanet's Turkish Kur'an Yolu… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/quran.tabulartext-generation1M<n<10M0 likes99 downloads6d agoHugging Face04nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face05nassimjp /afghanistan-post-2021-pashto-dataset Afghanistan Post-2021 Pashto Dataset Dataset Description This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for: Training and fine-tuning Pashto large language models (LLMs) Question-answering tasks Research on Afghanistan's post-2021 developments Low-resource language AI development The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.textquestion-answering1K<n<10K0 likes67 downloads9d agoHugging Face06nasa-impact /nasa-sde-IR-benchmark-20251024-v5 NASA SDE IR Benchmark v5 A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation. Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST. Code: NASA-IMPACT/st-training-workflow Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.texttext-retrieval100K<n<1M1 likes65 downloads4mo agoHugging Face07nassimjp /Pashto-OpenThoughts-15K-Reasoning Pashto-OpenThoughts-15K-Reasoning Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto. 📌 Dataset Description Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset. The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as: Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.texttext-generation1K<n<10K0 likes59 downloads2d agoHugging Face08nassimjp /pashto-instruct-dataset Pashto Instruct Dataset This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions. Dataset Structure Each sample in the dataset contains the following fields: id: Unique identifier for the sample. messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.texttext-generation1K<n<10K0 likes58 downloads26d agoHugging Face09nassimjp /Driving-License-Pashto-QA 🚗 Driving License Pashto QA Dataset (د موټر چلولو جواز - پښتو ډاټاسیټ) This dataset contains translated Pashto Questions and Answers related to Driving License exams and road traffic rules. It was originally sourced/translated from Persian driving theory test questions and formatted for fine-tuning Large Language Models (LLMs) and training Chat completions models. دا ډاټاسیټ د موټر چلولو د لایسنس/جواز او ترافیکي مقرراتو پښتو پوښتنې او ځوابونه لري، چې له فارسي منبع څخه په معیاري… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Driving-License-Pashto-QA.textquestion-answering1K<n<10K0 likes57 downloads2mo agoHugging Face10nasa-impact /nasa-smd-qa-benchmark NASA-QA Benchmark NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.textquestion-answeringn<1K2 likes52 downloads2y agoHugging Face11nassimjp /pashto-algebra د پښتو الجبرا پروژه 🤖🇦🇫📚🧠 د افغان نجونو لپاره ډالۍ 💝 "تعلیم یو حق دی، نه مرسته." ✨ هغو زړورو افغان نجونو ته چې له ښوونځي او کتابونو څخه محرومې دي — دا پروژه ستاسو لپاره ده! 🇦🇫❤️ 📖 د پروژې په اړه Pashto Algebra Dataset په پښتو ژبه کې لومړی او تر ټولو لوی ګام په ګام ریاضي ډیټاسیټ دی. دا پروژه د هغو افغان ماشومانو لپاره جوړه شوې چې په ځانګړې توګه نجونې چې په افغانستان کې له ښوونځي تګ څخه منع دي او هلکان چې په لرو پرتو سیمو کې اوسي. 🎯… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-algebra.text-generation1K<n<10K0 likes51 downloads4mo agoHugging Face12nassimjp /afghanistan-post-2021-pashto-conversation-3x 🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X nassimjp/afghanistan-post-2021-pashto-conversation-3x A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity. 📌 Overview This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset. Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.texttext-generationn<1K0 likes51 downloads9d agoHugging Face13nassimjp /Pashto-Social-Insight-Reasoning-Dataset Pashto Social Insight & Reasoning Dataset (PSIR) Overview The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning. Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.texttext-generation1K<n<10K0 likes51 downloads6d agoHugging Face14NASP /neteval-examNetEval is a NetOps evaluation suite for foundation models, consisting of 5269 multi-choice questions. Please check our paper for more details about NetEval. We hope NetEval could help developers track the progress and analyze the NetOps ability of their models. Citation Please cite our paper if you use our dataset. @misc{miao2023empirical, title={An Empirical Study of NetOps Capability of Pre-Trained Large Language Models}, author={Yukai Miao and Yu Bai and Li Chen and… See the full description on the dataset page: https://huggingface.co/datasets/NASP/neteval-exam.text-classification10K<n<100K6 likes50 downloads3y agoHugging Face15NasimBrz /SearchBench Dataset Card for SearchBench Dataset Summary SearchBench is a benchmark designed to evaluate Language Models' (LLMs) ability to solve state-based problems that require combinatorial search and backtracking. SearchBench problems require a systematic exploration of action paths and backtracking to feasible states, which poses a significant challenge for LLMs to solve end-to-end, due to their autoregressive next-token prediction architecture. The dataset is composed of five… See the full description on the dataset page: https://huggingface.co/datasets/NasimBrz/SearchBench.textquestion-answering1K<n<10K0 likes38 downloads2y agoHugging Face16nassimjp /Pashto-grammar-100 🇦🇫 Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.texttext-generationn<1K0 likes37 downloads10d agoHugging Face17nassimjp /Medic_Chat-Pashto 📦 Dataset Summary ژبه: Pashto ډول: Chat‑style SFT (Supervised Fine‑Tuning) موضوع: Traditional Chinese Medicine (TCM) ریکارډونه: شاوخوا 10.8k فورمټ: JSONL — messages: [{role, content}, ...] لایسنس: CC‑BY‑NC‑4.0 کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs 🧬 Data Structure هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده: { "messages": [ {"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.texttext-generation10K<n<100K0 likes35 downloads3mo agoHugging Face18nassimjp /questions 📝 Overview Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی. 🎯 Purpose دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی: Pashto instruction-tuning Pashto question-answering Pashto reasoning Pashto dialogue modeling Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.texttext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face19nassimjp /pashto-legal-qa-chat Pashto Legal QA Chat Dataset ⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.) Dataset Overview The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.texttext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face20nassimjp /pashto-math pashto-math Dataset Summary pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي. Dataset Structure هره نمونه د ChatML-style SFT په بڼه ده: { "id": "000005", "messages": [ { "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.texttext-generation1K<n<10K2 likes29 downloads2mo agoHugging Face21nassimjp /Pashto-Medical-o1-Reasoning-SFT-Dataset Pashto Medical o1 Reasoning SFT Dataset This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models. Dataset Structure The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response. Data Fields Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.texttext-generation10K<n<100K0 likes28 downloads11h agoHugging Face22nassimjp /pashto-sociology # Dataset Card for Pashto Sociology Dataset ## Dataset Description - **Homepage:** [N/A] - **Repository:** [Nassimjp/pashto-sociology](https://huggingface.co/datasets/nassimjp/pashto-sociology) - **Paper:** [N/A] - **Leaderboard:** [N/A] - **Point of Contact:** [N/A] ### Dataset Summary This dataset contains a collection of 100 sociological dialogue samples in the Pashto language. It is designed to facilitate research and development of conversational AI, natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-sociology.textquestion-answeringn<1K0 likes27 downloads1mo agoHugging Face23nassimjp /perfect-pashto-reasoning-sft Perfect Pashto Reasoning SFT Dataset د پښتو ژبې لپاره تر ټولو پاک او لوړ کیفیت لرونکی Reasoning Dataset 📌 الوتنه (Overview) دا ډېټاسیټ د Magpie-Pro-300K-Filtered dataset پر بنسټ جوړ شوی دی چې د Pashto LLM او AI ټولنې لپاره په بشپړ ډول نوي سره انجنیر شوی او پروسس شوی دی. ټول ډاټا په اتومي او لاین په لاین ډول ژباړل شوې او په لوړ کیفیت سره reformatted شوې ترڅو د alignment-handbook سره مستقیم مطابقت ولري. هدف یې د پښتو ژبې نوي نسل ماډلونو (لکه Rawanاو Ghanam لړۍ)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/perfect-pashto-reasoning-sft.texttext-generation100K<n<1M0 likes26 downloads4mo agoHugging Face24nassimjp /pashto-quotes-dataset Pashto Quotes Dataset with Chain-of-Thought Reasoning Dataset Description This dataset contains 990 Persian (Farsi/Dari) quotes from various philosophers, writers, and thinkers, each accompanied by: A Pashto translation of the quote 5 step-by-step reasoning steps (Chain-of-Thought) in Pashto explaining the quote's meaning A concise conclusion in Pashto summarizing the key insight Each entry is designed to help language models learn reasoning, translation, and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-quotes-dataset.texttext-generationn<1K0 likes25 downloads4mo agoHugging Face25nassimjp /pashto-fallacy-dataset Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ) The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts. The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face26nassimjp /pashto-stf-grammar-pairs Pashto SFT Grammar Pairs Dataset Description Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes. The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.texttext-generationn<1K0 likes24 downloads1mo agoHugging Face27nassimjp /tolanpohena 📘 Tolanpohena — Pashto Social Sciences (ټولنپوهنه) Dataset A curated Pashto dataset focused on social sciences, sociology, community studies, and human behavior. Designed for Pashto LLM training, SFT, and educational applications. 📑 Overview Tolanpohena is a high‑quality Pashto dataset containing questions, explanations, definitions, and conceptual discussions related to: ټولنه (Society) ټولنیز جوړښت (Social Structure) کلتور (Culture) ارزښتونه (Values)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/tolanpohena.texttext-generationn<1K0 likes24 downloads1mo agoHugging Face28nassimjp /Pashto-Reasoning-RescueBench Pashto-Reasoning-RescueBench A High-Quality Pashto Chain-of-Thought Dataset for Rescue, Survival & Emergency Preparedness Dataset Description Pashto-Reasoning-RescueBench is a Pashto-language dataset created for training language models with strong reasoning in survival, bushcraft, and emergency situations. Origin & Creation Process Base Dataset: Derived from mattwesney/CoT_Reasoning_Bushcraft_Survival The original questions were translated/adapted into… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Reasoning-RescueBench.texttext-generation1K<n<10K2 likes23 downloads4mo agoHugging Face29nassimjp /Pashto-Clean-100k-Pairs.QA Pashto‑Clean‑100k‑Pairs.QA A curated collection of 100,000 Pashto question–answer pairs, cleaned and normalized for general‑purpose Pashto NLP training.This dataset focuses on broad coverage, topic diversity, and clean formatting, without synthetic reasoning or long‑context generation. Dataset Summary Pashto-Clean-100k-Pairs.QA contains short, direct QA pairs across 70+ everyday topics: Daily life Community Education Work Nature Safety Culture… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Clean-100k-Pairs.QA.texttext-generation100K<n<1M0 likes23 downloads1mo agoHugging Face30nassimjp /Pashto-Ethical-Bench_Base Pashto Ethical Benchmark Base (Pashto-Ethical-Bench_Base) Overview Pashto-Ethical-Bench_Base is a high-quality, carefully curated dataset containing 3,604 instruction-response pairs in Pashto (پښتو). The dataset focuses on criminal law, evidence rules, investigation procedures, forensic science, presumption of innocence, and ethical/legal reasoning. It is designed to: Improve safety and alignment of Pashto-language LLMs Evaluate cultural and legal understanding Support… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Ethical-Bench_Base.texttext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.