CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nassimjp /pashto-emoji-dataset Pashto Emoji Dataset This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text. The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label. Dataset Structure The dataset is provided in the following format: text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.texttext-to-image100K<n<1M0 likes494 downloads27d agoHugging Face02nasrellahkharroubi /DarijaDz DarijaDZ DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content. Dataset Description Motivation Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.texttext-generation10M<n<100M10 likes324 downloads11d agoHugging Face03nassimjp /Pashto-Free-Hand-Reasoning-Dataset Pashto Free-Hand Reasoning SFT Dataset 🧠♻️ This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses. 🔄 The 3R Approach (Recycle, Reuse, Reason) Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy: Recycle: Taking older, simple, or raw legacy Pashto questions. Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.texttext-generation1K<n<10K0 likes228 downloads9d agoHugging Face04nassimjp /jits-legal-dataset JITS Legal Dataset A production-ready, deterministic pipeline for processing Indian legal judgments into structured, high-quality legal datasets — with comprehensive extraction, self-citation exclusion, and multi-act statutory section detection. Overview Disclaimer: This dataset is independently created for research and engineering use. It is not an official government or judicial release and does not constitute legal advice. The JITS Legal Dataset currently… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/jits-legal-dataset.texttext-classification1K<n<10K0 likes160 downloads4mo agoHugging Face05nasrellahkharroubi /DarijaDZ-DialectID DarijaDZ Dialect Identification DarijaDZ-DialectID is a labeled dataset for classifying Algerian online text into one of six dialect/language classes: darija, msa, arabize, french, english, code_switch. It is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija. Dataset Description Motivation Algeria's online text is a mix of several dialects and scripts -- Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.texttext-classification10K<n<100K0 likes112 downloads15d agoHugging Face06nasib-ullah /xmc-lfamazontitles-131ktext100K<n<1M0 likes93 downloads2y agoHugging Face07HamiltonMYu /NASA-EO-Bench NASA-EO-Bench A large-scale benchmark for geoscience dataset retrieval, derived from citation relationships in peer-reviewed NASA publications. Paper: Bringing Agentic Search to Earth Observation Data Discovery — CIKM '26, 10.1145/3799682.3841109 Overview Finding the right NASA Earth observation dataset for a given research need is hard even for domain experts. NASA-EO-Bench operationalises this task as an information retrieval problem: given a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/HamiltonMYu/NASA-EO-Bench.tabulartext-retrieval10K<n<100K0 likes82 downloads1mo agoHugging Face08nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face09nassimjp /afghanistan-post-2021-pashto-dataset Afghanistan Post-2021 Pashto Dataset Dataset Description This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for: Training and fine-tuning Pashto large language models (LLMs) Question-answering tasks Research on Afghanistan's post-2021 developments Low-resource language AI development The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.textquestion-answering1K<n<10K0 likes67 downloads9d agoHugging Face10nasa-impact /nasa-sde-IR-benchmark-20251024-v5 NASA SDE IR Benchmark v5 A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation. Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST. Code: NASA-IMPACT/st-training-workflow Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.texttext-retrieval100K<n<1M1 likes65 downloads4mo agoHugging Face11nassimjp /pashto-instruct-training-datasettext1K<n<10K0 likes62 downloads21d agoHugging Face12nassimjp /Pashto-OpenThoughts-15K-Reasoning Pashto-OpenThoughts-15K-Reasoning Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto. 📌 Dataset Description Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset. The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as: Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.texttext-generation1K<n<10K0 likes59 downloads2d agoHugging Face13nassimjp /pashto-instruct-dataset Pashto Instruct Dataset This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions. Dataset Structure Each sample in the dataset contains the following fields: id: Unique identifier for the sample. messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.texttext-generation1K<n<10K0 likes58 downloads26d agoHugging Face14nassimjp /Driving-License-Pashto-QA 🚗 Driving License Pashto QA Dataset (د موټر چلولو جواز - پښتو ډاټاسیټ) This dataset contains translated Pashto Questions and Answers related to Driving License exams and road traffic rules. It was originally sourced/translated from Persian driving theory test questions and formatted for fine-tuning Large Language Models (LLMs) and training Chat completions models. دا ډاټاسیټ د موټر چلولو د لایسنس/جواز او ترافیکي مقرراتو پښتو پوښتنې او ځوابونه لري، چې له فارسي منبع څخه په معیاري… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Driving-License-Pashto-QA.textquestion-answering1K<n<10K0 likes57 downloads2mo agoHugging Face15nassimjp /History-of-India-in-Pashto History of India in Pashto 📜 A clean, structured, and high‑quality Pashto dataset containing historical questions covering the entire span of Indian history — from the Indus Valley Civilization to the Delhi Sultanate, Bhakti movements, Maratha Empire, colonial era, independence, and global revolutions. This dataset is designed for: Pashto Question‑Answering (QA) Pashto Reasoning Benchmarks Historical knowledge modeling Chat-style LLM training Cultural and academic research… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/History-of-India-in-Pashto.textn<1K0 likes54 downloads2mo agoHugging Face16nassimjp /Da-Ploshi-OpenThoughts_Cache Da-Ploshi OpenThoughts Cache Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases. The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs. Dataset Structure Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.texttext-generation100K<n<1M0 likes53 downloads2d agoHugging Face17nasa-impact /nasa-smd-qa-benchmark NASA-QA Benchmark NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.textquestion-answeringn<1K2 likes52 downloads2y agoHugging Face18Nasaq-GP /Quran_tafsir Nasq Quranic Dataset (Arabic-English Tafsir) لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي). Columns: id: المعرف الفريد لكل آية (من 1 إلى 6236). surah_n: رقم السورة. ayah_n: رقم الآية داخل السورة. surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة). surah_name_english: اسم السورة باللغة الإنجليزية. ayah_text_ar: نص الآية. ayah_text_en: ترجمة نص الآية للإنجليزية. arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.tabulartranslation1K<n<10K0 likes52 downloads8mo agoHugging Face19nassimjp /afghanistan-post-2021-pashto-conversation-3x 🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X nassimjp/afghanistan-post-2021-pashto-conversation-3x A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity. 📌 Overview This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset. Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.texttext-generationn<1K0 likes51 downloads9d agoHugging Face20nassimjp /Pashto-Social-Insight-Reasoning-Dataset Pashto Social Insight & Reasoning Dataset (PSIR) Overview The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning. Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.texttext-generation1K<n<10K0 likes51 downloads6d agoHugging Face21nassimjp /pashto-reasoning-children-story-crafting-dataset Pashto Reasoning Children Story Crafting Dataset Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps. Dataset Overview & Methodology Language: Pashto (ps) Base Prompts: 100 unique core story prompts. Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.texttext-generationn<1K0 likes48 downloads4d agoHugging Face22nassimjp /gpt-oss-pashto-conversation-3xtext1K<n<10K0 likes44 downloads8d agoHugging Face23nasa-impact /EO-via-NLP Dataset Summary Toward Open Earth Science as Fast and Accessible as Natural Language This dataset was curated to accompany the EO-via-NLP code and the following paper: Ellis, M., Gurung, I., Ramasubramanian, M., & Ramachandran, R. (2025).Toward Open Earth Science as Fast and Accessible as Natural Language.arXiv:2505.15690 Supported Tasks This dataset was primarily designed for: Named Entity Recognition (NER) in earth science contexts. Languages… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/EO-via-NLP.textn<1K0 likes43 downloads1y agoHugging Face24nassimjp /pashto-reasoning-chat-dataset Pashto Reasoning Chat Dataset A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto. 📊 Dataset Structure Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities: system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.texttext-generationn<1K0 likes43 downloads5d agoHugging Face25nassimjp /Pashto-Brain-Extraction-Dataset 🧠 Pashto Brain Extraction Dataset A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer. Keep the brain 🧠 — throw away the mouth 🗣️ 🎯 Purpose A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer. Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.texttext-generationn<1K0 likes42 downloads10d agoHugging Face26nassimjp /Pashto-Instruct Pashto-Instruct Pashto-Instruct is a curated, high-quality instruction-tuning dataset designed specifically for the Pashto language. This repository is part of the iPashto.ai initiative, which aims to bridge the resource gap for the Pashto language in modern Large Language Models (LLMs). Dataset Overview This dataset provides structured instruction-response pairs tailored for Supervised Fine-Tuning (SFT), alignment, and conversational capabilities in Pashto. It… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Instruct.text10K<n<100K0 likes41 downloads28d agoHugging Face27nassimjp /convergent-wisdom-pashto Dataset Card for Convergent Wisdom (Pashto) This dataset explores the convergence of wisdom traditions, bridging Eastern philosophies (including the Bhagavad Gita) and Western philosophical thought, specifically tailored for the Pashto language. Uses This dataset is designed for training and fine-tuning language models to understand and generate philosophical discourse in Pashto, fostering cross-cultural dialogue and semantic analysis of timeless wisdom.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/convergent-wisdom-pashto.textsentence-similarity10K<n<100K0 likes40 downloads7d agoHugging Face28nassimjp /Pashto-Quran-Native-Reasoning-Dataset Pashto-Quran-Native-Reasoning-Dataset A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses. Overview Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto. The dataset is designed to help language models learn to: understand Quranic text and its Pashto meaning reason about the supplied content naturally distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.texttext-generationn<1K0 likes40 downloads4d agoHugging Face29nassimjp /pashto-wikipedia-sft Pashto Wikipedia SFT Dataset This dataset is a cleaned, structured, and context-anchored variant of the Pashto Wikipedia corpus, formatted explicitly for Supervised Fine-Tuning (SFT) and Continual Pre-Training (CPT). By restructuring raw encyclopedic data into explicit source-grounding prompts, this dataset trains Large Language Models (LLMs) to couple their factual generation directly with reference variables (URLs and Titles), mitigating hallucination tendencies in… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-wikipedia-sft.texttext-generation10K<n<100K1 likes39 downloads3mo agoHugging Face30NasimBrz /SearchBench Dataset Card for SearchBench Dataset Summary SearchBench is a benchmark designed to evaluate Language Models' (LLMs) ability to solve state-based problems that require combinatorial search and backtracking. SearchBench problems require a systematic exploration of action paths and backtracking to feasible states, which poses a significant challenge for LLMs to solve end-to-end, due to their autoregressive next-token prediction architecture. The dataset is composed of five… See the full description on the dataset page: https://huggingface.co/datasets/NasimBrz/SearchBench.textquestion-answering1K<n<10K0 likes38 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.