CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nassimjp /pashto-emoji-dataset Pashto Emoji Dataset This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text. The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label. Dataset Structure The dataset is provided in the following format: text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.texttext-to-image100K<n<1M0 likes494 downloads27d agoHugging Face02nasrellahkharroubi /DarijaDz DarijaDZ DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content. Dataset Description Motivation Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.texttext-generation10M<n<100M10 likes324 downloads11d agoHugging Face03nassimjp /Pashto-Free-Hand-Reasoning-Dataset Pashto Free-Hand Reasoning SFT Dataset 🧠♻️ This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses. 🔄 The 3R Approach (Recycle, Reuse, Reason) Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy: Recycle: Taking older, simple, or raw legacy Pashto questions. Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.texttext-generation1K<n<10K0 likes228 downloads9d agoHugging Face04nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face05nassimjp /afghanistan-post-2021-pashto-dataset Afghanistan Post-2021 Pashto Dataset Dataset Description This dataset contains 1,100+ high-quality Pashto-language questions covering Afghanistan's political, social, economic, and humanitarian situation after 2021. It is designed for: Training and fine-tuning Pashto large language models (LLMs) Question-answering tasks Research on Afghanistan's post-2021 developments Low-resource language AI development The questions are written in authentic, natural Pashto and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-dataset.textquestion-answering1K<n<10K0 likes67 downloads9d agoHugging Face06nassimjp /Pashto-OpenThoughts-15K-Reasoning Pashto-OpenThoughts-15K-Reasoning Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto. 📌 Dataset Description Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset. The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as: Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.texttext-generation1K<n<10K0 likes59 downloads2d agoHugging Face07nassimjp /pashto-instruct-dataset Pashto Instruct Dataset This is a curated instruction-tuning dataset for the Pashto language (ps), designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs). It contains multi-turn and single-turn conversational data, problem-solving prompts, and localized instructions. Dataset Structure Each sample in the dataset contains the following fields: id: Unique identifier for the sample. messages: A list of message objects representing the conversation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-instruct-dataset.texttext-generation1K<n<10K0 likes58 downloads26d agoHugging Face08nassimjp /Da-Ploshi-OpenThoughts_Cache Da-Ploshi OpenThoughts Cache Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases. The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs. Dataset Structure Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.texttext-generation100K<n<1M0 likes53 downloads2d agoHugging Face09Nasaq-GP /Quran_tafsir Nasq Quranic Dataset (Arabic-English Tafsir) لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي). Columns: id: المعرف الفريد لكل آية (من 1 إلى 6236). surah_n: رقم السورة. ayah_n: رقم الآية داخل السورة. surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة). surah_name_english: اسم السورة باللغة الإنجليزية. ayah_text_ar: نص الآية. ayah_text_en: ترجمة نص الآية للإنجليزية. arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.tabulartranslation1K<n<10K0 likes52 downloads8mo agoHugging Face10nassimjp /afghanistan-post-2021-pashto-conversation-3x 🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X nassimjp/afghanistan-post-2021-pashto-conversation-3x A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity. 📌 Overview This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset. Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.texttext-generationn<1K0 likes51 downloads9d agoHugging Face11nassimjp /Pashto-Social-Insight-Reasoning-Dataset Pashto Social Insight & Reasoning Dataset (PSIR) Overview The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning. Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.texttext-generation1K<n<10K0 likes51 downloads6d agoHugging Face12nassimjp /pashto-reasoning-children-story-crafting-dataset Pashto Reasoning Children Story Crafting Dataset Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps. Dataset Overview & Methodology Language: Pashto (ps) Base Prompts: 100 unique core story prompts. Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.texttext-generationn<1K0 likes48 downloads4d agoHugging Face13nassimjp /pashto-reasoning-chat-dataset Pashto Reasoning Chat Dataset A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto. 📊 Dataset Structure Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities: system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.texttext-generationn<1K0 likes43 downloads5d agoHugging Face14nassimjp /Pashto-Brain-Extraction-Dataset 🧠 Pashto Brain Extraction Dataset A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer. Keep the brain 🧠 — throw away the mouth 🗣️ 🎯 Purpose A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer. Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.texttext-generationn<1K0 likes42 downloads10d agoHugging Face15nassimjp /Pashto-Quran-Native-Reasoning-Dataset Pashto-Quran-Native-Reasoning-Dataset A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses. Overview Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto. The dataset is designed to help language models learn to: understand Quranic text and its Pashto meaning reason about the supplied content naturally distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.texttext-generationn<1K0 likes40 downloads4d agoHugging Face16nassimjp /pashto-wikipedia-sft Pashto Wikipedia SFT Dataset This dataset is a cleaned, structured, and context-anchored variant of the Pashto Wikipedia corpus, formatted explicitly for Supervised Fine-Tuning (SFT) and Continual Pre-Training (CPT). By restructuring raw encyclopedic data into explicit source-grounding prompts, this dataset trains Large Language Models (LLMs) to couple their factual generation directly with reference variables (URLs and Titles), mitigating hallucination tendencies in… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-wikipedia-sft.texttext-generation10K<n<100K1 likes39 downloads3mo agoHugging Face17nassimjp /Pashto-grammar-100 🇦🇫 Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.texttext-generationn<1K0 likes37 downloads10d agoHugging Face18nassimjp /da-tolanpoheni-pokhtany Da Tolanpoheni Pokhtany (د ټولنیزې پوښتنې) Overview da-tolanpoheni-pokhtany is a specialized dataset curated for Pashto language processing, natural language understanding, and related machine learning applications. Structure Data Fields text / standard fields containing the primary dataset contents. Data Instances An example instance from the dataset: { "text": "Sample text entry..." } Usage You can… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/da-tolanpoheni-pokhtany.texttext-generation1K<n<10K0 likes36 downloads6d agoHugging Face19nassimjp /Medic_Chat-Pashto 📦 Dataset Summary ژبه: Pashto ډول: Chat‑style SFT (Supervised Fine‑Tuning) موضوع: Traditional Chinese Medicine (TCM) ریکارډونه: شاوخوا 10.8k فورمټ: JSONL — messages: [{role, content}, ...] لایسنس: CC‑BY‑NC‑4.0 کارونې: Pashto medical assistants, TCM reasoning models, multilingual medical LLMs 🧬 Data Structure هره نمونه د user او assistant ترمنځ یوه طبي مکالمه ده: { "messages": [ {"role": "user", "content": "زه د معدې درد لرم، مهرباني وکړئ… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Medic_Chat-Pashto.texttext-generation10K<n<100K0 likes35 downloads3mo agoHugging Face20nassimjp /Sindhi-Reasoning-Chat-Dataset # 📘 **Sindhi‑Reasoning‑Chat‑Dataset** A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages. --- ## 🧠 **Dataset Summary** Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.texttext-generationn<1K0 likes33 downloads2d agoHugging Face21nassimjp /questions 📝 Overview Questions یو پاک، deduplicated، shuffled Pashto پوښتنو ډیټاسیټ دی چې د Pashto ژبې د پوښتنې–ځواب، reasoning، instruction-following، او general-purpose SFT لپاره کارول کېږي.دا ډیټاسیټ د مختلفو Pashto سرچینو څخه اخیستل شوی، پاک شوی، duplicate لرې شوي، او د ماډل د ښه عمومي کولو لپاره ګډوډ شوی دی. 🎯 Purpose دا ډیټاسیټ د لاندې کارونو لپاره جوړ شوی: Pashto instruction-tuning Pashto question-answering Pashto reasoning Pashto dialogue modeling Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/questions.texttext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face22nassimjp /pashto-legal-qa-chat Pashto Legal QA Chat Dataset ⚠️ محتاط (Caution): دا یو ماشین ژباړه ده انسانی سمون او بیا سفای ته اړتیا لری. (This is a machine translation and requires human editing and refinement.) Dataset Overview The Pashto Legal QA Chat Dataset is a conversational dataset structured specifically for fine-tuning Large Language Models (LLMs) on legal domains in the Pashto language. It adapts traditional legal question-answer pairs into a multi-turn chat format (messages… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-legal-qa-chat.texttext-generation10K<n<100K0 likes31 downloads2mo agoHugging Face23nassimjp /pashto-math pashto-math Dataset Summary pashto-math د Pashto ژبې لپاره یو پراخ، پاک، او ښوونیز ریاضي ډیټاسټ دی چې د کلمو مسئلې، محاسبې، منطقي استدلال، او ښوونیزو تمرینونو پراخ پوښښ لري. دا ډیټاسټ د Pashto LLMونو لپاره د reasoning وړتیا لوړولو هدف لري او د ښوونځي د ریاضي د کچې لپاره معیاري، منظم، او deterministic ځوابونه وړاندې کوي. Dataset Structure هره نمونه د ChatML-style SFT په بڼه ده: { "id": "000005", "messages": [ { "role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-math.texttext-generation1K<n<10K2 likes29 downloads2mo agoHugging Face24nassimjp /Pashto-Medical-o1-Reasoning-SFT-Dataset Pashto Medical o1 Reasoning SFT Dataset This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models. Dataset Structure The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response. Data Fields Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.texttext-generation10K<n<100K0 likes28 downloads8h agoHugging Face25nassimjp /pashto-sociology # Dataset Card for Pashto Sociology Dataset ## Dataset Description - **Homepage:** [N/A] - **Repository:** [Nassimjp/pashto-sociology](https://huggingface.co/datasets/nassimjp/pashto-sociology) - **Paper:** [N/A] - **Leaderboard:** [N/A] - **Point of Contact:** [N/A] ### Dataset Summary This dataset contains a collection of 100 sociological dialogue samples in the Pashto language. It is designed to facilitate research and development of conversational AI, natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-sociology.textquestion-answeringn<1K0 likes27 downloads1mo agoHugging Face26nassimjp /perfect-pashto-reasoning-sft Perfect Pashto Reasoning SFT Dataset د پښتو ژبې لپاره تر ټولو پاک او لوړ کیفیت لرونکی Reasoning Dataset 📌 الوتنه (Overview) دا ډېټاسیټ د Magpie-Pro-300K-Filtered dataset پر بنسټ جوړ شوی دی چې د Pashto LLM او AI ټولنې لپاره په بشپړ ډول نوي سره انجنیر شوی او پروسس شوی دی. ټول ډاټا په اتومي او لاین په لاین ډول ژباړل شوې او په لوړ کیفیت سره reformatted شوې ترڅو د alignment-handbook سره مستقیم مطابقت ولري. هدف یې د پښتو ژبې نوي نسل ماډلونو (لکه Rawanاو Ghanam لړۍ)… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/perfect-pashto-reasoning-sft.texttext-generation100K<n<1M0 likes26 downloads4mo agoHugging Face27nassimjp /pashto-dingding-poet Pashto DingDing Poet 🔔 A Pashto poetry error‑analysis and evaluation dataset 🧩 Overview Pashto DingDing Poet is a curated collection of low‑quality, unnatural, machine‑translated, or semantically incoherent Pashto poetry outputs.It is designed specifically for LLM evaluation, error analysis, and quality‑control research in Pashto NLP. This dataset helps models transition from: Pashto-looking text → Natural, meaningful Pashto. 🎯 Purpose This… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dingding-poet.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face28nassimjp /multilingual-proverb-reasoning 📖 Multilingual Proverb Reasoning Dataset This dataset contains 990 unique proverbs primarily in Japanese, with accompanying phonetic readings (Romaji), literal translations, and multilingual reasoning "Chain of Thought" (thought) fields. It is designed for training and evaluating LLMs on cultural nuance, metaphorical reasoning, and multilingual explanation tasks. 📊 Dataset Summary The dataset was curated and cleaned to remove 182 duplicates, resulting in a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/multilingual-proverb-reasoning.texttext-generationn<1K0 likes24 downloads4mo agoHugging Face29nassimjp /pashto-fallacy-dataset Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ) The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts. The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face30nassimjp /test_bank_ALPACA_pashto test_bank_ALPACA_pashto Overview test_bank_ALPACA_pashto د Pashto ژبې لپاره یو پاک، معیاري، ALPACA‑style dataset دی چې د instruction‑tuning موخو لپاره جوړ شوی. دا مجموعه د COLREGS Test Bank څخه Pashto ژباړل شوي او بیا ALPACA فورمټ ته restore شوي مثالونه لري. Dataset درې برخې لري: instruction — دنده یا پوښتنه input — اضافي معلومات یا انتخابونه (که موجود وي) output — د ماډل ځواب دا جوړښت د Pashto LLMs لپاره د reasoning، rewriting، summarization، او structured… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/test_bank_ALPACA_pashto.texttext-generation10K<n<100K0 likes24 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.