CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dpevzner /Cybersecurity_Reasoning_Dataset Cybersecurity Reasoning Dataset (Model-Agnostic) A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template; this dataset losslessly separates reasoning content from format, providing one neutral canonical corpus plus four per-family rendered training variants (Mistral/Llama, DeepSeek, ChatML, Gemma). Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.texttext-generationn<1K1 likes1.3k downloads2mo agoHugging Face02nassimjp /Pashto-Free-Hand-Reasoning-Dataset Pashto Free-Hand Reasoning SFT Dataset 🧠♻️ This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses. 🔄 The 3R Approach (Recycle, Reuse, Reason) Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy: Recycle: Taking older, simple, or raw legacy Pashto questions. Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.texttext-generation1K<n<10K0 likes228 downloads10d agoHugging Face03valendra /sherry-reasoning-effort-0.1-dataset Sherry Reasoning Effort Dataset (0.1) Math reasoning traces at 3 effort levels (low, medium, high) generated by Qwen3.6-35B-A3B (NVFP4) via API. Description Each problem was sampled from NuminaMath-CoT (numeric-answer problems only, proofs filtered out). For every problem, 3 candidates were sampled from the generator model with thinking enabled, verified against the ground truth, and the 3 verified traces with different reasoning depths were kept: the shortest… See the full description on the dataset page: https://huggingface.co/datasets/valendra/sherry-reasoning-effort-0.1-dataset.texttext-generation1K<n<10K0 likes53 downloads19d agoHugging Face04nassimjp /Pashto-Social-Insight-Reasoning-Dataset Pashto Social Insight & Reasoning Dataset (PSIR) Overview The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning. Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.texttext-generation1K<n<10K0 likes51 downloads6d agoHugging Face05nassimjp /pashto-reasoning-children-story-crafting-dataset Pashto Reasoning Children Story Crafting Dataset Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps. Dataset Overview & Methodology Language: Pashto (ps) Base Prompts: 100 unique core story prompts. Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.texttext-generationn<1K0 likes48 downloads5d agoHugging Face06nassimjp /pashto-reasoning-chat-dataset Pashto Reasoning Chat Dataset A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto. 📊 Dataset Structure Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities: system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.texttext-generationn<1K0 likes43 downloads5d agoHugging Face07RefinedNeuro /Qwen3-Reasoning-Distill-Q-A-Dataset Qwen3 Reasoning Distill Q&A Dataset Repository: RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset Authors Mehmet Can Farsak Serhat Atayeter License This dataset is released under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication. Dataset Summary This dataset contains question-answer pairs across six STEM subjects designed for Turkish-language reasoning tasks. It was generated using the qwen3-32b model and is intended for fine-tuning the RN_TR_R2… See the full description on the dataset page: https://huggingface.co/datasets/RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset.tabularquestion-answering10K<n<100K2 likes40 downloads1y agoHugging Face08nassimjp /Pashto-Quran-Native-Reasoning-Dataset Pashto-Quran-Native-Reasoning-Dataset A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses. Overview Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto. The dataset is designed to help language models learn to: understand Quranic text and its Pashto meaning reason about the supplied content naturally distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.texttext-generationn<1K0 likes40 downloads5d agoHugging Face09saberai /ccf-reasoning-dataset Cognitive Cascade Framework (CCF) Reasoning Dataset A high-quality dataset of structured reasoning examples using the Cognitive Cascade Framework (CCF), designed for training language models to perform systematic, multi-stage reasoning. Dataset Description This dataset contains problems across multiple domains (math, science, coding, creative reasoning) paired with detailed reasoning chains following the CCF methodology. Each example includes a complete reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/saberai/ccf-reasoning-dataset.texttext-generation1K<n<10K0 likes37 downloads10mo agoHugging Face10nassimjp /Sindhi-Reasoning-Chat-Dataset # 📘 **Sindhi‑Reasoning‑Chat‑Dataset** A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages. --- ## 🧠 **Dataset Summary** Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.texttext-generationn<1K0 likes33 downloads3d agoHugging Face11mattwesney /Reasoning_Problem_Solving_Datasetgated Reasoning and Problem-Solving Dataset (RPSD) Overview The Reasoning and Problem-Solving Dataset (RPSD) is a comprehensive, high-quality set of synthetically generated question-answer pairs (150k+) tailored for training AI systems in logical reasoning and problem-solving. It spans multiple domains, including core reasoning techniques, specialized fields like science, mathematics, engineering, computer science, and philosophy, along with practical, real-world… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/Reasoning_Problem_Solving_Dataset.texttext-generation100K<n<1M15 likes32 downloads2mo agoHugging Face12likhitjuttada /finance-reasoning-sft-dataset Personal Finance Reasoning Dataset A synthetic instruction-tuning dataset designed to teach language models to reason through personal finance and investing decisions using the mental frameworks from classic books in the genre. The goal is not recall of book content but principled reasoning: the model should apply frameworks to novel situations it has never seen. Source Books Principles were extracted from the following books: The Psychology of Money — Morgan Housel Rich… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/finance-reasoning-sft-dataset.texttext-generationn<1K0 likes31 downloads5mo agoHugging Face13nassimjp /Pashto-Medical-o1-Reasoning-SFT-Dataset Pashto Medical o1 Reasoning SFT Dataset This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models. Dataset Structure The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response. Data Fields Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.texttext-generation10K<n<100K0 likes28 downloads20h agoHugging Face14aditya-datasets /marathi-land-law-reasoning Marathi Land Legal Reasoning (100 High-Quality Rows) 📌 Dataset Description This is a high-precision reasoning dataset focused on the Maharashtra Land Revenue Code (MLRC) and property succession laws in India. It is specifically designed for fine-tuning Large Language Models (LLMs) to handle complex, domain-specific logic in the Marathi language. Unlike generic datasets, this collection focuses on "Chain-of-Thought" (CoT), providing a detailed thought block for every… See the full description on the dataset page: https://huggingface.co/datasets/aditya-datasets/marathi-land-law-reasoning.texttext-generationn<1K0 likes26 downloads9mo agoHugging Face15Zyroxx66 /Somali-Reasoning-Dataset Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴 This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language. 🌟 What makes this unique? This is a Hybrid Dataset that combines two powerful sources: The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.texttext-generation10K<n<100K0 likes21 downloads6mo agoHugging Face16Agnuxo /scientific-reasoning-dataset Scientific Reasoning Dataset Synthetic reasoning traces based on Agnuxo research. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generationn<1K0 likes15 downloads5mo agoHugging Face17mattwesney /ToT_Reasoning_Problem_Solving_Dataset_V2gated ToT-RPSD-V2 This dataset consists of 70,000 high-quality, synthetically generated Q&A pairs with a strong emphasis on reasoning (inspired by o1 type reasoning) and the use of "Train of Thought" methodologies. Each entry is meticulously structured into six key components: the question, answer, reasoning (detailing the thought process leading to the answer), a unique ID, topic tags, and a difficulty level. While the dataset strongly focuses on science and cognitive tasks, it… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/ToT_Reasoning_Problem_Solving_Dataset_V2.texttext-generation10K<n<100K7 likes14 downloads2y agoHugging Face18dpevzner /Cybersecurity_Reasoning_Dataset_MistralFamily_7bgated Cybersecurity Reasoning Dataset (v6.0) A high-fidelity, forensics-mapped training corpus for cybersecurity reasoning and SOC automation — formatted for the Mistral / Llama instruct family. Format-specific dataset. Every record uses the Alpaca-style ### Instruction: / ### Response: template native to Mistral/Llama instruct models. A model-agnostic version of this corpus (Mistral, DeepSeek, ChatML, and Gemma variants rendered from one neutral canonical source) is published… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset_MistralFamily_7b.texttext-generationn<1K1 likes10 downloads2mo agoHugging Face19aboros98 /reasoning_datasettextquestion-answering1K<n<10K1 likes9 downloads2y agoHugging Face20Alindstroem89 /step_reasoning_dataset Step Reasoning Dataset Structured Reasoning Decomposition for Language Models Overview The Step Reasoning Dataset is a synthetic dataset designed for training language models to decompose complex questions into structured reasoning plans. Instead of directly generating final answers, the dataset focuses on: Breaking problems into reasoning steps Identifying dependencies between steps Generating structured JSON decompositions Distinguishing factual, logical… See the full description on the dataset page: https://huggingface.co/datasets/Alindstroem89/step_reasoning_dataset.textquestion-answeringn<1K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.