datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cybersecurity_Reasoning_Dataset
Cybersecurity Reasoning Dataset (Model-Agnostic)
A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original
corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template;
this dataset losslessly separates reasoning content from format, providing one
neutral canonical corpus plus four per-family rendered training variants
(Mistral/Llama, DeepSeek, ChatML, Gemma).
Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.unified-reasoning-dataset
Unified Reasoning Dataset
A 94,860-row English SFT collection that normalizes four synthetic reasoning and instruction datasets into one consistent schema.
Quick start
from datasets import load_dataset
dataset = load_dataset(
"j0no12/unified-reasoning-dataset",
split="train",
)
print(dataset.column_names)
# ['thinking', 'instruction', 'response', 'source']
print(dataset[0])
Dataset summary
Property
Value
Split
train only
Rows… See the full description on the dataset page: https://huggingface.co/datasets/j0no12/unified-reasoning-dataset.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.physics-reasoning-dataset
📚 Flux Physics Reasoning Dataset
This dataset contains detailed physics reasoning scenarios designed to train Small Language Models (SLMs) and Liquid Neural Networks in physical intuition.
📄 Format
The dataset is provided in Parquet format (train.parquet) for efficient loading. Each row contains:
prompt: The physics question or scenario description.
answer: The correct physical explanation and answer.
concept: The underlying physics principle (e.g., "Conservation of… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/physics-reasoning-dataset.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.twi-llm-reasoning-dataset-1k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Reasoning Dataset
A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.logic-problems-reasoning-dataset
Dataset Card for my-distiset-a26cd729
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-a26cd729/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/logic-problems-reasoning-dataset.reasoning-gym-verl-datasets
reasoning-gym-verl-datasets
This dataset contains procedurally generated reasoning tasks from the Reasoning Gym (r-gym) framework, structured and pre-processed in parquet format for training models with veRL.
These datasets were used to train MauroPello/Qwen3-1.7B-RL-final using GRPO (Group Relative Policy Optimization).
Dataset Splits & Structure
Split Name
Path
Size (Examples)
Description
train
train.parquet
100,000
Raw training set containing… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/reasoning-gym-verl-datasets.reasoning-dataset
ArgChains Reasoning Dataset
Dataset Description
This dataset contains narrative reasoning examples used for evaluation in the ArgChains hybrid neuro-symbolic reasoning framework. The dataset consists of fictional narratives involving ethical dilemmas.
The gold-standard chains are present in gold_standard_chains.txt for each story. They were generated using expert annotations.
The dataset is publicly available for research and non-commercial use.
sherry-reasoning-effort-0.1-dataset
Sherry Reasoning Effort Dataset (0.1)
Math reasoning traces at 3 effort levels (low, medium, high) generated by Qwen3.6-35B-A3B (NVFP4) via API.
Description
Each problem was sampled from NuminaMath-CoT (numeric-answer problems only, proofs filtered out).
For every problem, 3 candidates were sampled from the generator model with thinking enabled,
verified against the ground truth, and the 3 verified traces with different reasoning depths were kept:
the shortest… See the full description on the dataset page: https://huggingface.co/datasets/valendra/sherry-reasoning-effort-0.1-dataset.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.python-reasoning-dataset
Dataset Card for my-distiset-986461
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-986461/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/python-reasoning-dataset.pashto-reasoning-children-story-crafting-dataset
Pashto Reasoning Children Story Crafting Dataset
Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps.
Dataset Overview & Methodology
Language: Pashto (ps)
Base Prompts: 100 unique core story prompts.
Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.pashto-reasoning-chat-dataset
Pashto Reasoning Chat Dataset
A specialized chain-of-thought and multi-turn instruction dataset designed for training culturally grounded, sociologically aware, and reasoning-capable conversational AI agents in Pashto.
📊 Dataset Structure
Each sample in the dataset follows a structured conversational and reasoning format to support advanced alignment and chain-of-thought capabilities:
system: Fixed persona instructions (e.g., sociological context, cultural… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-chat-dataset.Qwen3-Reasoning-Distill-Q-A-Dataset
Qwen3 Reasoning Distill Q&A Dataset
Repository: RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset
Authors
Mehmet Can Farsak
Serhat Atayeter
License
This dataset is released under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication.
Dataset Summary
This dataset contains question-answer pairs across six STEM subjects designed for Turkish-language reasoning tasks. It was generated using the qwen3-32b model and is intended for fine-tuning the RN_TR_R2… See the full description on the dataset page: https://huggingface.co/datasets/RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset.high-reasoning-dataset-v1
High-Reasoning Dataset v1
2,139 premium Q&A pairs autonomously generated by a multi-AI knowledge distillation system. Each answer includes deep reasoning (CoT), mathematical foundations, production-ready code, historical context, edge cases, and real-world production incidents.
Made by Aditya Wakharkar | Tantra AI Labs
What is this dataset?
This is a synthetic training dataset created entirely by two AI models talking to each other 24/7 — no humans in the loop.… See the full description on the dataset page: https://huggingface.co/datasets/tantra-ai-labs/high-reasoning-dataset-v1.Pashto-Quran-Native-Reasoning-Dataset
Pashto-Quran-Native-Reasoning-Dataset
A specialized Pashto dataset designed for Quranic understanding, native reasoning, and natural conversational responses.
Overview
Pashto-Quran-Native-Reasoning-Dataset contains Quran-focused conversational training examples in Pashto.
The dataset is designed to help language models learn to:
understand Quranic text and its Pashto meaning
reason about the supplied content naturally
distinguish between text, translation… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Quran-Native-Reasoning-Dataset.reasoning-dataset
Reasoning & Thinking Dataset (RL/SFT Combined)
Overview
This dataset is a compiled collection of various reasoning, math, coding, and creative writing datasets designed for training reasoning models (System 2 thinking). It contains two main subsets:
RL (Reinforcement Learning): High-quality ground truth pairs augmented with task_type and rubrics for reward modeling.
SFT (Supervised Fine-Tuning): Instruction-following and thinking process data (with <think> tags).
Total… See the full description on the dataset page: https://huggingface.co/datasets/comoZ/reasoning-dataset.medical-reasoning-dataset
Dataset Card for my-distiset-2021d421
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-2021d421/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/medical-reasoning-dataset.math-python-reasoning-dataset
Dataset Card for my-distiset-3c1699f5
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-3c1699f5/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-python-reasoning-dataset.ccf-reasoning-dataset
Cognitive Cascade Framework (CCF) Reasoning Dataset
A high-quality dataset of structured reasoning examples using the Cognitive Cascade Framework (CCF), designed for training language models to perform systematic, multi-stage reasoning.
Dataset Description
This dataset contains problems across multiple domains (math, science, coding, creative reasoning) paired with detailed reasoning chains following the CCF methodology. Each example includes a complete reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/saberai/ccf-reasoning-dataset.Deepseek-mcq-reasoning-dataset
Turkish Reasoning Dataset
A Turkish reasoning dataset generated from alibayram/turkish_mmlu using DeepSeek-V3.2 (deepseek-reasoner). Each sample contains a multiple-choice academic question paired with a step-by-step rationale and internal thinking trace.
Dataset Summary
Source: Turkish MMLU (academic exam questions from TUS, KPSS, YKS, etc.)
Size: 1,000 samples
Language: Turkish
Generator Model: DeepSeek-V3.2 (deepseek-reasoner)
Purpose: Fine-tuning language models for… See the full description on the dataset page: https://huggingface.co/datasets/AhmetSemih/Deepseek-mcq-reasoning-dataset.Sindhi-Reasoning-Chat-Dataset
# 📘 **Sindhi‑Reasoning‑Chat‑Dataset**
A high‑quality, reasoning‑focused, instruction‑tuned conversational dataset for Sindhi (سنڌي), designed to support modern NLP research and LLM training for one of South Asia’s most underrepresented languages.
---
## 🧠 **Dataset Summary**
Sindhi is a **low‑resource South Asian language** spoken by **over 30 million people**, primarily in Sindh, Pakistan. Despite its rich cultural and literary heritage, Sindhi has **very limited publicly available… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Sindhi-Reasoning-Chat-Dataset.Reasoning_Problem_Solving_Dataset
Reasoning and Problem-Solving Dataset (RPSD)
Overview
The Reasoning and Problem-Solving Dataset (RPSD) is a comprehensive, high-quality set of synthetically generated question-answer pairs (150k+) tailored for training AI systems in logical reasoning and problem-solving. It spans multiple domains, including core reasoning techniques, specialized fields like science, mathematics, engineering, computer science, and philosophy, along with practical, real-world… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/Reasoning_Problem_Solving_Dataset.Nepali-Datasets-Reasoning-Grounding-V1Copyright 2026 Sandesh Bastola
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language… See the full description on the dataset page: https://huggingface.co/datasets/Matrix-Man-Lab/Nepali-Datasets-Reasoning-Grounding-V1.finance-reasoning-sft-dataset
Personal Finance Reasoning Dataset
A synthetic instruction-tuning dataset designed to teach language models to reason through personal finance and investing decisions using the mental frameworks from classic books in the genre. The goal is not recall of book content but principled reasoning: the model should apply frameworks to novel situations it has never seen.
Source Books
Principles were extracted from the following books:
The Psychology of Money — Morgan Housel
Rich… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/finance-reasoning-sft-dataset.Pashto-Medical-o1-Reasoning-SFT-Dataset
Pashto Medical o1 Reasoning SFT Dataset
This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models.
Dataset Structure
The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response.
Data Fields
Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.marathi-land-law-reasoning
Marathi Land Legal Reasoning (100 High-Quality Rows)
📌 Dataset Description
This is a high-precision reasoning dataset focused on the Maharashtra Land Revenue Code (MLRC) and property succession laws in India. It is specifically designed for fine-tuning Large Language Models (LLMs) to handle complex, domain-specific logic in the Marathi language.
Unlike generic datasets, this collection focuses on "Chain-of-Thought" (CoT), providing a detailed thought block for every… See the full description on the dataset page: https://huggingface.co/datasets/aditya-datasets/marathi-land-law-reasoning.team-truthowl-mixed-reasoning-dataset
Team P11 Mixed Reasoning Dataset
📊 Dataset description
HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。
推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。
対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。
🎯 Purpose
Competition: 松尾研LLMコンペ 2025
Target Model: DeepSeek-R1-Distill-Qwen-32B
Training Method: QLoRA Fine-tuning(4bit NF4, double quant)
📦 Composition
Math Hard(MATH Level≥3, HARDMath)
Math Mid(GSM8K, MetaMathQA)
Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.
