datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Math-Reasoning
Math-Reasoning
Dataset Description
Mathematical problem-solving, rewriting, and dialogue data for reasoning-oriented language-model training. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets
IFM/TxT360-v2… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Math-Reasoning.qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised
Qwen3-30B-A3B Reasoning SFT Prepacked Nemotron Math v4 CoT 4k-12k
This dataset is a train-ready, offline-prepacked SFT corpus for full supervised
fine-tuning of Qwen/Qwen3-30B-A3B-Base into a math reasoning model.
Source And Filtering
Source dataset: nvidia/Nemotron-SFT-Math-v4
Source revision: a94e56aeddcf6e75d28c8bd210f40fa62309288d
Source split: train
Intended subset: cot
Preferred source during selection: AoPS
Length filter: 4,000 to 12,000 supervised… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised.math-reasoning-ift-pairs
Reasoning-IFT Pairs (Math Domain)
Paper | Project Page
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain).
It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data.
We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.severity_ablation_mathGaokao2023-Math-En
Data Summary
This is a compilation of math test questions and answers drawn from the 2023 Chinese National College Entrance Examination, the 2023 American Mathematics Competitions, and the 2023 American College Testing. For simplicity, we refer to it as Gaokao2023.
olympiad_style_integer_math_reasoning
Olympiad Math Reasoning Traces
Version: v1.0.2
Release date: 2026-04-19
64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.MATH-Reasoning-Paths
News
🌟🌟🌟 Try this dataset in our HuggingFace Space!
🥳🥳🥳 Thrilled to share that this NeurIPS paper was selected as 🏆 #1 Paper of the Day on Oct. 20th!
Sampled Reasoning Paths for the MATH dataset
This dataset contains sampled reasoning paths for the MATH dataset, released as part of the NeurIPS 2025 paper: "A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning" (Arxiv).
Overview
We generated… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/MATH-Reasoning-Paths.reasoning-math-advanced-1m
🧠 Reasoning Math Advanced 1M
📖 Dataset Summary
Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks.
A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.math_reasoning_dataset_3MAlphaMath-Trainset
Dataset Card for AlphaMath Almost Zero
This is the round 3 training data for AlphaMath Almost Zero: Process Supervision Without Process. The solution process was automatically generated by the model in round 2, without GPT or Human annotations.
Dataset Details
The question-answer pairs are extracted from the train split of GSM8k and MATH.
Both positive and negative examples are included, for training both policy and value models.
CoRA_math_reasoning_benchmark_scalingCoRA_math_reasoning_benchmarkDAPO-MATH-17k-oss-reasoning
DAPO-MATH-17k-oss-reasoning
This dataset contains reasoning trajectories produced by gpt-oss-120b on BytedTsinghua-SIA/DAPO-Math-17k.
Under different reasoning efforts, we observe different token usage.
Effort Level
Avg Tokens
Low
1300
Medium
2936
High
8419
Keywords appearance frequency indicates reasoning efforts of the LLM.
Keyword
Low
Medium
High
All (L+M+H)
wait
40.5%
69.0%
87.3%
65.6%
double check
0.1%
2.4%
0.8%
1.9%
check
57.9%
87.3%… See the full description on the dataset page: https://huggingface.co/datasets/thuzhizhi/DAPO-MATH-17k-oss-reasoning.CoRA_math_reasoning_benchmark_finalmathreasoning
MathReasoning
The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/mathreasoning.open-math-reasoning-cot-cleanverified-math-reasoning
verified-math-reasoning (CargoDash flagship recipe)
A CargoDash framework demonstration. 999-row showcase of three-layer,
program-verified, vote-stratified math reasoning traces — the dataset is
small on purpose (its job is to prove the framework works on real
production LLM endpoints, not to be a serious math benchmark). Each row
carries three independent chain-of-thought solutions to the same problem
(from DeepSeek, Doubao, and Qwen3.5) plus a programmatically extracted
\boxed{}… See the full description on the dataset page: https://huggingface.co/datasets/Bdyskov/verified-math-reasoning.arabic-math-reasoning-synth
Arabic Math Reasoning (synthetic) — مسائل رياضيات عربية مع خطوات الحل
120,462 Arabic grade-school math word problems, each with a step-by-step derivation and a
concluding sentence. Generated with gemma-3-12b-it and Qwen3.8-27B-Uncensored-NVFP4 and
arithmetically verified — every equation the reasoning states was re-evaluated, and rows whose
own arithmetic does not check out were dropped.
generator
rows
share
gemma-3-12b-it
80,480
66.8%
Qwen3.8-27B-Uncensored-NVFP4… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-math-reasoning-synth.MathReasoningCoRA_math_reasoning_benchmark_DPO_hintsmath-intuition-reasoning-traces
math-intuition reasoning traces
Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded
by each problem family's own verifier.
Questions come from
amphora/math-intuition-20260908-402-easy-10
— 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id
in that dataset, so prompts and the instance cache can be joined from it.
Generation settings
Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.math-reasoningmath_reasoning_benchmark_scaling_hint-genarxiv_rlad_math_reasoning_benchmark_hintsMath_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.open-math-reasoning-cot-clean-v2gsm8k-math-reasoning-spanish
