CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /Math-Reasoning Math-Reasoning Dataset Description Mathematical problem-solving, rewriting, and dialogue data for reasoning-oriented language-model training. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Math-Reasoning.texttext-generation1B<n<10B22 likes21k downloads20d agoHugging Face02ar0cket1 /qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised Qwen3-30B-A3B Reasoning SFT Prepacked Nemotron Math v4 CoT 4k-12k This dataset is a train-ready, offline-prepacked SFT corpus for full supervised fine-tuning of Qwen/Qwen3-30B-A3B-Base into a math reasoning model. Source And Filtering Source dataset: nvidia/Nemotron-SFT-Math-v4 Source revision: a94e56aeddcf6e75d28c8bd210f40fa62309288d Source split: train Intended subset: cot Preferred source during selection: AoPS Length filter: 4,000 to 12,000 supervised… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised.1 likes1.2k downloads2mo agoHugging Face03Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes697 downloads3mo agoHugging Face04stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes629 downloads2mo agoHugging Face05AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes414 downloads5d agoHugging Face06reasoning-proj /severity_ablation_mathtabular10K<n<100K0 likes390 downloads1y agoHugging Face07MARIO-Math-Reasoning /Gaokao2023-Math-En Data Summary This is a compilation of math test questions and answers drawn from the 2023 Chinese National College Entrance Examination, the 2023 American Mathematics Competitions, and the 2023 American College Testing. For simplicity, we refer to it as Gaokao2023. textn<1K7 likes372 downloads2y agoHugging Face08mihailgribov /olympiad_style_integer_math_reasoning Olympiad Math Reasoning Traces Version: v1.0.2 Release date: 2026-04-19 64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.tabulartext-generation10K<n<100K0 likes355 downloads5mo agoHugging Face09AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes339 downloads5d agoHugging Face10WNJXYK /MATH-Reasoning-Paths News 🌟🌟🌟 Try this dataset in our HuggingFace Space! 🥳🥳🥳 Thrilled to share that this NeurIPS paper was selected as 🏆 #1 Paper of the Day on Oct. 20th! Sampled Reasoning Paths for the MATH dataset This dataset contains sampled reasoning paths for the MATH dataset, released as part of the NeurIPS 2025 paper: "A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning" (Arxiv). Overview We generated… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/MATH-Reasoning-Paths.text-generation30 likes279 downloads11mo agoHugging Face11naimulislam /reasoning-math-advanced-1m 🧠 Reasoning Math Advanced 1M 📖 Dataset Summary Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks. A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.texttext-generation1M<n<10M0 likes269 downloads9mo agoHugging Face12vinhpx /math_reasoning_dataset_3Mtext1M<n<10M2 likes263 downloads1y agoHugging Face13MARIO-Math-Reasoning /AlphaMath-Trainset Dataset Card for AlphaMath Almost Zero This is the round 3 training data for AlphaMath Almost Zero: Process Supervision Without Process. The solution process was automatically generated by the model in round 2, without GPT or Human annotations. Dataset Details The question-answer pairs are extracted from the train split of GSM8k and MATH. Both positive and negative examples are included, for training both policy and value models. text100K<n<1M16 likes229 downloads2y agoHugging Face14CohenQu /CoRA_math_reasoning_benchmark_scalingtext1K<n<10K0 likes189 downloads1y agoHugging Face15CohenQu /CoRA_math_reasoning_benchmarktext1K<n<10K0 likes187 downloads1y agoHugging Face16thuzhizhi /DAPO-MATH-17k-oss-reasoning DAPO-MATH-17k-oss-reasoning This dataset contains reasoning trajectories produced by gpt-oss-120b on BytedTsinghua-SIA/DAPO-Math-17k. Under different reasoning efforts, we observe different token usage. Effort Level Avg Tokens Low 1300 Medium 2936 High 8419 Keywords appearance frequency indicates reasoning efforts of the LLM. Keyword Low Medium High All (L+M+H) wait 40.5% 69.0% 87.3% 65.6% double check 0.1% 2.4% 0.8% 1.9% check 57.9% 87.3%… See the full description on the dataset page: https://huggingface.co/datasets/thuzhizhi/DAPO-MATH-17k-oss-reasoning.text10K<n<100K3 likes172 downloads4mo agoHugging Face17CohenQu /CoRA_math_reasoning_benchmark_finaltext1K<n<10K0 likes168 downloads1y agoHugging Face18169Pi /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/mathreasoning.texttext-generation1M<n<10M6 likes165 downloads1y agoHugging Face19Ilia2003Mah /open-math-reasoning-cot-cleantext100K<n<1M1 likes165 downloads9mo agoHugging Face20Bdyskov /verified-math-reasoning verified-math-reasoning (CargoDash flagship recipe) A CargoDash framework demonstration. 999-row showcase of three-layer, program-verified, vote-stratified math reasoning traces — the dataset is small on purpose (its job is to prove the framework works on real production LLM endpoints, not to be a serious math benchmark). Each row carries three independent chain-of-thought solutions to the same problem (from DeepSeek, Doubao, and Qwen3.5) plus a programmatically extracted \boxed{}… See the full description on the dataset page: https://huggingface.co/datasets/Bdyskov/verified-math-reasoning.text-generationn<1K1 likes136 downloads4mo agoHugging Face21oddadmix /arabic-math-reasoning-synth Arabic Math Reasoning (synthetic) — مسائل رياضيات عربية مع خطوات الحل 120,462 Arabic grade-school math word problems, each with a step-by-step derivation and a concluding sentence. Generated with gemma-3-12b-it and Qwen3.8-27B-Uncensored-NVFP4 and arithmetically verified — every equation the reasoning states was re-evaluated, and rows whose own arithmetic does not check out were dropped. generator rows share gemma-3-12b-it 80,480 66.8% Qwen3.8-27B-Uncensored-NVFP4… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-math-reasoning-synth.texttext-generation100K<n<1M0 likes126 downloads25d agoHugging Face22khaled123 /MathReasoningtexttable-question-answering1K<n<10K5 likes116 downloads3y agoHugging Face23CohenQu /CoRA_math_reasoning_benchmark_DPO_hintstextn<1K0 likes95 downloads1y agoHugging Face24amphora /math-intuition-reasoning-traces math-intuition reasoning traces Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded by each problem family's own verifier. Questions come from amphora/math-intuition-20260908-402-easy-10 — 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id in that dataset, so prompts and the instance cache can be joined from it. Generation settings Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.tabulartext-generation10K<n<100K1 likes95 downloads11d agoHugging Face25dongboklee /math-reasoningtext10K<n<100K0 likes92 downloads9mo agoHugging Face26CohenQu /math_reasoning_benchmark_scaling_hint-gentextn<1K0 likes85 downloads1y agoHugging Face27CohenQu /arxiv_rlad_math_reasoning_benchmark_hintstext10K<n<100K0 likes84 downloads1y agoHugging Face28miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes83 downloads1y agoHugging Face29Ilia2003Mah /open-math-reasoning-cot-clean-v2text1M<n<10M0 likes81 downloads18d agoHugging Face30dvilasuero /gsm8k-math-reasoning-spanishtabularn<1K0 likes79 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.