CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /Math-Reasoning Math-Reasoning Dataset Description Mathematical problem-solving, rewriting, and dialogue data for reasoning-oriented language-model training. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Math-Reasoning.texttext-generation1B<n<10B22 likes25k downloads21d agoHugging Face02Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes725 downloads3mo agoHugging Face03stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes648 downloads2mo agoHugging Face04AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes419 downloads6d agoHugging Face05reasoning-proj /severity_ablation_mathtabular10K<n<100K0 likes392 downloads1y agoHugging Face06mihailgribov /olympiad_style_integer_math_reasoning Olympiad Math Reasoning Traces Version: v1.0.2 Release date: 2026-04-19 64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.tabulartext-generation10K<n<100K0 likes365 downloads5mo agoHugging Face07MARIO-Math-Reasoning /Gaokao2023-Math-En Data Summary This is a compilation of math test questions and answers drawn from the 2023 Chinese National College Entrance Examination, the 2023 American Mathematics Competitions, and the 2023 American College Testing. For simplicity, we refer to it as Gaokao2023. textn<1K7 likes362 downloads2y agoHugging Face08AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes357 downloads6d agoHugging Face09vinhpx /math_reasoning_dataset_3Mtext1M<n<10M2 likes286 downloads1y agoHugging Face10naimulislam /reasoning-math-advanced-1m 🧠 Reasoning Math Advanced 1M 📖 Dataset Summary Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks. A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.texttext-generation1M<n<10M0 likes286 downloads9mo agoHugging Face11MARIO-Math-Reasoning /AlphaMath-Trainset Dataset Card for AlphaMath Almost Zero This is the round 3 training data for AlphaMath Almost Zero: Process Supervision Without Process. The solution process was automatically generated by the model in round 2, without GPT or Human annotations. Dataset Details The question-answer pairs are extracted from the train split of GSM8k and MATH. Both positive and negative examples are included, for training both policy and value models. text100K<n<1M16 likes233 downloads2y agoHugging Face12CohenQu /CoRA_math_reasoning_benchmark_scalingtext1K<n<10K0 likes190 downloads1y agoHugging Face13CohenQu /CoRA_math_reasoning_benchmarktext1K<n<10K0 likes188 downloads1y agoHugging Face14thuzhizhi /DAPO-MATH-17k-oss-reasoning DAPO-MATH-17k-oss-reasoning This dataset contains reasoning trajectories produced by gpt-oss-120b on BytedTsinghua-SIA/DAPO-Math-17k. Under different reasoning efforts, we observe different token usage. Effort Level Avg Tokens Low 1300 Medium 2936 High 8419 Keywords appearance frequency indicates reasoning efforts of the LLM. Keyword Low Medium High All (L+M+H) wait 40.5% 69.0% 87.3% 65.6% double check 0.1% 2.4% 0.8% 1.9% check 57.9% 87.3%… See the full description on the dataset page: https://huggingface.co/datasets/thuzhizhi/DAPO-MATH-17k-oss-reasoning.text10K<n<100K3 likes173 downloads4mo agoHugging Face15CohenQu /CoRA_math_reasoning_benchmark_finaltext1K<n<10K0 likes169 downloads1y agoHugging Face16169Pi /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/mathreasoning.texttext-generation1M<n<10M6 likes161 downloads1y agoHugging Face17oddadmix /arabic-math-reasoning-synth Arabic Math Reasoning (synthetic) — مسائل رياضيات عربية مع خطوات الحل 120,462 Arabic grade-school math word problems, each with a step-by-step derivation and a concluding sentence. Generated with gemma-3-12b-it and Qwen3.8-27B-Uncensored-NVFP4 and arithmetically verified — every equation the reasoning states was re-evaluated, and rows whose own arithmetic does not check out were dropped. generator rows share gemma-3-12b-it 80,480 66.8% Qwen3.8-27B-Uncensored-NVFP4… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-math-reasoning-synth.texttext-generation100K<n<1M0 likes128 downloads26d agoHugging Face18Ilia2003Mah /open-math-reasoning-cot-cleantext100K<n<1M1 likes105 downloads9mo agoHugging Face19khaled123 /MathReasoningtexttable-question-answering1K<n<10K5 likes104 downloads3y agoHugging Face20CohenQu /CoRA_math_reasoning_benchmark_DPO_hintstextn<1K0 likes97 downloads1y agoHugging Face21amphora /math-intuition-reasoning-traces math-intuition reasoning traces Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded by each problem family's own verifier. Questions come from amphora/math-intuition-20260908-402-easy-10 — 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id in that dataset, so prompts and the instance cache can be joined from it. Generation settings Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.tabulartext-generation10K<n<100K1 likes97 downloads12d agoHugging Face22dongboklee /math-reasoningtext10K<n<100K0 likes93 downloads9mo agoHugging Face23CohenQu /math_reasoning_benchmark_scaling_hint-gentextn<1K0 likes86 downloads1y agoHugging Face24CohenQu /arxiv_rlad_math_reasoning_benchmark_hintstext10K<n<100K0 likes85 downloads1y agoHugging Face25LangAGI-Lab /magpie-reasoning-v1-20k-math-verifiable-step-by-step-rationale-alpaca-formattext10K<n<100K5 likes84 downloads2y agoHugging Face26Ilia2003Mah /open-math-reasoning-cot-clean-v2text1M<n<10M0 likes84 downloads19d agoHugging Face27miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes81 downloads1y agoHugging Face28dvilasuero /gsm8k-math-reasoning-spanishtabularn<1K0 likes80 downloads1y agoHugging Face29shangshang /math-reasoning-zh Math Reasoning Chinese Dataset Overview A high-quality Chinese math word problem reasoning dataset with 500 problems featuring detailed Chain-of-Thought reasoning processes. All problems are algorithmically verified for 100% correctness. Dataset Structure Field Description Example problem_id Unique identifier math_0001 question Problem text 商店原价800元的商品打75折后... chain_of_thought Step-by-step reasoning 先算打折后价格:800 × 75% = 600元...… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/math-reasoning-zh.textn<1K0 likes79 downloads22d agoHugging Face30RedStar-Reasoning /math_datasettext1K<n<10K1 likes77 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.