CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes654 downloads2mo agoHugging Face02Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes633 downloads3mo agoHugging Face03naimulislam /reasoning-math-advanced-1m 🧠 Reasoning Math Advanced 1M 📖 Dataset Summary Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks. A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.texttext-generation1M<n<10M0 likes278 downloads9mo agoHugging Face04mihailgribov /olympiad_style_integer_math_reasoning Olympiad Math Reasoning Traces Version: v1.0.2 Release date: 2026-04-19 64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.tabulartext-generation10K<n<100K0 likes265 downloads5mo agoHugging Face05169Pi /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/mathreasoning.texttext-generation1M<n<10M6 likes155 downloads1y agoHugging Face06Jackrong /GPT-OSS-120B-Distilled-Reasoning-math GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.textquestion-answering1K<n<10K9 likes78 downloads1y agoHugging Face07est-ai /math-reasoning-sft Mathematical Reasoning SFT Dataset This dataset contains mathematical reasoning problems and solutions in instruction-following format, designed for supervised fine-tuning of language models. Dataset Structure The dataset follows the Alpaca format with three fields: instruction: Mathematical problem statement input: Empty string (not used) output: Detailed solution with step-by-step reasoning and final answer in \boxed{} format Example { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/est-ai/math-reasoning-sft.texttext-generation1K<n<10K4 likes75 downloads1y agoHugging Face08miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes71 downloads1y agoHugging Face09erayalp /easy_turkish_math_reasoning Easy Turkish Math Reasoning Dataset Summary The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT). Use Case Primarily used for: Bootstrapping reasoning ability in Turkish for compact LLMs. Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.textquestion-answering1K<n<10K7 likes65 downloads1y agoHugging Face10erayalp /medium_turkish_math_reasoning Dataset Summary The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning. Use Case This dataset is ideal for: Continuing SFT after foundational training with simpler problems. Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.textquestion-answering1K<n<10K5 likes65 downloads1y agoHugging Face11aloks16 /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/aloks16/mathreasoning.texttext-generation1M<n<10M0 likes64 downloads8mo agoHugging Face12HSH-Intelligence /verified-math-reasoning-3k HSH Verified Math Reasoning — Fine-Tuning Ready A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.texttext-generation1K<n<10K0 likes58 downloads3mo agoHugging Face13sdiazlor /math-python-reasoning-dataset Dataset Card for my-distiset-3c1699f5 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-3c1699f5/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-python-reasoning-dataset.texttext-generationn<1K3 likes37 downloads2y agoHugging Face14Thiraput01 /Math-reasoning-Opus4.6-typhoon-translated Dataset Card for Math-reasoning-Opus4.6-typhoon-translated Dataset Description This dataset is a Thai-translated version of the Crownelius/Opus-4.6-Reasoning-3300x dataset. It is designed to train and evaluate mathematical reasoning capabilities in Thai language models. The original English dataset was translated into Thai using the scb10x/typhoon-translate1.5-4b model, providing high-quality, localized mathematical problems, step-by-step thinking processes, and… See the full description on the dataset page: https://huggingface.co/datasets/Thiraput01/Math-reasoning-Opus4.6-typhoon-translated.texttext-generation1K<n<10K0 likes28 downloads6mo agoHugging Face15NNEngine /Reasoning-Heavy-Math-ML-Explanations Reasoning-Heavy Math & ML Explanations Dataset: NNEngine/Reasoning-Heavy-Math-ML-Explanations Version: wikipedia_reasoning_final_v1.0 License: CC-BY-SA 4.0 Author: Shivam Sharma (Independent Researcher) Dataset Summary Reasoning-Heavy Math & ML Explanations is a high-quality, reasoning-oriented dataset derived exclusively from English Wikipedia. The dataset focuses on explicit human-authored reasoning and explanations in mathematics and machine learning–related domains… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Reasoning-Heavy-Math-ML-Explanations.texttext-generation1K<n<10K0 likes21 downloads9mo agoHugging Face16sumeetrm /math-reasoning-benchmark [!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. Math Reasoning Benchmark Evaluating LLMs on Chained Multi-Step Mathematical Reasoning Leaderboard Overview The Math Reasoning Benchmark evaluates language models on their ability to solve chained multi-step mathematical problems. Each question is a directed acyclic graph (DAG) of math sub-problems ("nodes"), where… See the full description on the dataset page: https://huggingface.co/datasets/sumeetrm/math-reasoning-benchmark.textquestion-answeringn<1K0 likes20 downloads5mo agoHugging Face17est-ai /math-reasoning-dpo Mathematical Reasoning DPO Dataset This dataset contains mathematical reasoning problems with chosen and rejected responses, designed for Direct Preference Optimization (DPO) and preference learning of language models. Dataset Structure The dataset follows the ShareGPT format for DPO training with three main fields: conversations: List of conversation turns leading up to the response chosen: Preferred response with detailed reasoning and correct solution rejected: Less… See the full description on the dataset page: https://huggingface.co/datasets/est-ai/math-reasoning-dpo.texttext-generation1K<n<10K0 likes19 downloads1y agoHugging Face18BakeAI /BakeAI_Reasoning_Math_L4_2603_Previewgated BakeAI Reasoning Math L4 2603 Preview Dataset Summary This dataset contains 50 challenging, competition-level mathematics reasoning problems with detailed reference solutions, structured grading rubrics, and anonymized model evaluation results. Each problem: Requires multi-step reasoning, proof construction, or complex computation Includes a structured rubric with point-by-point grading criteria Contains a frontier model attempt that was evaluated against the… See the full description on the dataset page: https://huggingface.co/datasets/BakeAI/BakeAI_Reasoning_Math_L4_2603_Preview.textquestion-answeringn<1K1 likes16 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.