CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes681 downloads3mo agoHugging Face02stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes652 downloads2mo agoHugging Face03mihailgribov /olympiad_style_integer_math_reasoning Olympiad Math Reasoning Traces Version: v1.0.2 Release date: 2026-04-19 64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.tabulartext-generation10K<n<100K0 likes351 downloads5mo agoHugging Face04naimulislam /reasoning-math-advanced-1m 🧠 Reasoning Math Advanced 1M 📖 Dataset Summary Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks. A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.texttext-generation1M<n<10M0 likes281 downloads9mo agoHugging Face05169Pi /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/mathreasoning.texttext-generation1M<n<10M6 likes160 downloads1y agoHugging Face06Bdyskov /verified-math-reasoning verified-math-reasoning (CargoDash flagship recipe) A CargoDash framework demonstration. 999-row showcase of three-layer, program-verified, vote-stratified math reasoning traces — the dataset is small on purpose (its job is to prove the framework works on real production LLM endpoints, not to be a serious math benchmark). Each row carries three independent chain-of-thought solutions to the same problem (from DeepSeek, Doubao, and Qwen3.5) plus a programmatically extracted \boxed{}… See the full description on the dataset page: https://huggingface.co/datasets/Bdyskov/verified-math-reasoning.text-generationn<1K1 likes112 downloads4mo agoHugging Face07miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes77 downloads1y agoHugging Face08Jackrong /GPT-OSS-120B-Distilled-Reasoning-math GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.textquestion-answering1K<n<10K9 likes77 downloads1y agoHugging Face09est-ai /math-reasoning-sft Mathematical Reasoning SFT Dataset This dataset contains mathematical reasoning problems and solutions in instruction-following format, designed for supervised fine-tuning of language models. Dataset Structure The dataset follows the Alpaca format with three fields: instruction: Mathematical problem statement input: Empty string (not used) output: Detailed solution with step-by-step reasoning and final answer in \boxed{} format Example { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/est-ai/math-reasoning-sft.texttext-generation1K<n<10K4 likes75 downloads1y agoHugging Face10erayalp /easy_turkish_math_reasoning Easy Turkish Math Reasoning Dataset Summary The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT). Use Case Primarily used for: Bootstrapping reasoning ability in Turkish for compact LLMs. Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.textquestion-answering1K<n<10K7 likes73 downloads1y agoHugging Face11HSH-Intelligence /verified-math-reasoning-3k HSH Verified Math Reasoning — Fine-Tuning Ready A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.texttext-generation1K<n<10K0 likes68 downloads3mo agoHugging Face12erayalp /medium_turkish_math_reasoning Dataset Summary The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning. Use Case This dataset is ideal for: Continuing SFT after foundational training with simpler problems. Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.textquestion-answering1K<n<10K5 likes63 downloads1y agoHugging Face13aloks16 /mathreasoning MathReasoning The MathReasoning Dataset is a large-scale, high-quality dataset (~3.13M rows) focused on mathematics, logical reasoning, and problem-solving. It is primarily generated through synthetic distillation techniques, complemented by curated open-source educational content. The dataset is designed to train and evaluate language models in mathematical reasoning, quantitative problem-solving, and structured chain-of-thought tasks across domains from basic arithmetic to… See the full description on the dataset page: https://huggingface.co/datasets/aloks16/mathreasoning.texttext-generation1M<n<10M0 likes57 downloads8mo agoHugging Face14sdiazlor /math-python-reasoning-dataset Dataset Card for my-distiset-3c1699f5 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-3c1699f5/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/math-python-reasoning-dataset.texttext-generationn<1K3 likes37 downloads2y agoHugging Face15Caplin43 /ai-reasoning-math-dataset 🧮 AI Reasoning Math Dataset Dataset containing math word problems with step-by-step reasoning and final answers. Designed for: Chain-of-thought training Reasoning model fine-tuning Math QA benchmarking 📊 Dataset Statistics Train: 5,000 samples Validation: 1,000 samples Test: 1,000 samples Total: 7,000 samples 📄 Data Format { "question": "If a train travels 60 km in 1.5 hours, what is its average speed?", "reasoning": "Average speed = distance /… See the full description on the dataset page: https://huggingface.co/datasets/Caplin43/ai-reasoning-math-dataset.text-generation1K<n<10K0 likes26 downloads7mo agoHugging Face16Thiraput01 /Math-reasoning-Opus4.6-typhoon-translated Dataset Card for Math-reasoning-Opus4.6-typhoon-translated Dataset Description This dataset is a Thai-translated version of the Crownelius/Opus-4.6-Reasoning-3300x dataset. It is designed to train and evaluate mathematical reasoning capabilities in Thai language models. The original English dataset was translated into Thai using the scb10x/typhoon-translate1.5-4b model, providing high-quality, localized mathematical problems, step-by-step thinking processes, and… See the full description on the dataset page: https://huggingface.co/datasets/Thiraput01/Math-reasoning-Opus4.6-typhoon-translated.texttext-generation1K<n<10K0 likes25 downloads6mo agoHugging Face17NNEngine /Reasoning-Heavy-Math-ML-Explanations Reasoning-Heavy Math & ML Explanations Dataset: NNEngine/Reasoning-Heavy-Math-ML-Explanations Version: wikipedia_reasoning_final_v1.0 License: CC-BY-SA 4.0 Author: Shivam Sharma (Independent Researcher) Dataset Summary Reasoning-Heavy Math & ML Explanations is a high-quality, reasoning-oriented dataset derived exclusively from English Wikipedia. The dataset focuses on explicit human-authored reasoning and explanations in mathematics and machine learning–related domains… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Reasoning-Heavy-Math-ML-Explanations.texttext-generation1K<n<10K0 likes23 downloads9mo agoHugging Face18sumeetrm /math-reasoning-benchmark [!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. Math Reasoning Benchmark Evaluating LLMs on Chained Multi-Step Mathematical Reasoning Leaderboard Overview The Math Reasoning Benchmark evaluates language models on their ability to solve chained multi-step mathematical problems. Each question is a directed acyclic graph (DAG) of math sub-problems ("nodes"), where… See the full description on the dataset page: https://huggingface.co/datasets/sumeetrm/math-reasoning-benchmark.textquestion-answeringn<1K0 likes23 downloads5mo agoHugging Face19Uunan /turkish-math-reasoning Turkish Math Reasoning Dataset (ChatML) Bu veri seti, Türkçe dilinde yapay zeka modellerinin matematiksel mantık yürütme (reasoning/chain-of-thought) yeteneklerini geliştirmek için özel olarak hazırlanmıştır. Veri setindeki problemlerin tamamı Toplama (+) ve Çıkarma (-) işlemlerinden oluşmaktadır. Özellikler Doğrulanmış İçerik: Veri setindeki her bir satır algoritmik olarak kontrol edilmiş, matematiksel olarak hatalı olan (yanlış işlem yapan) model çıktıları… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-math-reasoning.text-generation10K<n<100K0 likes20 downloads2mo agoHugging Face20est-ai /math-reasoning-dpo Mathematical Reasoning DPO Dataset This dataset contains mathematical reasoning problems with chosen and rejected responses, designed for Direct Preference Optimization (DPO) and preference learning of language models. Dataset Structure The dataset follows the ShareGPT format for DPO training with three main fields: conversations: List of conversation turns leading up to the response chosen: Preferred response with detailed reasoning and correct solution rejected: Less… See the full description on the dataset page: https://huggingface.co/datasets/est-ai/math-reasoning-dpo.texttext-generation1K<n<10K0 likes19 downloads1y agoHugging Face21BakeAI /BakeAI_Reasoning_Math_L5_2603_Previewgated BakeAI Reasoning Math L5 2603 Preview Dataset Summary This dataset contains 50 challenging, university-level mathematics reasoning problems with detailed reference solutions, structured grading rubrics, and anonymized model evaluation results. Each problem: Requires multi-step reasoning, proof construction, or complex computation Includes a structured rubric with point-by-point grading criteria Contains a frontier model attempt that was evaluated against the… See the full description on the dataset page: https://huggingface.co/datasets/BakeAI/BakeAI_Reasoning_Math_L5_2603_Preview.question-answeringn<1K1 likes19 downloads7mo agoHugging Face22BakeAI /BakeAI_Reasoning_Math_L4_2603_Previewgated BakeAI Reasoning Math L4 2603 Preview Dataset Summary This dataset contains 50 challenging, competition-level mathematics reasoning problems with detailed reference solutions, structured grading rubrics, and anonymized model evaluation results. Each problem: Requires multi-step reasoning, proof construction, or complex computation Includes a structured rubric with point-by-point grading criteria Contains a frontier model attempt that was evaluated against the… See the full description on the dataset page: https://huggingface.co/datasets/BakeAI/BakeAI_Reasoning_Math_L4_2603_Preview.textquestion-answeringn<1K1 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.