datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.math-reasoning-benchmark
[!NOTE]
IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset.
Math Reasoning Benchmark
Evaluating LLMs on Chained Multi-Step Mathematical Reasoning
Leaderboard
Overview
The Math Reasoning Benchmark evaluates language models on their ability to solve chained multi-step mathematical problems. Each question is a directed acyclic graph (DAG) of math sub-problems ("nodes"), where… See the full description on the dataset page: https://huggingface.co/datasets/sumeetrm/math-reasoning-benchmark.
