datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math_full_minus_math500
MATH (minus MATH-500)
This dataset is derived from the original MATH dataset by Hendrycks et al.
(qwedsacf/competition_math) with all problems from the MATH-500 benchmark set removed.
Construction
Source: 12,500 problems from the MATH dataset by Hendrycks et al. (qwedsacf/competition_math)
Benchmark held out: 500 problems from the MATH-500 dataset (HuggingFaceH4/MATH-500)
Matching criterion: exact match on the problem field (see… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/math_full_minus_math500.math-500-th
Math-500-th
A Thai translation of MATH-500: the 500-problem subset of the MATH benchmark used
in OpenAI's Let's Verify Step by Step. Every row corresponds 1:1, in order, to a
row of the English original, so the Thai and English scores of a model are directly
comparable.
Source and licence
Original benchmark
hendrycks/math — MIT
500-problem subset
openai/prm800k — MIT
File we translated from
HuggingFaceH4/MATH-500
This dataset
MIT, see LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/iapp/math-500-th.math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on MATH500
This dataset contains 99 self-consistency generations per question for the
MATH500 benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: HuggingFaceH4/MATH-500
Model:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs.MATH-500-Overall
MATH-500-Overall
About the dataset
This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct.
Brief information
Number of rows: 500
Type of dataset files: parquet
Type of dataset: text, alpaca with system prompts
Language: English
License: MIT
Structure:
math¯¯¯¯¯⌉
school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.math500-deepseek-r1-distill-qwen-1.5b
Dataset Card for "math500-deepseek-r1-distill-qwen-1.5b"
Dataset Summary
This dataset is a distilled version of the MATH500 dataset, augmented with reasoning-based responses generated by the deepseek-r1-distill-qwen-1.5b language model. The dataset is designed to evaluate and improve the mathematical reasoning capabilities of LLMs through step-by-step solutions and final answers.
Each example consists of:
The original problem statement from MATH500
The reference solution… See the full description on the dataset page: https://huggingface.co/datasets/jsm0424/math500-deepseek-r1-distill-qwen-1.5b.math500
math500
A subset of 500 mathematical problems from the MATH dataset, covering algebra, precalculus, number theory, and geometry.
Dataset Structure
This dataset is in Hugging Face datasets format. Load it with:
from datasets import load_dataset
dataset = load_dataset("Tyrion279/math500")
math-500-ptpt
Math-500-PT
Portuguese machine translation of Math-500, a benchmark of 500 challenging math problems across various topics.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/HuggingFaceH4/MATH-500
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/math-500-ptpt.math500_best_of_n
Dataset Summary
This dataset contains mathematical problem-solving responses generated using two decoding methods: Greedy and Best-of-N for 20 problems spanning three difficulty levels (1, 2, 3) from the [MATH500 dataset](https://huggingface.co/datasets/HuggingFaceH4/MATH-500.
Languages
The dataset content is entirely in English.
Dataset Structure
Each instance contains the following fields:
problem: math problem given as an input
answer: a ground truth answer… See the full description on the dataset page: https://huggingface.co/datasets/Jakh0103/math500_best_of_n.math500-enhanced
Math500 Enhanced Dataset
This dataset contains LLM-enhanced versions of mathematical problems with step-by-step reasoning solutions.
Dataset Statistics
Examples: 500 (500 enhanced with LLM)
Enhancement Rate: 100.0%
Data Fields
question: The mathematical problem statement
solution: LLM-enhanced step-by-step solution
original_solution: Original solution text (for reference)
answer: Final numerical answer
level: Problem difficulty level
type: Problem… See the full description on the dataset page: https://huggingface.co/datasets/rachitbansal-harvard/math500-enhanced.math500-deepseek-r1-distill-qwen-14b
Dataset Card for "math500-deepseek-r1-distill-qwen-14b"
Dataset Summary
This dataset is a distilled version of the MATH500 dataset, augmented with reasoning-based responses generated by the deepseek-r1-distill-qwen-14b language model. The dataset is designed to evaluate and improve the mathematical reasoning capabilities of LLMs through step-by-step solutions and final answers.
Each example consists of:
The original problem statement from MATH500
The reference solution… See the full description on the dataset page: https://huggingface.co/datasets/jsm0424/math500-deepseek-r1-distill-qwen-14b.
