datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
math500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
MATH500https://github.com/openai/prm800k/blob/main/prm800k/math_splits/test.jsonl
MATH-500MATH-500-multilingual
MATH-500 Multilingual Problem Set 🌍➗
A multilingual subset from OpenAI's MATH benchmark. Perfect for testing math skills across languages, this dataset includes same problems in English, French, Italian, Turkish and Spanish.
🌐 Available Languages
English 🇬🇧
French 🇫🇷
Italian 🇮🇹
Turkish 🇹🇷
Spanish 🇪🇸
📂 Source & Attribution
Original Dataset: Sourced from HuggingFaceH4/MATH-500.
🚀 Quick Start
Load the dataset… See the full description on the dataset page: https://huggingface.co/datasets/bezir/MATH-500-multilingual.R-HORIZON-Math500R-HORIZON-Math500
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-Math500.math500-floatExtracted 316 examples with plain‐decimal answers from the original dataset loaded by datasets.load_dataset("HuggingFaceH4/MATH-500", split="test"). The processing code demonstrates in the notebook file.
MATH-500-Russian
Карточка датасета MATH-500-Russian
Перевод датасета HuggingFaceH4/MATH-500 на русский язык,
был выполнен моделью qwen2.5:32b через
скрипты EvilFreelancer/datasets-translator.
Данный набор данных содержит подмножество из 500 задач из теста MATH, который OpenAI создал для статьи Let's Verify
Step by Step и переведённых на русский язык.
Подробности в их репозиторий на GitHub.
formal_math500
Dataset Card for Formal Problem-Solving Benchmarks
This dataset is part of the official implementation of Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving, accepted as an ICML 2026 Spotlight paper.
Links:
Paper: https://openreview.net/forum?id=hgMZraPlSv
Project: https://github.com/Purewhite2019/formal_problem_solving_main
Overview
The benchmark supports three evaluation settings:
Formal Problem-Solving (FPS): Given a… See the full description on the dataset page: https://huggingface.co/datasets/purewhite42/formal_math500.MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
MATH-500
Dataset Card for MATH-500
This dataset contains 12,000 training problems and 500 test problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
Compared with the existing repository, this version additionally includes the training split.
MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
MATH500MATH_500_MMLU_Pro
Description
The MATH_500 and MMLU_Pro datasets combined in a custom format.
Format for the data
Field
Type
Description
unique_id
string
MD5 hash of the entire record (all other columns), computed on the canonical JSON representation with sorted keys.
question
string
The full question text.
category
string
The subject or category of the question.
choices
array of strings | null
List of answer choices for multiple-choice questions; null for open-ended… See the full description on the dataset page: https://huggingface.co/datasets/HPC-Boys/MATH_500_MMLU_Pro.math500-preference-pairs-fable
MATH-500 Preference Pairs (Fable-generated)
500 preference pairs covering all 500 problems of MATH-500, generated by Anthropic's Claude Fable 5 for reward-model training in a math-RLHF project (Qwen2.5-7B, PPO/GRPO on AWS EKS).
Format
{
"idx": 0,
"problem": "Convert the point $(0,3)$ ...",
"chosen": "<complete correct solution with full reasoning>",
"rejected_1": "<plausible-but-wrong solution, error mode A>",
"rejected_2": "<plausible-but-wrong solution… See the full description on the dataset page: https://huggingface.co/datasets/weivzhang/math500-preference-pairs-fable.MATH500math-500_6reason-14B-MATH500MATH-500
Dataset Card for MATH-500
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
gpt4o-math500-t0
gpt4o-math500-t0
This dataset repo contains external-model traces used in the Illusion-of-Reasoning analysis.
Domain: MATH-500
Model: GPT-4o
Target temperature label: 0
Rows in data.jsonl: 4000
Source artifact: /n/fs/similarity/Illusion-of-Reasoning/artifacts/results/gpt4o-math-portkey/step0000_test.jsonl
Notes:
Full G=8 traces from local artifact root.
Rows are uploaded as raw JSONL records from the local analysis artifacts.
qwq-MATH500-94math500_stop_stringsMath500-basedeepseek-r1-math500-t03
deepseek-r1-math500-t03
This dataset repo contains external-model traces used in the Illusion-of-Reasoning analysis.
Domain: MATH-500
Model: DeepSeek-R1
Target temperature label: 0.3
Rows in data.jsonl: 4000
Source artifact: /n/fs/similarity/Illusion-of-Reasoning/artifacts/results/deepseek-r1-openrouter-temp03/step0000_test.jsonl
Notes:
Full G=8 traces from local artifact root.
Rows are uploaded as raw JSONL records from the local analysis artifacts.
MATH-500-self-rewarding使用self-rewarding方法微调的模型,在math-500上的结果
模型:qwen2.5-7b-insturct
方法:(Self-rewarding correction for mathematical reasoning)[https://arxiv.org/pdf/2502.19613]
MATH500-qwen-math-instruct-stepdeepseek-r1-math500-t005
deepseek-r1-math500-t005
This dataset repo contains external-model traces used in the Illusion-of-Reasoning analysis.
Domain: MATH-500
Model: DeepSeek-R1
Target temperature label: 0.05
Rows in data.jsonl: 4000
Source artifact: /n/fs/similarity/Illusion-of-Reasoning/artifacts/results/deepseek-r1-openrouter-temp005/step0000_test.jsonl
Notes:
Full G=8 traces from local artifact root.
Rows are uploaded as raw JSONL records from the local analysis artifacts.
deepseek-r1-math500-t1
deepseek-r1-math500-t1
This dataset repo contains external-model traces used in the Illusion-of-Reasoning analysis.
Domain: MATH-500
Model: DeepSeek-R1
Target temperature label: 1
Rows in data.jsonl: 4000
Source artifact: /n/fs/similarity/Illusion-of-Reasoning/artifacts/results/deepseek-r1-openrouter-temp1/step0000_test.jsonl
Notes:
Full G=8 traces from local artifact root.
Rows are uploaded as raw JSONL records from the local analysis artifacts.
Math500-instruct
