datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olympiad_style_integer_math_problems
Olympiad Math Corpus
Version: v2.1.1
Release date: 2026-05-03
59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs.
Loading
from datasets import load_dataset
ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train")
lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.olympiad_style_integer_math_reasoning
Olympiad Math Reasoning Traces
Version: v1.0.2
Release date: 2026-04-19
64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.verified-math-olympiad-trajectories
Verified Math Olympiad Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's math olympiad reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), answer-verifier evaluation, process-supervision candidates, judge training, proof criticism, and private evaluations.
The goal is not merely to provide final-answer math examples. Each record is a structured olympiad reasoning object containing a normalized problem… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-math-olympiad-trajectories.Olympiads_hard
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 21525
Filtered size: 21408
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_hard.Olympiads_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 13284
Filtered size: 13240
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.Llama-3.2-1B-Instruct-uPRM-70B-T80-olympiadbench-best_of_n-completionsQwen2.5-1.5B-Instruct-uPRM-32B-T80-olympiadbench-best_of_n-completionsQwen2.5-1.5B-Instruct-uPRM-70B-T80-olympiadbench-best_of_n-completionsLlama-3.2-1B-Instruct-uPRM-32B-T80-olympiadbench-best_of_n-completionsolympiad-books-open-source
olympiad-books-open-source
Chunked content from 12 open-source mathematics textbooks, suitable for retrieval (RAG), embedding, and math reasoning research.
Source code: github.com/yoonholee/olympiad-books-open-source-pipeline
Books
Book
Author(s)
License
Source
An Infinitely Large Napkin
Evan Chen
CC BY-SA 4.0 / GPL v3
GitHub
Mathematical Reasoning: Writing and Proof
Ted Sundstrom
CC BY-NC-SA 3.0
GitHub
Exploring Combinatorial Mathematics
Richard Grassl… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/olympiad-books-open-source.Qwen2.5-7B-Instruct-uPRM-32B-T80-olympiadbench-best_of_n-completionsevolved-math-problems-OlympiadBench-from-deepseek-r1-0528-freeOlympiads_medium_filteredQwen2.5-7B-Instruct-uPRM-32B-T80-olympiadbench-dvts-completionsolympiads-proof-graderbencholympiadbench-a-subproblemolympiads-ref-cleaned-v1Olympiads_hard_filteredolympiads-ref-cleanedb2_math_fasttext_pos_brando_olympiad_neg_lap1official_mathqwen3_math_ai_olympiadbench_4xtrain_fasttext_classifier_seed_math_best_olympiadb2_train_fasttext_math_pos_brando_olympiad_neg_lap1official_mathb2_train_fasttext_math_pos_brando_olympiad_neg_lap1official_math_fixb2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
21.3
66.5
80.8
29.4
42.3
42.1
14.4
4.8
7.0
AIME24
Average Accuracy: 21.33% ± 1.17%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_eval_636d.b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_10k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
20.7
60.0
79.4
27.6
40.9
39.9
20.8
4.3
8.5
AIME24
Average Accuracy: 20.67% ± 1.40%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_10k_eval_636d.b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_10kdeepseek_r1_math_ai_olympiadbench_4xqwen_math_math_ai_olympiadbench_8xb2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_3k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
18.0
61.3
80.8
29.0
42.5
35.7
23.7
4.4
5.9
AIME24
Average Accuracy: 18.00% ± 1.65%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
20.00%
6
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_brando_olympiad_neg_lap1official_math_3k_eval_636d.
