datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.AIME-trajectory
AIME Trajectory Dataset
Model-generated solution trajectories for AIME (American Invitational Mathematics Examination) problems. Each row is one model response to a single problem, including the hidden chain-of-thoughts (when available), and the final response.
Dataset Summary
Split
Rows
Unique Problems
Years
Model(s)
Has reasoning_content
Accuracy
train
1,258
875
1983–2023
deepseek-r1
Yes
100%
test
180
30
2024
Multiple (see below)
No
3.3%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/AIME-trajectory.qwen3-8b-aime-2009-2024-16x
AIME Reasoning Traces · Qwen3-8B
7,680 reasoning traces for 480 AIME problems, with 16 sampled responses per problem. The corpus covers AIME I and AIME II from 2009 through 2024 and includes both correct and incorrect answers.
We created this dataset for How Should Incorrect Traces Be Used in Supervised Fine-Tuning? It is the source for the paper's larger AIME experiment, which compares correct and incorrect supervision across two disjoint sets of 106 problems.
Code and… See the full description on the dataset page: https://huggingface.co/datasets/suryadv/qwen3-8b-aime-2009-2024-16x.AIME24-25_CoT_Verification
Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
📌 Dataset Summary
This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.AIME_1983_2026_Kimi_K3
AIME 1983–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.SHP
🚢 Stanford Human Preferences Dataset (SHP)
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice.
The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training… See the full description on the dataset page: https://huggingface.co/datasets/aimeelizq/SHP.difficulty-aime_2025-generations
Generations Dataset: aime_2025
Paper: LLMs Encode Their Failures: Predicting Success from Pre-Generation ActivationsCode: GitHub
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-aime_2025-generations.AIME-24-25-26-clustered
AIME 2024-2026, clustered
90 AIME problems from 2024, 2025 and 2026, embedded and assigned to the three
semantic clusters used by CRE-Router.
Intended as a held-out test set for query routing: three years of 30 problems
each, spread evenly across clusters, so accuracy can be broken down by year.
Figure rule
Every problem is self-contained. Where a figure is needed to solve a problem it
is included as source, Asymptote or plain text; where a figure was decorative… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/AIME-24-25-26-clustered.AIME_2026_Kimi_K3
AIME 2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2026_Kimi_K3.AIME_2025_Kimi_K3
AIME 2025 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.
AIME… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2025_Kimi_K3.aime24-tr
AIME 2024 (Turkish) Dataset
This dataset contains the Turkish translations of problems from the 2024 American Invitational Mathematics Examination (AIME). It is intended to serve as a benchmark for evaluating the advanced mathematical reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The questions were translated into Turkish using GPT-5, then manually verified and corrected. Additional quality checks were performed to identify formatting issues, LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/aime24-tr.aime25-tr
AIME 2025 (Turkish) Dataset
This dataset contains the Turkish translations of problems from the 2025 American Invitational Mathematics Examination (AIME). It is intended to serve as a benchmark for evaluating the advanced mathematical reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The questions were translated into Turkish using GPT-5, then manually verified and corrected. Additional quality checks were performed to identify formatting issues, LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/aime25-tr.aime-1983-2024-ptpt
AIME-PT (1983-2024)
Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024.
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.aime2025-merging-leaderboard
AIME 2025 Evaluation Leaderboard
Evaluation results for 8 models on AIME 2025 with 30 problems and 32 rollouts per problem.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@32
avg@32
Loop Rate
allenai/Olmo-3.1-7B-RL-Zero-Math
80.0%
38.0%
2.9%
pmahdavi/Olmo-3.1-7B-Math-Code
73.3%
36.4%
2.4%… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/aime2025-merging-leaderboard.gemma-4-31b-it_aime-all
google/gemma-4-31b-it — aime-all
Model outputs from the micro-creativity inference suite.
Model: google/gemma-4-31b-it
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_aime-all.llama-3.1-8b-instruct_aime-all
meta-llama/Llama-3.1-8B-Instruct — aime-all
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_aime-all.mistral-small-3.2-24b-instruct-2506_aime-all
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — aime-all
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_aime-all.nanbeige4-3b-thinking-2511_aime-all
Nanbeige/Nanbeige4-3B-Thinking-2511 — aime-all
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_aime-all.olmo-3-7b-instruct_aime-all
allenai/OLMo-3-7B-Instruct — aime-all
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Instruct
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-instruct_aime-all.qwen3-30b-a3b_aime-all
Qwen/Qwen3-30B-A3B — aime-all
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-30B-A3B
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full model output… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-30b-a3b_aime-all.olmo-3-7b-think_aime-all
allenai/OLMo-3-7B-Think — aime-all
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_aime-all.gpt-oss-20b_aime-all
openai/gpt-oss-20b — aime-all
Model outputs from the micro-creativity inference suite.
Model: openai/gpt-oss-20b
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full model output… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gpt-oss-20b_aime-all.gemma-3-27b-it_aime-all
google/gemma-3-27b-it — aime-all
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_aime-all.Qwen3-0.6B-AIME-2023-2024-2025-sampling64
Model : Qwen3-0.6B
Original Dataset : AIME2023, AIME2024, AIME2025
Prompt:
{"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem}
Sampling Parameters :
num_sampling=64
max_tokens=38912
temperature=0.6
top_p=0.95
top_k=20
min_p=0
‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py)
qwen3-8b_aime-all
Qwen/Qwen3-8B — aime-all
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-8B
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full model output string… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_aime-all.Qwen3-32B-AIME-2023-2024-2025-sampling64
Model : Qwen3-32B
Original Dataset : first 24 queries in AIME2023 (시간 없어서 뒤에꺼 못함.)
Prompt:
{"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem}
Sampling Parameters :
num_sampling=64
max_tokens=38912
temperature=0.6
top_p=0.95
top_k=20
min_p=0
‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py)
