datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amc23aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.amc23AMC-23numina_162k_amc_aime_problemsnumina_amc_aime_deepseek_r1_responsesamc_filteredamc12-full
AMC12 Dataset (Research-Oriented)
A structured dataset derived from the AMC 12 (American Mathematics Competitions), designed for LLM training, evaluation, and reinforcement learning (RL) on mathematical reasoning tasks.
This repository contains all AMC 12 problems from 2000–2025, making it one of the most complete AMC12 datasets available for research.
📘 Introduction
The AMC 12 is a 25-question, 75-minute multiple-choice examination aimed at high school… See the full description on the dataset page: https://huggingface.co/datasets/edev2000/amc12-full.bespokelabs-sky-t1-numina-amc-aime-subset-unfilteredamc_aime_filtered2024_AMC12All problems copyrighted by the Mathematical Association of America's American Mathematics Competitions
Source:
https://artofproblemsolving.com/wiki/index.php/2024_AMC_12A_Problems
https://artofproblemsolving.com/wiki/index.php/2024_AMC_12B_Problems
Removed problems with figures:
12A: problem 14,18,22
12B: problem 7, 19
amc12-2025-non-figure
AMC 12 2025 (non-figure questions)
This dataset is a curated collection of AMC 12 2025 A/B problems that do not include figures, prepared for use in an academic study and released to support reproducibility.
Dataset Description
Each row corresponds to one AMC 12 problem and includes the following fields:
id: Stable identifier of the form amc12-2025-a1, amc12-2025-b25
year: 2025
contest: AMC 12
problem_number: Problem label from the source (e.g., A1–A25, B1–B25)… See the full description on the dataset page: https://huggingface.co/datasets/sonthenguyen/amc12-2025-non-figure.amc_aime_math_220k_subsetamc_aime_math_220k_filtered_single_genamc-booksamc3kamc_aime_linksR-HORIZON-AMC23
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.amc23amc_aime_diagramsamc_aime_figuresamc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
AMCrawlamc_aime_imgsamc-ruler-qwen35-32k
AMC RULER 32k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 32,768 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-32k.amc-ruler-qwen35-16k
AMC RULER 16k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 16,384 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-16k.amc_olympiads_math_mergedamc2kDataset Sources:
AMC 8 - AMC 10 - AMC 12
Both problems and solutions were scraped from their original URLs, preserving LaTeX format.
amc12_22-24amc2023
