datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.amc23amc_filteredR-HORIZON-AMC23
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.amc-ruler-qwen35-32k
AMC RULER 32k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 32,768 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-32k.amc-ruler-qwen35-16k
AMC RULER 16k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 16,384 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-16k.amc2023DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1difficulty-E2H-AMC-generations
Generations Dataset: E2H-AMC
LLM-generated solutions across train/validation/test splits for multiple models.
Columns
Column
Type
Description
problem
str
Problem statement
generated_solutions
list
Generated solutions with scores
success_rate
float
Fraction of correct generations
majority_vote_is_correct
int (0/1)
Whether majority vote is correct
k
int
Number of samples generated
temperature
float
Sampling temperature
max_len
int
Maximum… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeGitta/difficulty-E2H-AMC-generations.amc25amc2023_no_solfew_shot_amcmath_eval_suite-amcamc-ruler-qwen35-4k
AMC RULER 4k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 4,096 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-4k.AMC23_evalchemyamc-ruler-qwen35-128k
AMC RULER 128k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 131,072 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-128k.aime25_amc23_gpqa_olymp_cotamc23_dup32amcimbalanced_amc_datasetamc-ruler-qwen35-64k
AMC RULER 64k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 65,536 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-64k.amc23amc-ruler-qwen35-8k
AMC RULER 8k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 8,192 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-8k.amc_repeated_4xamc23-aR-HORIZON-AMC23amc0.1krecept
Dataset Card for Recept
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/amcoff/recept.amc23-enrichedDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-mistral
