datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIME-Plus-Plus
AIME++ Sample
AIME++ is Ulam AI's exact-answer mathematical reasoning environment. It keeps one of the most useful properties of AIME-style evaluation—a compact, deterministic answer in the integer range 0–999—and extends it across four levels of mathematical depth, from competition-style problems to research-level challenges.
This repository contains a 157-problem, MIT-licensed sample of Ulam AI's much larger problem catalog. Every problem has a canonical integer answer and a… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/AIME-Plus-Plus.amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
AIME-trajectory
AIME Trajectory Dataset
Model-generated solution trajectories for AIME (American Invitational Mathematics Examination) problems. Each row is one model response to a single problem, including the hidden chain-of-thoughts (when available), and the final response.
Dataset Summary
Split
Rows
Unique Problems
Years
Model(s)
Has reasoning_content
Accuracy
train
1,258
875
1983–2023
deepseek-r1
Yes
100%
test
180
30
2024
Multiple (see below)
No
3.3%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/AIME-trajectory.AIME25amc_aime_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
aimo3-math-dataset
AIMO3 Math Dataset
Training data for AI Mathematical Olympiad Progress Prize 3.
Files
train_cot.jsonl - Chain-of-Thought examples
train_tir.jsonl - Tool-Integrated Reasoning examples
Author
Ryan J Cardwell (Archer Phoenix) - AIMO3 Competitor
prompt-slimmer-slm
Prompt Slimmer SLM — Demo Dataset
Synthetic examples for experimenting with prompt rewriting and sentence selection. Exported without changing the examples or their original splits from the shared GitHub codebase.
Model · Project page
Configuration
Train
Validation
Test
Purpose
rewrites-expanded (default)
41
2
2
Expanded rewriting dataset: 45 examples
rewrites
9
2
2
Original dataset used by the first adapter
selector
256
64
64
KEEP/DROP labels for source spans… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-slimmer-slm.aime2026-en
AIME 2026 · English — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in English, for evaluating whether a model can
reason in English (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-zh · aime2026-es · aime2026-fr · aime2026-ar · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-en.AI-MO-NuminaMath-TIR-korean-240918
IMPORTANT NOTE
This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST)
I'm taking a short break due to personal reasons. I'll be back in a month.
TODO-LIST
Finish translation
Translation
I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end.
Dataset Card for NuminaMath CoT
Dataset Summary
Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.repremover-xl
repremover-xl
7,281 anti-repetition roleplay conversations. Each row is a real RP
conversation containing an assistant turn that largely repeated an earlier
turn, cut at that point, with the repeating turn replaced by a rewritten
continuation. The intent: supervision located at the decision point where
models tend to loop.
Built for spoomplesmaxx-thrasher-24B
to reconstruct the idea behind the gated Dans-Prosemaxx-RepRemover-1. The
full generation script (gen_repremover_xl.py)… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/repremover-xl.aime24-official
AIME 2024 — official wording, figures retained
All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II),
transcribed from the official exam text with every figure retained as Asymptote source.
This exists because the circulating text-only versions of AIME 2024 are not faithful to the
official problems, and at least one problem in them cannot be solved as written.
Why this dataset exists
While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.aime2026-ar
AIME 2026 · Arabic — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in Arabic, for evaluating whether a model can
reason in Arabic (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-en · aime2026-zh · aime2026-es · aime2026-fr · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-ar.ai-ml-instruction-dataset
AI/ML Engineering Instruction Dataset
Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.aime2026-zh
AIME 2026 · Chinese — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in Chinese, for evaluating whether a model can
reason in Chinese (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-en · aime2026-es · aime2026-fr · aime2026-ar · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-zh.aime2026-es
AIME 2026 · Spanish — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in Spanish, for evaluating whether a model can
reason in Spanish (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-en · aime2026-zh · aime2026-fr · aime2026-ar · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-es.aime2026-fr
AIME 2026 · French — parallel multilingual math benchmark
The 2026 AIME competition (30 problems) in French, for evaluating whether a model can
reason in French (not pivot to English) and still solve competition math. Each item forces
target-language reasoning and carries a rule-based numeric ground-truth answer. One of six parallel
languages (en/zh/es/fr/ar/ru); companion sets: aime2026-en · aime2026-zh · aime2026-es · aime2026-ar · aime2026-ru.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/96kevinli29/aime2026-fr.dadbench
DadBench
A benchmark for dad-joke quality across language models — and a case study in how hard it is to measure a subjective objective without getting Goodharted.
DadBench scores model-written dad jokes on one question: would a real dad be proud to inflict this? It exists because the DadBench project RL-trained a model (PunTune, a fine-tune of Thinking Machines' Inkling) to tell dad jokes, and needed an evaluation rigorous enough to tell whether training actually worked. The… See the full description on the dataset page: https://huggingface.co/datasets/aimatey/dadbench.AIME_2024_to_Hard
AIME 2024 Augmented Dataset
Dataset Description
This dataset is an augmented version of the AIME 2024 dataset. It contains problems from the American Invitational Mathematics Examination (AIME) 2024, enriched with harder versions of the original problems and detailed Chain-of-Thought (CoT) solutions. This dataset is specifically designed to test and improve the mathematical reasoning capabilities of Large Language Models.
Dataset Details
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/UR-xiaoyang/AIME_2024_to_Hard.AIME-2025-prompt-only
AIME-2025-prompt-only
Prompt-only eval extraction from MathArena/aime_2025.
Generic Doubleword requests: 30
Request model placeholder: [MODEL]
Latin-OCR-Artifacts
Latin sentences sourced from The Latin Library, converted to images, were subsequently degraded via OCRODEG.
OCR (via Kraken and Tesseract) transcriptions were generated. The dataset was then augmented with several synthetic noise patterns in order to emulate the more severe corruption found in many older digitizations.
If you use this in your work, please cite:
@misc{mccarthy2025LACOROCR,
author = {McCarthy, A. M.},
title = {{Latin OCR Artifacts}},
year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/aimgo/Latin-OCR-Artifacts.aima-tasks-v1
aima-tasks-v1
30 classical-AI tasks for the
aima-env RL
environment, on the topics of Artificial Intelligence: A Modern Approach (Russell & Norvig):
uninformed and informed search, adversarial search with alpha-beta pruning, constraint
satisfaction, exact probability, and Markov decision processes.
field
meaning
task_id
ai-000 … ai-029
category
search / adversarial / csp / probability / mdp
prompt
the question, the tie-breaking rule, and the exact shape of the… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/aima-tasks-v1.aimperum_kaappiyangal-kundalakesi
📙 Kundalakesi Dataset (குண்டலகேசி தரவுத்தொகுப்பு)
🧾 Dataset Summary
Kundalakesi (குண்டலகேசி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics).
This epic is distinct in its strong focus on religious debate, renunciation, and philosophical transformation. The work survives only in a fragmentary form, with a limited number of verses available today.
This dataset presents a structured digital collection of all the extant verses of Kundalakesi, preserved in their… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-kundalakesi.moh_8_fake_rollouts
MOH-8 Fake Rollouts
480 math competition problems, each with 8 candidate solution rollouts from OSS 120B.
A controlled number of rollouts per problem are correct — use this to train/test a
verifier model that must identify which solutions are right.
Source
Problems and rollouts sampled from aimosprite/training-data-oss120b (the oss128-fixed-FINAL.jsonl file).
Only polymath-source problems in the 2/8–4/8 pass rate range (32–64 correct out of 128 attempts).
4 problems… See the full description on the dataset page: https://huggingface.co/datasets/aimosprite/moh_8_fake_rollouts.ko-o3-mini-high-aime-2022_4openai의 o3-mini high 를 이용하여 생성하였습니다.
참고하시고 사용바랍니다.
AIME24AIME25-RL2KletterMix
Dataset Card
KletterMix is a large German-language text dataset released as sharded JSONL files, introduced in the paper KletterMix: Climbing Toward High-Quality German Pretraining Data. Each row contains the text itself, a stable row identifier, a cluster assignment, a GPT-2 token count, and a proxy score.
This full release combines the deduplicated KletterMix data with the remaining scored KletterMix examples. It supersedes the smaller KletterMix-12B review-time subset.… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/KletterMix.KletterMix-12B-0.60
KletterMix-12B-0.60
KletterMix-12B-0.60 is the quality-filtered 0.60 release variant of KletterMix-12B, the German pretraining corpus accompanying KletterMix: Climbing Toward High-Quality German Pretraining Data.
The dataset contains German-language text examples selected with a target-language proxy score threshold of proxy_score >= 0.60. It follows the same public schema as KletterMix/KletterMix-12B and is intended for language-model pretraining, annealing experiments, data… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/KletterMix-12B-0.60.ClimbMix-split
ClimbMix Split
climbmix-split reorganizes the detokenized NVIDIA ClimbMix source corpus into
source-oriented splits. ClimbMix is described as being built from Nemotron-CC
and SmolLM-Corpus. Because SmolLM-Corpus is the smaller and directly
identifiable component, we used exact normalized-text matching against
SmolLM-Corpus to recover the SmolLM-derived portions. The remaining rows are
provided as the residual nemotron-cc split.
The data rows are unchanged from the detokenized… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/ClimbMix-split.
