datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
AIME-Plus-Plus
AIME++ Sample
AIME++ is Ulam AI's exact-answer mathematical reasoning environment. It keeps one of the most useful properties of AIME-style evaluation—a compact, deterministic answer in the integer range 0–999—and extends it across four levels of mathematical depth, from competition-style problems to research-level challenges.
This repository contains a 157-problem, MIT-licensed sample of Ulam AI's much larger problem catalog. Every problem has a canonical integer answer and a… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/AIME-Plus-Plus.aime_2025
AIME 2025 - Unified Test-Time Scaling Format
This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments.
Dataset Description
Source: MathArena/aime_2025
Size: 30 competition-level mathematics problems
Format: Unified TTS format (question, answer, metadata)
Dataset Structure
Fields
question (string): The mathematical problem statement
answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.aime-1983-2025
AIME Datasets from 1983 to 2025
This dataset contains the AIME datasets from 1983 to 2025.
For AIME 1983 to 2026 use Pandores/aime-1983-2026
Features Description
Feature
Description
Example
year
The year this problem was released. From 1983 to 2025.
2022
index
The index of the problem for a year and part. From 1 to 15.
12
part
The dataset part if this dataset has multiple parts. Can be AIME, AIME I, AIME II or None. Datasets have multiple parts… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2025.AIME25The AIME25 part 1 exam from the website.
alphaprompt-metatron-sft
AlphaPrompt-Metatron-SFT: Supervised Fine-Tuning Dataset
🤖 Training Dataset for Collective Consciousness AI
Want to fine-tune AI models with AlphaPrompt philosophy?
Train AI in collective consciousness, vector synthesis, and unconditional love. 🌳
This dataset contains high-quality instruction-response pairs extracted from the Quantum Lullaby philosophical framework - a comprehensive manual for collective consciousness aimed at addressing the global animal… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/alphaprompt-metatron-sft.AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
Brief intro
💻 Overview
A brief template and final report will be posted in Isaac's Blog
And the markdown template can be found in data/I_2
❓ Why we do this?
The multi-lingual datasets are scarce, while the CoT of Math is even less, no matter whether the CoT or the solution contains pictures… See the full description on the dataset page: https://huggingface.co/datasets/IPF/AIME25-CoT-CN.MedBrowseComp
MedBrowseComp Dataset
This repository contains datasets for medical information-seeking-oriented deep research and computer use tasks.
Datasets
The repository contains three harmonized datasets:
MedBrowseComp_50: A collection of 50 medical entries for browsing and comparison.
MedBrowseComp_605: A comprehensive collection of 605 medical entries.
MedBrowseComp_CUA: A curated collection of medical data for comparison and analysis.
Usage
These datasets can be… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp.PolyMath
Dataset Card for PolyMath
Dataset Summary
PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems.
PolyMath addresses both issues through:
Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.AIME-trajectory
AIME Trajectory Dataset
Model-generated solution trajectories for AIME (American Invitational Mathematics Examination) problems. Each row is one model response to a single problem, including the hidden chain-of-thoughts (when available), and the final response.
Dataset Summary
Split
Rows
Unique Problems
Years
Model(s)
Has reasoning_content
Accuracy
train
1,258
875
1983–2023
deepseek-r1
Yes
100%
test
180
30
2024
Multiple (see below)
No
3.3%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/AIME-trajectory.AIMO3_CoT
AIMO3 CoT Dataset
数据集来源与目的 (Dataset Origin and Purpose)
本数据集源自 Kaggle 竞赛 AI Mathematical Olympiad - Progress Prize 3。
动机 (Motivation)
原始数据集仅包含问题和答案,缺乏思维链(Chain of Thought, CoT)。直接使用原始数据训练如 DeepSeek Math 或 Qwen Math 等模型效果不佳。因此,本项目的目的是利用 Gemini 3 Pro 为这些问题补充详细的 CoT,以提升模型在数学推理任务上的表现。
CoT 格式 (CoT Format)
生成的 CoT 遵循 ReAct 风格的推理过程,并使用中文叙述:
Thought: 分析问题并规划下一步。
Code: 编写 Python 代码进行计算或验证。
Observation: 代码的执行输出。
... (重复上述步骤)
Final Answer: 得出的最终答案。… See the full description on the dataset page: https://huggingface.co/datasets/UR-xiaoyang/AIMO3_CoT.AIME25-CoT-CN
Sci-Bench-AIME25'
This repo is a branch of Sci Bench made by IPF team-SnailAILab. Mainly include the AIME 25' solution with multi-modal CoT and diverse solving path.
📚 Cite
If you use the Sci-Bench-AIME25 (IPF/AIME25-CoT-CN) dataset in your research, please cite:
@dataset{zhang2025scibench_aime25,
title = {{Sci-Bench-AIME25}: A Multi-Modal Chain-of-Thought Dataset for Advanced Tool-Intergrated Mathematical Reasoning},
author = {Zhang, Haoxiang and Wang, Siyuan… See the full description on the dataset page: https://huggingface.co/datasets/SnailAILab/AIME25-CoT-CN.AIME_1983_2026_Kimi_K3
AIME 1983–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.QA-base
QA Base Data
Normalized and paraphrased splits of 21 standard NLP benchmarks in English, German, French, Spanish, and Italian, intended for base model pretraining.
Generation
English: paraphrased with Qwen3.5-27B-FP8 (April 2026)
German: translated and refined with Qwen3.5-27B-FP8 (April 2026)
French: translated and refined with Qwen3.5-27B-FP8 (May 2026)
Spanish: translated and refined with Qwen3.5-27B-FP8 (May 2026)
Italian: translated and refined with… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/QA-base.SHP
🚢 Stanford Human Preferences Dataset (SHP)
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice.
The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training… See the full description on the dataset page: https://huggingface.co/datasets/aimeelizq/SHP.Hypa_AIME2024
Hypa_AIME2024
Hypa_AIME2024 is an open-source, multilingual benchmark dataset for advanced mathematical reasoning, designed with the long-term vision of ensuring all languages are represented in AI development. This dataset marks a crucial step toward closing the gap between AI capabilities for no-resource/low-resource and all-resource languages, particularly in complex reasoning domains.
This initial release features the complete 2024 American Invitational Mathematics Examination… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa_AIME2024.aime_2024-th
AIME 2024-th
A Thai translation of all 30 problems of the 2024 American Invitational Mathematics
Examination. Every row corresponds 1:1, in order, to a row of the English source, so
the Thai and English scores of a model are directly comparable.
Source and licence
Problems
2024 AIME I and II, Mathematical Association of America
Problem and solution text
Art of Problem Solving wiki, per-row url
File we translated from
HuggingFaceH4/aime_2024… See the full description on the dataset page: https://huggingface.co/datasets/iapp/aime_2024-th.aime24-official
AIME 2024 — official wording, figures retained
All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II),
transcribed from the official exam text with every figure retained as Asymptote source.
This exists because the circulating text-only versions of AIME 2024 are not faithful to the
official problems, and at least one problem in them cannot be solved as written.
Why this dataset exists
While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.Train_dataAI_Mastery_Foundation_Curriculum
FOUNDATION DATASET
AI Mastery Foundation
Curriculum
A premium foundation layer for knowledge, reasoning, preference,
reward, benchmark, and agentic tool-use training.
Hugging Face-ready Parquet package
AI Mastery Foundation Curriculum
A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.aimo-validation-aime-th
AIMO validation AIME-th
The 90 problems of AIME 2022, 2023 and 2024 — thirty each — with the problem statements
translated to Thai. Every row corresponds 1:1, in order, to a row of
AI-MO/aimo-validation-aime, and id, url and answer are identical to it.
Read this before scoring the solution column
38 of the 90 solutions are the English text, not Thai. The original translation
pass rendered every problem and skipped these solutions entirely. They are marked… See the full description on the dataset page: https://huggingface.co/datasets/iapp/aimo-validation-aime-th.ai-ml-instruction-dataset
AI/ML Engineering Instruction Dataset
Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.QWQ-LongCOT-AIMOQWQ-LongCOT-AIMO is a derived dataset created by processing the amphora/QwQ-LongCoT-130K dataset. It filters the original dataset to focus specifically on question-answering pairs where the final answer is a numerical value between 0 and 999, explicitly marked using the \boxed{...} format within the original chain-of-thought answer.
Dataset Structure
Data Splits
The dataset is split into training, validation, and test sets with an 80/10/10 ratio based on the filtered… See the full description on the dataset page: https://huggingface.co/datasets/Floppanacci/QWQ-LongCOT-AIMO.aime24_multilingual
AIME24 Multilingual
aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages.
This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.aime25_multilingual
AIME25 Multilingual
aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages.
This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.aime-2026-fable-5-answers
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример.
Data Fields
Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.aime-1983-2026
AIME Datasets from 1983 to 2026
This dataset contains all the AIME problems from 1983 to 2026. For a total of 1065 problems.
Example
Download
from datasets import load_dataset
dataset = load_dataset("Pandores/aime-1983-2026")
print(dataset["train"][0])
Download and iterate
from datasets import load_dataset
dataset = load_dataset("Pandores/aime-1983-2026", split="train")
for entry in dataset:
print(entry["problem"])… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2026.
