datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime25
AIME 25
American Invitational Mathematics Examination (AIME) 2025
Citation
If you use the AIME25 dataset in your research, please consider citing it as follows:
@misc{aime25,
title={American Invitational Mathematics Examination (AIME) 2025},
author={Zhang, Yifan and Math-AI, Team},
year={2025},
}
aime26
AIME 26
American Invitational Mathematics Examination (AIME) 2026
Citation
If you use the AIME26 dataset in your research, please consider citing it as follows:
@misc{aime26,
title={American Invitational Mathematics Examination (AIME) 2026},
author={Zhang, Yifan and Math-AI, Team},
year={2026},
}
AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
aime24_nofiguresThe 30 problems from AIME 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded.
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_nofigures.PhD
[CVPR2025 Highlight] PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
preprint
🔥 PhD-webdataset
To enhance usability and integration with evaluation frameworks like lmm-eval, we are pleased to offer a packaged version in webdataset format. This packaged version is designed to facilitate easier deployment and testing. For further details and access, please refer to our repository PhD-webdataset.
Please note that the data in both repositories is completely… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD.aime-2024AIME-Plus-Plus
AIME++ Sample
AIME++ is Ulam AI's exact-answer mathematical reasoning environment. It keeps one of the most useful properties of AIME-style evaluation—a compact, deterministic answer in the integer range 0–999—and extends it across four levels of mathematical depth, from competition-style problems to research-level challenges.
This repository contains a 157-problem, MIT-licensed sample of Ulam AI's much larger problem catalog. Every problem has a canonical integer answer and a… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/AIME-Plus-Plus.AIME_Deepseek_Cleanamc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
AIME_1983_2024prompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedAIME25The AIME25 part 1 exam from the website.
AIME-trajectory
AIME Trajectory Dataset
Model-generated solution trajectories for AIME (American Invitational Mathematics Examination) problems. Each row is one model response to a single problem, including the hidden chain-of-thoughts (when available), and the final response.
Dataset Summary
Split
Rows
Unique Problems
Years
Model(s)
Has reasoning_content
Accuracy
train
1,258
875
1983–2023
deepseek-r1
Yes
100%
test
180
30
2024
Multiple (see below)
No
3.3%… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/AIME-trajectory.aime25_nofiguresprompt-swap-medium12-e2-mxfp4-mergedaime_nofiguresThe 90 problems from AIME 2022, 2023, 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded.
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime_nofigures.AIME-COD
Dataset Summary
AIME-COD is a synthetic dataset created to solve problems from the American Invitational Mathematics Examination (AIME) using chain of draft reasoning, as proposed in the Chain of Draft: Thinking Faster by Writing Less paper.
The dataset was generated using Curator and synthetic reasoning produced by Gemini 2.0 Flash. Problems are sourced from the gneubig/aime-1983-2024 dataset.
Dataset Details
Purpose and Scope
The AIME-COD… See the full description on the dataset page: https://huggingface.co/datasets/KingNish/AIME-COD.aime-2025AIME25aimeeARC-Bench
ARC-Bench: An Open-Ended Autonomous-Research Benchmark Across Five Scientific Domains
The benchmark released with AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration.
ARC-Bench is a 55-topic, open-ended autonomous-research benchmark spanning
five scientific domains. Each topic is not a fixed-input/fixed-output task — it is a
research question plus a structured briefing. A research agent (or a human) must
take a topic from question →… See the full description on the dataset page: https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench.ai-models-2026
AI Models & Releases 2026
AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.otis-mock-aime-24-25Problems from the 2024 and 2025 editions of the OTIS Mock AIME exam.
The problems were written by students from the Olympiad Training for Individual Study (OTIS) program.
The dataset contains 45 problems from 3 exams:
15 problems from Mock AIME 2024
15 problems from Mock AIME 2025 (I)
15 problems from Mock AIME 2025 (II)
amc_aime_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
AIME-22-25AIME22-24 + AIME25-PartI + AIME25-PartII
aime24_figuresThe 30 problems from AIME 2024 with all ASY code for figures.
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto},
year={2025},
eprint={2501.19393},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_figures.HeisenVec
HeisenVec: A Large-Scale Dataset for Text-to-SVG Generation
Overview
HeisenVec is a large-scale dataset designed to advance research in vector graphics generation from natural language descriptions. Unlike traditional image generation datasets focused on raster images, HeisenVec targets the structured domain of Scalable Vector Graphics (SVG). The dataset includes 2.2 million SVGs, each paired with four complementary textual descriptions generated by multimodal models.… See the full description on the dataset page: https://huggingface.co/datasets/aimagelab/HeisenVec.ai-military-2026
AI Military & Government 2026
AI in defense, government contracts, policy. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c =… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-military-2026.aimodel-sft-v1
