datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.pretraining_v1-omega_bookspeacock-data-public-datasets-idcNuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.iDATA
Dataset
A dataset of AI + EDA
iDATA is a dataset of AI + EDA, which can be used to train AI models for design PPA prediction, PPA-aware physical design, and related tasks.
Dataset structure
Describe the dataset structure.
aes/
├── iEDA_route_process_data/ # Process data exported by iEDA-iRT 2D routing
├── syn_netlist/ # The synthesized netlist files、sdc files
├── place/ # The place stage def、sdc、vectors
└── route/ # The route… See the full description on the dataset page: https://huggingface.co/datasets/AiEDA/iDATA.common_voice_17_0olympiads
AI-MO Olympiad Reference Dataset
This dataset contains a structured collection of Olympiad problems and their solutions,
organized by competition. Contains high quality data, prioritizing "official" solutions to problems.
Structure
<competition name>/ # Problems and solutions from the International Mathematical Olympiad
├── raw/ # Raw problem/solution statements (.pdf)
│ ├── file1.pdf
│ ├── file2.pdf
├── download_script/ # the scripts used to… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/olympiads.cad-1000-hours
CAD-1K Open v2 - 1,018.1229 Hours
509 end-to-end, single-display Windows CAD task recordings across seven CAD software families.
Each task contains:
task_desc.json - task prompt, application, reference-input paths, and expected deliverables
input_files/ - reference inputs named input.ext or input_N.ext
output_files/ - submitted CAD deliverables and supplemental outputs named output.ext or output_N.ext
rubrics.json - task-specific evaluation criteria
task_overview.pdf - review… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-1000-hours.AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.AIME_2024
AIME 2024 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2024. AIME is a prestigious high school mathematics competition known for its challenging mathematical problems.
Dataset Details
Format: JSONL
Size: 30 records
Source: AIME 2024 I & II
Language: English
Data Fields
Each record contains the following fields:
ID: Problem identifier (e.g., "2024-I-1" represents Problem 1… See the full description on the dataset page: https://huggingface.co/datasets/Maxwell-Jia/AIME_2024.xperience-10m
⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified.
Interactive Intelligence from Human Xperience
Xperience-10M
Dataset Summary
Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.aime25
AIME 25
American Invitational Mathematics Examination (AIME) 2025
Citation
If you use the AIME25 dataset in your research, please consider citing it as follows:
@misc{aime25,
title={American Invitational Mathematics Examination (AIME) 2025},
author={Zhang, Yifan and Math-AI, Team},
year={2025},
}
PostTrainBench-Trajectories
PostTrainBench Agent Traces
Agent traces from PostTrainBench (GitHub), a benchmark that measures CLI agents' ability to post-train base LLMs.
Task
Each agent is given:
A pre-trained base LLM to fine-tune
An evaluation script for a specific benchmark
10 hours on an NVIDIA H100 80GB GPU
The agent must autonomously improve the model's performance on the target benchmark using any post-training strategy it chooses (SFT, LoRA, RLHF, prompt engineering for data… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/PostTrainBench-Trajectories.L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school
90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany
6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU
Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals
Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.RobotDesign1M
RobotDesign1M: A Large-scale Dataset for Robot Design Understanding
RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval.
📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.pretraining_v1-omegacovost2This is a partial copy of CoVoST2 dataset.
The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer.
The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger.
As such, not all the data is included: Only the validation and test subsets are available.
From the XX_EN subsets, only fr, es, and zh-CN are included.
aime_2024
Dataset card for AIME 2024
This dataset consists of 30 problems from the 2024 AIME I and AIME II tests. The original source is AI-MO/aimo-validation-aime, which contains a larger set of 90 problems from AIME 2022-2024.
ai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.NuminaMath-1.5
Dataset Card for NuminaMath 1.5
Dataset Summary
This is the second iteration of the popular NuminaMath dataset, bringing high quality post-training data for approximately 900k competition-level math problems. Each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-1.5.HLE-Verified
HLE-Verified
A Systematic Verification and Structured Revision of Humanity’s Last Exam
Overview
Humanity’s Last Exam (HLE) is a high-difficulty, multi-domain benchmark designed to evaluate advanced reasoning capabilities across diverse scientific and technical domains.
Following its public release, members of the open-source community raised concerns regarding the reliability of certain items. Community discussions and informal replication attempts suggested that some… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/HLE-Verified.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use-large.fineweb-edu-fortified
Fineweb-Edu-Fortified
The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in
Airtrain
What is it?
Fineweb-Edu-Fortified is a dataset derived from
Fineweb-Edu by applying exact-match
deduplication across the whole dataset and producing an embedding for each row. The number of times
the text from each row appears is also included as a count column. The embeddings were produced
using TaylorAI/bge-micro
Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.sangraha
Sangraha
Sangraha is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations.
More information:
For detailed information on the curation and cleaning process of Sangraha, please checkout our paper on Arxiv;
Check out the scraping and cleaning pipelines used to curate Sangraha on GitHub;
Getting Started
For… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/sangraha.leaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.aimo-validation-aime
Dataset Card for AIMO Validation AIME
All 90 problems come from AIME 22, AIME 23, and AIME 24, and have been extracted directly from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-aime.aime_2025
AIME 2025
This dataset contains 30 problems from the 2025 AIME tests, including:
AIME I: 15 problems
AIME II: 15 problems
airlens-live
AirLens Live Data
Live data layer for AirLens, an open
air-quality monitoring platform. Updated by scheduled GitHub Actions pipelines.
Layout mirrors the former Supabase Storage buckets:
Path
Content
Cadence
aq-data/current-*-grid.json
Global pollutant grids (PM2.5/PM10/O3/NO2/CO)
hourly
aq-data/timeline/
GEFS-Aerosols PM2.5 frames, -24h..+24h, 3h step
every 3h
aq-data/predictions/grid_latest.json
AOD→PM2.5 model predictions (p10-p90 + DQSS)
every 3h… See the full description on the dataset page: https://huggingface.co/datasets/Robeedau/airlens-live.aime_2026
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME 2026 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (int64): Gold final answer.
problem (string): Problem statement, usually stored as LaTeX source.
Source… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2026.
