datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime25
AIME 25
American Invitational Mathematics Examination (AIME) 2025
Citation
If you use the AIME25 dataset in your research, please consider citing it as follows:
@misc{aime25,
title={American Invitational Mathematics Examination (AIME) 2025},
author={Zhang, Yifan and Math-AI, Team},
year={2025},
}
aime26
AIME 26
American Invitational Mathematics Examination (AIME) 2026
Citation
If you use the AIME26 dataset in your research, please consider citing it as follows:
@misc{aime26,
title={American Invitational Mathematics Examination (AIME) 2026},
author={Zhang, Yifan and Math-AI, Team},
year={2026},
}
cornstack-python-v1
CoRNStack Python Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-python-v1.minervamathtelegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.AiAppAiice
Dataset
Aiice benchmark dataset for Arctic sea ice concentration (SIC) forecasting,
based on OSI-SAF satellite products (CC BY 4.0).
Coverage
Period: October 1978 – April 2026
Resolution: 25 km spatial, daily temporal
Grid: 432×432 (Lambert Azimuthal Equal Area, EPSG:6931)
Source products
Product
Source
Period
OSI-450-a
SMMR, SSM/I, SSMIS
1978–2020
OSI-430-a
SSMIS
2021–Jul 2025
OSI-438
AMSR2
Jul 2025–present… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/Aiice.aime24_nofiguresThe 30 problems from AIME 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded.
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_nofigures.AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.mosaic-combine-all
Mosaic format for combine all dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all
load it,
from streaming import LocalDataset
import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.Evo-BenchEvo-Bench: Can Language Models Improve Agent Harness?
A benchmark for measuring the intrinsic harness-evolving capability of language models.
Overview of the Evo-Bench evaluation pipeline.
✨ Highlights
608 harness-sensitive tasks from five established benchmarks, spanning
Search, Office, and General agent domains with disjoint validation and
evaluation suites.
Harness-guided benchmark construction selects tasks that respond to
harness improvements… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench.SCPWiki-Cleaned-PDF-ArchivesAgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko1,†,*, Alexandra Souly2,*
Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,*
Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,*
1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor
†EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford
Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.ai-arxiv2-chunksIndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.AIGVDBenchhplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.pre-flight-06
Aviation Operations Knowledge LLM Benchmark Dataset
This dataset contains multiple-choice questions designed to evaluate Large Language Models' (LLMs) knowledge of aviation operations, regulations, and technical concepts. It serves as a specialized benchmark for assessing aviation domain expertise.
📄 Paper: Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge (Brooker and Hughes, 2026). The benchmark is runnable via inspect_evals as the… See the full description on the dataset page: https://huggingface.co/datasets/AirsideLabs/pre-flight-06.tau2-bench-verified-airline
tau2-bench-verified — airline domain (mirror)
Mirror of the airline domain from
amazon-agi/tau2-bench-verified
(MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c.
Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch
the verified airline tasks + environment DB at runtime.
Contents
tasks/test.jsonl — 50 verified airline tasks. Each row has a single
task_json string column holding one verbatim tau2 v2 task object
(id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.mosaic-starcoder-filtered
Mosaic format for filtered starcoder dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.MERA
WARNING! This is the deprecated version. The new MERA datasets are now here!
MERA (Multimodal Evaluation for Russian-language Architectures)
Summary
MERA (Multimodal Evaluation for Russian-language Architectures) is a new open benchmark for the Russian language for evaluating fundamental models.
MERA benchmark brings together all industry and academic players in one place to study the capabilities of fundamental models, draw attention to AI problems, develop… See the full description on the dataset page: https://huggingface.co/datasets/ai-forever/MERA.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.enemThe ENEM 2022, 2023 and 2024 datasets encompass all multiple-choice questions from the last two editions of the Exame Nacional do Ensino Médio (ENEM), the main standardized entrance examination adopted by Brazilian universities. The datasets have been created to allow the evaluation of both textual-only and textual-visual language models. To evaluate textual-only models, we incorporated into the datasets the textual descriptions of the images that appear in the questions' statements from the… See the full description on the dataset page: https://huggingface.co/datasets/maritaca-ai/enem.C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.PhD
[CVPR2025 Highlight] PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
preprint
🔥 PhD-webdataset
To enhance usability and integration with evaluation frameworks like lmm-eval, we are pleased to offer a packaged version in webdataset format. This packaged version is designed to facilitate easier deployment and testing. For further details and access, please refer to our repository PhD-webdataset.
Please note that the data in both repositories is completely… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD.
