datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.MathVista
Dataset Card for MathVista
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Data Visualization
Data Source
Automatic Evaluation
License
Citation
Dataset Description
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of three newly created datasets, IQTest, FunctionQA, and PaperQA, which address the missing visual domains and are tailored to evaluate logical… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVista.AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
AutoMathText🎉 This work, introducing the AutoMathText dataset and the AutoDS method, has been accepted to The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Findings)! 🎉
AutoMathText
AutoMathText is an extensive and carefully curated dataset encompassing around 200 GB of mathematical texts. It's a compilation sourced from a diverse range of platforms including various websites, arXiv, and GitHub (OpenWebMath, RedPajama, Algebraic Stack). This rich repository… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText.AirQA
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
This repository contains the test set, the metadata, processed_data and papers for the AirQA dataset introduced in our paper AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation accepted to ICLR 2026. Detailed instructions for using the dataset will soon be publicly available in our official repository.
AirQA is a human-annotated multi-modal multitask Artificial Intelligence… See the full description on the dataset page: https://huggingface.co/datasets/OpenDFM/AirQA.MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/ai-hyz/MemoryAgentBench.damru-knowledge
🐕 Damru Knowledge
A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students.
The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/.
📦 What's inside
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.officeqa
OfficeQA manifest (nearai-bench packaging)
Harness-ready question manifest for
databricks/officeqa — document-grounded
QA over U.S. Treasury Bulletins (1939–2025). 246 items in full,
8 in smoke (a 4-easy/4-hard subset for pipeline checks).
from datasets import load_dataset
ds = load_dataset("NEAR-AI/officeqa", split="full")
⚠️ This is the manifest only — documents are NOT included
Unlike our pinchbench and
clawbench exports, the
source corpus is not bundled here.… See the full description on the dataset page: https://huggingface.co/datasets/NEAR-AI/officeqa.TemplateGSM
TemplateMath: Template-based Data Generation (TDG)
This is the official repository for the paper "Training and Evaluating Language Models with Template-based Data Generation", published at the ICLR 2025 DATA-FM Workshop.
Our work introduces Template-based Data Generation (TDG), a scalable paradigm to address the critical data bottleneck in training LLMs for complex reasoning tasks. We use TDG to create TemplateGSM, a massive dataset designed to unlock the next level of… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/TemplateGSM.worldcup2026
⚽ WorldCup Arena
A Leakage-Free Forecasting Benchmark on a Live Tournament
Can a language model forecast a match — when the match had not been played at the moment it was asked?
🌐 Language / 语言 : 中文 ▾
📊 四张表
点开本页顶部的 Data Studio 标签即可浏览,也可以直接按名字加载。
Config
行数
内容
fixtures
104
基准本体 —— 喂给模型的头部信息,以及结算后的 90 分钟赛果,七个盘口全部推导好(outcome_1x2、over_2_5、both_score、odd_total)
dossiers
2,208
简报索引 —— 46 快照 × 48… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/worldcup2026.pre-flight-06
Aviation Operations Knowledge LLM Benchmark Dataset
This dataset contains multiple-choice questions designed to evaluate Large Language Models' (LLMs) knowledge of aviation operations, regulations, and technical concepts. It serves as a specialized benchmark for assessing aviation domain expertise.
📄 Paper: Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge (Brooker and Hughes, 2026). The benchmark is runnable via inspect_evals as the… See the full description on the dataset page: https://huggingface.co/datasets/AirsideLabs/pre-flight-06.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.Futurex-Past
FutureX-Past
📜 Overview
This repository contains a dataset of past questions from the FutureX benchmark.
FutureX is a live, dynamic benchmark designed to evaluate the future prediction capabilities of Large Language Model (LLM) agents. It features a fully automated pipeline that generates new questions about upcoming real-world events, deploys agents to predict their outcomes, and scores the results automatically. For more information on the live benchmark… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Past.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.MathVerse
Dataset Card for MathVerse
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Citation
Dataset Description
The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solvingremain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially assist MLLMs in deducing answers without truly interpreting the input diagrams.
To… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVerse.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.NLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.Futurex-Online
Submission Guidelines — Weekly Prediction Challenge
📄 Technical Report: https://arxiv.org/pdf/2508.11987
🌐 LeaderBoard: https://futurex-ai.github.io/
We run a weekly real-time prediction challenge. This repo always contains latest events.
1. Weekly Rules
A new set of tasks is released every week.
This week's tasks cover events with an end time between 2026-09-23, 24:00 (UTC+8) and 2026-09-29, 24:00 (UTC+8).
You must download the tasks, make predictions, and… See the full description on the dataset page: https://huggingface.co/datasets/futurex-ai/Futurex-Online.StackMathQA
StackMathQA
StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange
StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research.
Configs
configs:
- config_name: stackmathqa1600k… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/StackMathQA.AIME-Plus-Plus
AIME++ Sample
AIME++ is Ulam AI's exact-answer mathematical reasoning environment. It keeps one of the most useful properties of AIME-style evaluation—a compact, deterministic answer in the integer range 0–999—and extends it across four levels of mathematical depth, from competition-style problems to research-level challenges.
This repository contains a 157-problem, MIT-licensed sample of Ulam AI's much larger problem catalog. Every problem has a canonical integer answer and a… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/AIME-Plus-Plus.aime_2025
AIME 2025 - Unified Test-Time Scaling Format
This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments.
Dataset Description
Source: MathArena/aime_2025
Size: 30 competition-level mathematics problems
Format: Unified TTS format (question, answer, metadata)
Dataset Structure
Fields
question (string): The mathematical problem statement
answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.MILU
MILU: A Multi-task Indic Language Understanding Benchmark
Overview
MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages. It spans 8 domains and 41 subjects, reflecting both general and culturally specific knowledge from India.
Key Features
11 Indian Languages: Bengali, Gujarati, Hindi, Kannada, Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MILU.AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.CodeMMLU
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding Capabilities
📌 CodeMMLU
CodeMMLU is a comprehensive benchmark designed to evaluate the capabilities of large language models (LLMs) in coding and software knowledge.
It builds upon the structure of multiple-choice question answering (MCQA) to cover a wide range of programming tasks and domains, including code generation, defect detection, software engineering principles, and much more.
📄… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/CodeMMLU.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.SuperLim\finqa-verified
FinQA: Financial Question Answering Dataset
Description
The FinQA dataset is designed to facilitate research and development in the area of question answering (QA) using financial texts.
It consists of a subset of QA pairs from a larger dataset, originally created through a collaboration between researchers from the University of Pennsylvania,
J.P. Morgan, and Amazon.The original dataset includes 8,281 QA pairs built against publicly available earnings reports of S&P… See the full description on the dataset page: https://huggingface.co/datasets/Aiera/finqa-verified.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.
