datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webfaqWebFAQ Q&A Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.SWE-QA-Pro-Bench
SWE-QA-Pro Bench (A Repository-level QA Benchmark Built from Diverse Long-tail Repositories)
💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro
📢 News
🚀 [2026-5-19] The evaluation code is released on GitHub.
🔥 [2026-3-23] SWE-QA-Pro Bench is publicly released! The model and code will be released soon.
Introduction
SWE-QA-Pro Bench is a repository-level question answering dataset designed to evaluate whether models can perform grounded, agentic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-Bench.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.AIME25The AIME25 part 1 exam from the website.
LOOMBench
🔬 LOOMBench: Long-Context Language Model Evaluation Benchmark
🎯 Framework Overview
LOOMBench is a streamlined evaluation suite derived from our comprehensive long-context evaluation framework. It represents the gold standard for efficient long-context language model assessment.
✨ Key Highlights
📊 16 Diverse Benchmarks: Carefully curated from extensive benchmark collections.
⚡ Efficient Evaluation: Optimized for unified loading and evaluation.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/LOOMBench.PHYBench
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
[🌐 Project]
[📄 Paper]
[💻 Code]
[🏆 Leaderboard]
[🌟 Overview]
[🔧 Data Details]
[🚩 Citation]
New Updates
2025.4.25: We release our code of EED Score. View and star on our github page!
2025.5.15: We have significantly improved the paper and experiments, including diversified experimental discussions and in-depth error analysis. The updated website is now live at… See the full description on the dataset page: https://huggingface.co/datasets/Eureka-Lab/PHYBench.drug_label_approved_openfda
KEMIRIX OpenFDA Clinical Drug Dataset
Built for KEMIRIX — Africa's first Clinical Decision Support AI
Developer: Emmanuel Bain Oduwo | TechFryz Ltd. | Nairobi, Kenya
Generated: May 2026
Configurations
clean (default): instruction + output only, fully cleaned, ready for fine-tuning Kemirix
raw: full metadata schema, original generated data
Usage
from datasets import load_dataset
# Clean data for training Kemirix
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Oduwo/drug_label_approved_openfda.HealMed
HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems.
The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.Trust-Data
Dataset Card for Trust framework
Description
Repository: https://github.com/declare-lab/trust-align
Paper: https://arxiv.org/abs/2409.11242
Data Summary
The Trust-score evaluation dataset includes the top 100 GTR-retrieved results for ASQA, QAMPARI, and ExpertQA, along with the top 100 BM25-retrieved results for ELI5. The answerability of each question is assessed based on its accompanying documents.
The Trust-align training dataset comprises 19K high-quality… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/Trust-Data.rq-bench
RQ-Bench: A Benchmark for Grounded Research Question Generation
RQ-Bench evaluates whether language models can read background literature and propose the same kinds of research questions that a human author actually went on to investigate.
Each example pairs a held-out research question (RQ) — distilled from a real arXiv paper (the target paper) — with the full text of the prior-work papers that the target paper cites as motivation. A model is shown only the cited references and… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/rq-bench.dnd5e-srd-qa
D&D 5.2.1 SRD RAG Evaluation Dataset
A high-quality Question-Answering (QA) dataset built by the Datapizza AI Lab from the Dungeons & Dragons 5th Edition System Reference Document (SRD) version 5.2.1, designed to evaluate Retrieval Augmented Generation (RAG) systems.
Dataset Summary
This dataset contains 56 question-answer pairs across two difficulty tiers (Easy and Medium), each designed to test different aspects of RAG system capabilities. The dataset is built from 20… See the full description on the dataset page: https://huggingface.co/datasets/datapizza-ai-lab/dnd5e-srd-qa.swen-1-data
Swen-1 Dataset
Conversational and mathematical reasoning data collected by Sorika Labs from Swen Dual-Engine (Swen-1.1-Instruct & Swen-1-Math).
IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.EnvFactory-SFT-FILTERED
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
## Overview
EnvFactory-SFT-FILTERED is a filtered supervised fine-tuning (SFT) dataset containing 53,400 tool-use trajectories synthesized using the EnvFactory framework. This dataset is designed for SFT training of tool-use agents.
The dataset contains high-quality multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED.supervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.Fineweb-InstructWe convert the pre-training corpus from Fineweb-Edu (https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) to instruction following format. We select a subset with quality filter and then use GPT-4 to extract instruction-following pairs. The dataset contains roughly 16M instruction pairs. The basic concept is similar to MAmmoTH2 (https://arxiv.org/abs/2405.03548).
Citation
If you use dataset useful, please cite the following paper:
@article{yue2024mammoth2… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Fineweb-Instruct.SWE-QA-Pro-SFT-Trajectories
SWE-QA-Pro SFT Trajectories
💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro
Introduction
SWE-QA-Pro SFT Trajectories is a set of agentic tool-use trajectories for repository-level question answering, used as the supervised fine-tuning (SFT) data in the SWE-QA-Pro training recipe.
Each item is a multi-turn trajectory in which an agent answers a repository-grounded question by exploring the codebase with read-only tools rather than relying on memorized knowledge. The… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-SFT-Trajectories.HALT_Benchmark_0.1_v1
HALT Benchmark Dataset v1.0
HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse
Overview
HALT is a benchmark for evaluating bounded agentic decision-making under partial observability,
constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity
workflows, where acting too early, failing to escalate, or over-escalating can all be costly.
The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.EnvFactory-SFT-ALL
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
## Overview
EnvFactory-SFT-ALL is the complete supervised fine-tuning (SFT) dataset containing 26,500 tool-use trajectories synthesized using the EnvFactory framework. This dataset includes all generated trajectories before filtering.
The dataset contains multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL.tau2-bench-ko
τ²-bench 한국어 번역 공개판
τ²-bench v0.2.0의 retail·airline·telecom 사용자 시나리오를 한국어로 번역
원본: sierra-research/tau2-bench v0.2.0
번역 및 검수 모델: gpt-5.6-sol, 일부 모호한 문장은 사람이 검수
번역 범위: 278 task, 1,004 user-scenario field
파일
파일
행 수
설명
data/retail.jsonl
114
유통 에이전트 태스크
data/airline.jsonl
50
항공 에이전트 태스크
data/telecom.jsonl
114
통신 에이전트 태스크
metadata.json
-
생성 설정·파일 SHA-256
build_dataset.py
-
원본 task에 번역 필드를 결합해 jsonl 생성하는 스크립트
LICENSE
-
원본 τ²-bench의 MIT License 사본… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tau2-bench-ko.GSM8k_MOREDataset introduced in the paper: Evaluating LLMs' Mathematical Competency through Ontology-guided Perturbations.
This dataset was created by randomly sampling five questions from GSM8K and perturbing them using an ontology.
qa-dataset-k1000
QA Dataset K1000 — The First Drop of Ink
Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026).
Paper · Full text (v2) · Hugging Face paper page
The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.LabCraft-Eval
LabCraft-Eval
LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI
agents execute benign molecular-microbiology protocols inside a seeded
laboratory simulator with task-dependent stochasticity. It pairs task prompts
and tool-accessible lab operations with deterministic, multi-axis trajectory
scoring.
This Hugging Face dataset export is generated from the GitHub repository:
https://github.com/jang1563/LabCraft-Eval.git
Release
Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.AlgGeoTest
Welcome to AlgGeoTest created by PKU-DS-LAB!
Citation Information
Paper Link: https://arxiv.org/abs/2508.02208
Dataset Description
AlgGeoTest is the first benchmark specifically designed to evaluate LLMs' comprehension of Algebraic Geometry—a frontier domain of modern mathematics that occupies a central position within the contemporary mathematical landscape.
AlgGeoTest was created by implementing Proof2Hybrid—the first fully-automated framework for… See the full description on the dataset page: https://huggingface.co/datasets/PKU-DS-LAB/AlgGeoTest.claude-fable-derm
Claude Fable Derm
Claude Fable Derm is a dataset of patient dermatology questions and their raw, unprompted
responses from Claude Fable. Each question is generated by crossing a real dermatological
topic with one of Paul Ekman's six basic emotions, producing emotionally distinct framings
of the same underlying medical concern. Answers are collected with no system prompt or
role instruction, capturing how the model responds to a patient question exactly as it
would in the wild.… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/claude-fable-derm.LogiOR
Overview
LogiOR, a comprehensive benchmark dataset comprising 92 logistics and supply chain optimization problems, which was developed over two months under the guidance of three Operations Research (OR) experts. The problems are adapted from classical OR solver test datasets, textbook examples, research papers, and real-world applications. LogiOR covers a broad spectrum of optimization types including Linear Programming (LP), Integer Linear Programming (ILP), Mixed-Integer… See the full description on the dataset page: https://huggingface.co/datasets/LabMem012/LogiOR.FineBadmintonBenchmark
FineBadmintonBenchmark
Fine-grained badminton video question answering benchmark.
Dataset Structure
hf_video_clips_qa/: one clip per QA item, named by video_uid (for example video_000001.mp4).
finebadmintonbenchmark/: annotation JSON files.
each item contains video_uid
each QA item maps to exactly one video clip through video_uid
Citation
@inproceedings{he2025finebadminton,
title={Finebadminton: A multi-level dataset for fine-grained badminton video… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/FineBadmintonBenchmark.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.ATANTV1.0-corpus
ATANT Narrative Test Corpus
Automated Test for Acceptance of Narrative Truth, v1.0
The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time.
Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710)
Standard repository: github.com/Kenotic-Labs/ATANT
Author: Samuel Sameer Tanguturi
Affiliation: Kenotic Labs
Published: April 2026
Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.
