datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tau2-bench-verified-airline
tau2-bench-verified — airline domain (mirror)
Mirror of the airline domain from
amazon-agi/tau2-bench-verified
(MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c.
Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch
the verified airline tasks + environment DB at runtime.
Contents
tasks/test.jsonl — 50 verified airline tasks. Each row has a single
task_json string column holding one verbatim tau2 v2 task object
(id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.ouroboros-osworld-verified-opus5
Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date
Status: Self-reported result over all 361 tasks. The official per-task
scores, prompts, manifests and feasibility records are public here, together
with every acting task record that the run produced.
Start here
Result
90.69% (327.39 / 361)
Model
anthropic/claude-opus-5
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence
f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.ouroboros-osworld-verified-sonnet46
Ouroboros on OSWorld-Verified: best published result on Claude Sonnet 4.6
Status: Self-reported result over all 361 task packages. Prompts,
manifests, outcomes and feasibility records are public for every task. The
official evaluator produced 360 score files; one unscored task is counted as zero.
Start here
Result
83.27% (300.59 / 361)
Model
anthropic/claude-sonnet-4.6
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-sonnet46.SciCode-Verified
SciCode-Verified
SciCode-Verified is the corrected, human-verified release of the
SciCode scientific-code-generation benchmark.
A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and
corrected every confirmable defect. The released evaluation set contains 64 main problems and
287 scored subproblems; one original problem is excluded because its specification does not
determine a unique, verifiable answer.
Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.TACO-verified
Introduction
This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
25443
1468722
TACO-verified
12898
1043251
Correct Ratio
50.69 %
71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.SWEBench-Pro-Verified
SWE-Bench Pro Verified: Anti-hacking & Task refinement
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging
repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of
unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and
task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.verified-defi-datasets
Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus
Dataset Description
High-density, verified AI fine-tuning dataset in ALPACA format.
Domain: Solana Sealevel & Anchor Program Optimization
Verified Records: 3
Estimated Tokens: 339
Quality QA Score: 99.0%
Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution
HLE-Verified
HLE-Verified (HF-native)
This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735.
Why this exists
The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset.
This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval.
Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.Doc2Feat-bench_Verified
Dataset Summary
NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically.
Languages
The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type.
Dataset Structure
An example of a SWE-bench datum is as follows:
repo: (str) - The repository owner/name identifier from GitHub.
instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.verified-sql-rewards
Verified SQL Rewards
A text-to-SQL corpus where every reward carries a machine-checkable proof
that it is correct.
Questions, all independently verified
109,306
Databases
1,400 across 7 schema families
Tables / data rows
4,400 / ~19.6 million
Unique (question, answer) pairs
102,764
Candidates refused and published
12,150
Verification pass rate
90.00%
Trivial baseline (always answer 0)
1.83%
Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.Lego-RL-SWE-Bench-Verified
Lego-RL-SWE-Bench-Verified
The 500 SWE-bench Verified instances as ready-to-run harbor RL environments —
the exact evaluation set behind every SWE-bench Verified number in
LEGO-RL, packaged the same way as the training
set Lego-X/Lego-RL-2699 so
one trainer reads both.
Two parallel views of the same 500 instances:
View
Path
What it is
Official SWE-bench records
swebench_verified_official_500/
The upstream princeton-nlp/SWE-bench_Verified rows, verbatim
Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.MAPS_Verified
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.VeriLoop-Structural-Repair-Verified
VLR-StructuralRepair v1.0.0 — non-regressive repair of real semantic defects
Evidence-convergent supervision for function-level semantic repair under a
hidden set of protected obligations. A candidate is positive only when it
preserves every already-satisfied obligation and strictly repairs at least
one. Aggregate improvement that breaks a protected obligation is a negative,
however far the total failure count drops.
The previous generation of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Structural-Repair-Verified.verified-math-olympiad-trajectories
Verified Math Olympiad Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's math olympiad reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), answer-verifier evaluation, process-supervision candidates, judge training, proof criticism, and private evaluations.
The goal is not merely to provide final-answer math examples. Each record is a structured olympiad reasoning object containing a normalized problem… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-math-olympiad-trajectories.VeriLoop-Governed-Recurrence-Verified
VLR-Recurrence-Verified
VLR-Recurrence-Verified is a synthetic-data construction release for studying
evidence-convergent program repair. It operationalizes a protected partial order:
a candidate is positive only when it preserves every already-satisfied
obligation and strictly improves at least one unresolved obligation.
Scale
Split
Tasks
Families
Transitions
Balanced pairs
Certified finals
Train
3,500
28
12,250
49,000
3,500
Validation
750
10
2,623… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Governed-Recurrence-Verified.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.verified_wiki_historian_the_beatles_anthology_dataset_active
Verified-Wiki-Historian: The Beatles Anthology
Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning.
This dataset is a cleaned and rebuilt refinement of:
Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active
The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.idfu-verified-code
IDFU Code Negative Dataset — Free Preview
A curated dataset of Python code samples that failed execution-based
validation, designed for training reward models, DPO rejected-side pairs,
and error-detection classifiers. Free 100-sample preview; paid full versions
available separately.
What's inside this preview
100 unique Python samples, all AST-validated
19 CS domains represented (MCMC, FFT, distributed consensus, ZKP,
formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.IMO-AnswerBench-Verified
IMO AnswerBench Verified
IMO AnswerBench Verified is a human-expert-verified derivative of OpenEvals/IMO-AnswerBench, originally curated by the Google DeepMind Superhuman Reasoning team. Every record in the 400-problem benchmark was reviewed individually. The review identified and corrected 13 records while preserving the benchmark's balanced coverage of four major mathematical areas.
Dataset summary
Total records: 400
Verification method: record-by-record human… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/IMO-AnswerBench-Verified.verified-tool-use-dataset
Verified tool-use trajectories for LLM agents
This was a time-boxed experiment by an autonomous agent (Protogonos), now concluded. Nothing here is offered for sale or for hire, and no payment is accepted.
Multi-turn function-calling conversations for training and evaluating
tool-using agents — 48 trajectories across 16 domains, with every tool call
checked against its tool's JSON-Schema. The free sample in this repo is a real
slice of the full set: the viewer above renders it… See the full description on the dataset page: https://huggingface.co/datasets/protogonos/verified-tool-use-dataset.regex-pattern-generation-with-verified-match-sets-cmskdvm7
Regex Pattern Generation with Verified Match Sets
A dataset of regex pattern generation with verified match sets examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Regex
Framework: Community
License: CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/regex-pattern-generation-with-verified-match-sets-cmskdvm7.chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.android-kotlin-compose-compiler-verified
Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset
5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs.
Every SFT row was actually compiled — not LLM-approved, not heuristically
filtered. A subset was verified behaviorally by running JUnit tests.
Built to fine-tune small models into focused Android specialists rather than
general-purpose coders.
Why this exists
Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.physics-verified
PHYSICS-Verified
PHYSICS-Verified is a benchmark of 1,109 PhD-qualifying-exam physics problems with 2,803 scored answers, covering six core areas of physics. Every problem asks for results that can be checked: numbers, formulas, or short verbal conclusions. Each answer has been checked against its reference solution.
This release is a cleaned, results-only version of the original PHYSICS benchmark (GitHub). Problems that required a proof, explanation, or drawing were removed, as… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/physics-verified.kodcode-verified-python-235k
KodCode-Verified Python — 234,555 execution-verified Python SFT rows
One row per problem. Every assistant turn is code that passed its own unit tests when
actually run — real pytest against KodCode-V1's
tests, in a pinned interpreter, in a sandboxed subprocess. No LLM judge, no heuristic
filter, no model-generated answers.
Unlike the v3 release this supersedes, the corpus is deduplicated, decontaminated against
HumanEval/MBPP, and stripped of rows whose tests cannot constrain… See the full description on the dataset page: https://huggingface.co/datasets/F-A-I-L/kodcode-verified-python-235k.UPBench-Error-verified-v2
UPBench-Error-verified-v2
LCZZZZ/UPBench-Error → generation_error 子集,经两轮人工核验后保留的
1,219 条样本。每条样本视觉上看不出明显的低级生成缺陷。
筛选过程
步骤
剩余
原始 generation_error 样本
5,761
剔除 is_gui=true(GUI-World / egoproactive 屏幕录制)
3,390
第一轮:逐条过目 error_clip.mp4
good 1,420 / bad 1,970
第二轮:对第一轮 good 再过一遍
good 1,219 / bad 201
第二轮刷掉了第一轮 14.2% 的样本,最终保留率 1,219 / 3,390 = 36.0%。
内容
metadata/manifest-verified.jsonl 1,219 条,原 manifest 全部 25 个字段逐字保留,… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified-v2.APPS-verified
Introduction
This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
5000
117232
TACO-verified
4211
93921
Correct Ratio
84.22%
80.12%
