datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tau2-bench-verified-airline
tau2-bench-verified — airline domain (mirror)
Mirror of the airline domain from
amazon-agi/tau2-bench-verified
(MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c.
Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch
the verified airline tasks + environment DB at runtime.
Contents
tasks/test.jsonl — 50 verified airline tasks. Each row has a single
task_json string column holding one verbatim tau2 v2 task object
(id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.ouroboros-osworld-verified-opus5
Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date
Status: Self-reported result over all 361 tasks. The official per-task
scores, prompts, manifests and feasibility records are public here, together
with every acting task record that the run produced.
Start here
Result
90.69% (327.39 / 361)
Model
anthropic/claude-opus-5
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence
f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.SciCode-Verified
SciCode-Verified
SciCode-Verified is the corrected, human-verified release of the
SciCode scientific-code-generation benchmark.
A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and
corrected every confirmable defect. The released evaluation set contains 64 main problems and
287 scored subproblems; one original problem is excluded because its specification does not
determine a unique, verifiable answer.
Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.ouroboros-osworld-verified-sonnet46
Ouroboros on OSWorld-Verified: best published result on Claude Sonnet 4.6
Status: Self-reported result over all 361 task packages. Prompts,
manifests, outcomes and feasibility records are public for every task. The
official evaluator produced 360 score files; one unscored task is counted as zero.
Start here
Result
83.27% (300.59 / 361)
Model
anthropic/claude-sonnet-4.6
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-sonnet46.SWEBench-Pro-Verified
SWE-Bench Pro Verified: Anti-hacking & Task refinement
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging
repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of
unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and
task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.TACO-verified
Introduction
This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
25443
1468722
TACO-verified
12898
1043251
Correct Ratio
50.69 %
71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.VeriWebVeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
[!NOTE]
This project was originally named VeriGUI. As our initial data collection focused on web-based tasks that primarily involve information-seeking rather than GUI interaction, we now define this part as the standalone VeriWeb benchmark, while desktop and other GUI-oriented scenarios will be released as a separate benchmark (in progress). We apologize for any resulting confusion.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/2077AIDataFoundation/VeriWeb.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.HLE-Verified
HLE-Verified (HF-native)
This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735.
Why this exists
The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset.
This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval.
Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.verified-defi-datasets
Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus
Dataset Description
High-density, verified AI fine-tuning dataset in ALPACA format.
Domain: Solana Sealevel & Anchor Program Optimization
Verified Records: 3
Estimated Tokens: 339
Quality QA Score: 99.0%
Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution
k-verify-benchmark-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
K-Verify Benchmark v1
Author: Yachay / SZL Holdings · Version: 1.0.0 · Items: 100
K-Verify measures whether an AI's claimed factual answer is verifiable via a
receipt chain — not just whether it is correct. It is the first benchmark we
know of that scores provenance and honest refusal as… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/k-verify-benchmark-v1.nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.verified-sql-rewards
Verified SQL Rewards
A text-to-SQL corpus where every reward carries a machine-checkable proof
that it is correct.
Questions, all independently verified
109,306
Databases
1,400 across 7 schema families
Tables / data rows
4,400 / ~19.6 million
Unique (question, answer) pairs
102,764
Candidates refused and published
12,150
Verification pass rate
90.00%
Trivial baseline (always answer 0)
1.83%
Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.gspc-verify
GSPC Verify
This dataset carries inputs and pointers for verifying published GSPC evidence. It is a reader surface, not a mill, ranking engine, or certificate. The live board remains the authority: GET https://councilof.ai/api/gspc.
Lid: 23 axes measured · 14 model fleets · 3 public leader scores · 9 fact runs · TIE is TIE · not a certificate.
Do not freeze a MEASURED/UNMEASURED table here. Re-read the live endpoints; fetch failure is UNCHECKABLE, never zero. The verifier… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-verify.verified_wiki_historian_the_beatles_anthology_dataset_active
Verified-Wiki-Historian: The Beatles Anthology
Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning.
This dataset is a cleaned and rebuilt refinement of:
Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active
The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.Lego-RL-SWE-Bench-Verified
Lego-RL-SWE-Bench-Verified
The 500 SWE-bench Verified instances as ready-to-run harbor RL environments —
the exact evaluation set behind every SWE-bench Verified number in
LEGO-RL, packaged the same way as the training
set Lego-X/Lego-RL-2699 so
one trainer reads both.
Two parallel views of the same 500 instances:
View
Path
What it is
Official SWE-bench records
swebench_verified_official_500/
The upstream princeton-nlp/SWE-bench_Verified rows, verbatim
Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.MAPS_Verified
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.LLMVerify-Verifier
LLMVerify-Verifier
Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995).
This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.Doc2Feat-bench_Verified
Dataset Summary
NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically.
Languages
The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type.
Dataset Structure
An example of a SWE-bench datum is as follows:
repo: (str) - The repository owner/name identifier from GitHub.
instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.VeriTime
VeriTime: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning
This is the dataset associated with our paper:
Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning
Jiahui Zhou, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Lin Li, Zhuomin Chen, Jian Lou, See-Kiong Ng
ICML 2026 | Paper
Dataset Construction Pipeline: TSRgen
TSRgen is an… See the full description on the dataset page: https://huggingface.co/datasets/HayleyZhou1113/VeriTime.VerifyBench
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
Yuchen Yan1,2,*,
Jin Jiang2,3,
Zhenbang Ren1,4,
Yijun Li1,
Xudong Cai1,
Yang Liu2,
Xin Xu5,
Mengdi Zhang2,
Jian Shao1,†,
Yongliang Shen1,†,
Jun Xiao1,
Yueting Zhuang1
1Zhejiang University
2Meituan Group
3Peking university
4University of Electronic Science and Technology of China
5The Hong Kong University of Science and Technology
ICLR 2026… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-REAL/VerifyBench.tr-rss-haber-akisi-verisi
TR-RSS Haber Akışı Verisi
TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır.
Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.MoatlessTools-Agent-Verifier-Train-Datasft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.physics-verified
PHYSICS-Verified
PHYSICS-Verified is a benchmark of 1,109 PhD-qualifying-exam physics problems with 2,803 scored answers, covering six core areas of physics. Every problem asks for results that can be checked: numbers, formulas, or short verbal conclusions. Each answer has been checked against its reference solution.
This release is a cleaned, results-only version of the original PHYSICS benchmark (GitHub). Problems that required a proof, explanation, or drawing were removed, as… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/physics-verified.
