datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWEBench-Pro-Verified
SWE-Bench Pro Verified: Anti-hacking & Task refinement
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging
repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of
unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and
task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.LLMVerify-Verifier
LLMVerify-Verifier
Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995).
This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.tr-rss-haber-akisi-verisi
TR-RSS Haber Akışı Verisi
TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır.
Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.verified-math-olympiad-trajectories
Verified Math Olympiad Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's math olympiad reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), answer-verifier evaluation, process-supervision candidates, judge training, proof criticism, and private evaluations.
The goal is not merely to provide final-answer math examples. Each record is a structured olympiad reasoning object containing a normalized problem… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-math-olympiad-trajectories.verizon-codex-trace
Verizon Codex Trace
A sanitized Codex session trace of GPT-5.5/Codex working through a Verizon billing and trade-in support flow.
The trace includes the original JSONL session structure, assistant/user turns, tool calls, shell output, selected Computer Use browser screenshots, and the Verizon live-chat/bill context needed to understand the agent's work.
What's included
One saved Codex session JSONL under sessions/
Assistant messages, user messages, tool calls… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/verizon-codex-trace.APPS-verified
Introduction
This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
5000
117232
TACO-verified
4211
93921
Correct Ratio
84.22%
80.12%
sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.ReForm-DafnyComp-Benchmark
Re:Form Datasets
This repository contains the datasets used in the paper Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny.
The Re:Form project introduces a framework for code to specification generation using large language models, based on Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This work systematically explores ways to reduce human priors in scalable formal software verification by… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-DafnyComp-Benchmark.vericoding
Vericoding
A benchmark for vericoding: formally verified program synthesis
Sergiu Bursuc, Theodore Ehrenborg, Shaowei Lin, Lacramioara Astefanoaei, Ionel Emilian Chiosa, Jure Kukovec, Alok Singh, Oliver Butterley, Adem Bizid, Quinn Dougherty, Miranda Zhao, Max Tan, Max Tegmark
We present and test the largest benchmark for vericoding, LLM-generation of formally verified code from formal specifications - in contrast to vibe coding, which generates potentially buggy code from a… See the full description on the dataset page: https://huggingface.co/datasets/beneficial-ai-foundation/vericoding.verified-facts-sample-100
DeepInquiry Verified Facts (Sample-100)
A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus.
This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.swe-verified-gemini3-flash-trajectories
SWE-bench Verified — Gemini-3-flash agent trajectories (graded, 3 samples/instance)
Agent trajectories from gemini-3-flash-preview (high reasoning, temperature 0.8) run with the
OpenHands agent on SWE-bench Verified, in Modal sandboxes. For each of 100 instances
we sampled multiple trajectories and graded them with the SWE-bench harness; this dataset holds the
3 graded samples per instance = 296 trajectories, 198 resolved (67%).
pass@1 ≈ 66/100, pass@3 (oracle) = 73/100.… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swe-verified-gemini3-flash-trajectories.hw-verify
hw-verify — a hardware-security verification dataset with controls
Every positive example ships beside a deliberately broken counterpart, so a model or tool is graded against controls instead of against itself.
Try the checker that generated this data:
🔒 hw-verify Space — paste Verilog,
get a verdict, in your browser, no install.
Install
pip install datasets
30-second quickstart
from datasets import load_dataset
rtl =… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/hw-verify.swe-bench-verified-raw-traces-qwen3-coder
SWE-bench Verified raw mini-SWE-agent traces
Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified.
The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That
public dataset contains SWE-bench Verified questions and Kimi-generated answers; this
dataset uses only its instance IDs. The trajectory contents here are local
mini-SWE-agent/Qwen traces.
Files
data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/parsaidp/swe-bench-verified-raw-traces-qwen3-coder.hw-verify-paths
hw-verify-paths
▶ Try the checker in your browser · Docs & overview
Dependency graphs and witness paths for constant-time RTL analysis — the reasoning,
not just the label.
The companion dataset records
what each design is: CONSTANT_TIME or LEAKY. This one records why. For every
fixture it carries the full signal dependency graph, and for every leaky one the
concrete chains of signals that carry a secret to the observation.
Why witness paths and not just verdicts… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/hw-verify-paths.regexgym-verified-traces
RegexGym-Verified-Traces
Reasoning traces for writing regexes from examples. Each record shows a task (some strings that
should match, some that shouldn't), the teacher's chain-of-thought, and the regex it landed on.
Every trace here actually solved the task's hidden holdout — the regex was run against
examples the teacher never saw, and only exact solves were kept. The ground-truth regexes aren't
in the released records; the model has to earn its answer.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/regexgym-verified-traces.numeric-claim-verifier
Numeric Claim Verifier (Science) — Adaption AutoScientist
Programmatically verified prompt/completion pairs for scientific and statistical numeric claim verification.
Labels
correct — claim matches ground-truth tables
wrong_direction — trend/sign reversed
wrong_magnitude — right direction, wrong size (25–70% offset)
unverifiable — no matching source row (real entity + absent metric)
Sources
Our World in Data CO₂ / Energy
WHO GHO life expectancy… See the full description on the dataset page: https://huggingface.co/datasets/mishface123/numeric-claim-verifier.Z3-Verified-Reasoning-Graphs
Z3-Verified Constraint Reasoning Dataset
5k Baseline · Production-Ready · Zero Label Noise
The Problem This Solves
Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack.
Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different:
❌ No LLM-generated reasoning — zero hallucinations, zero label noise
✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.swe-bench-verified-raw-traces-qwen3-coder
SWE-bench Verified raw mini-SWE-agent traces
Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified.
The raw/easy split uses exactly the 194 instance IDs from
parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench
Verified questions and Kimi-generated answers; this dataset uses only its
instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces.
Files
data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/nikitamounier/swe-bench-verified-raw-traces-qwen3-coder.brainblast-verified-footgun-corpus
Brainblast — Verified SDK Footgun Corpus (free sample)
The only code-training data that ships with a machine-checkable proof. Each
record is a real insecure→fixed code footgun with a replayable RED→GREEN
receipt: a deterministic checker fails the insecure version and passes the fixed
one. You don't trust the labels — you replay the proof.
This repo is a free 40-record sample (receipt-only tier). The full corpus is
4,183 proven records across 154 SDKs and 9 vulnerability classes… See the full description on the dataset page: https://huggingface.co/datasets/dsb117/brainblast-verified-footgun-corpus.verified-analytics-tasks
Verified Analytics Tasks
150+ mainly small analytics and data-engineering tasks. The point of the set is
the answer key: every task ships its own automated checker, and every gold
answer was run through that checker and scored a clean 1.0 before the task was
allowed in. So the labels are more like "here's the checker, score it
yourself" instead of "just trust me bro."
I wanted to create a synthetic dataset inspired by this paper:
Autodata: An agentic data scientist to create… See the full description on the dataset page: https://huggingface.co/datasets/Eve39570/verified-analytics-tasks.clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k
データ件数: 140,980
平均トークン数: 602
最大トークン数: 2,040
合計トークン数: 84,894,510
ファイル形式: JSONL
ファイル分割数: 2
合計ファイルサイズ: 256.3 MB
加工内容:
データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。
response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。
prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.VeriSoftBench
VeriSoftBench
VeriSoftBench is a benchmark for evaluating neural theorem provers on software verification tasks in Lean 4.
The dataset contains 500 theorem-proving tasks drawn from 23 real-world Lean 4 repositories spanning compiler verification, type system formalization, applied verification (zero-knowledge proofs, smart contracts), semantic frameworks, and more.
📄 Paper (arXiv): https://arxiv.org/html/2602.18307v1💻 Full benchmark + pipeline + setup:… See the full description on the dataset page: https://huggingface.co/datasets/maxRyeery/VeriSoftBench.bullet-swebench-verified
Bullet on SWE-bench Verified — 479/500 = 95.8%
Results for the Bullet coding agent on all 500 instances of
SWE-bench Verified, graded by the official swebench.harness.run_evaluation
scorer. Every instance was attempted and graded; there are no empty patches.
479 / 500 resolved = 95.8%
Run with the Bullet harness on gpt-5.6-sol at high reasoning effort, one attempt per instance.
By repository
repository
resolved
django
223/231
96.5%
sympy
73/75
97.3%… See the full description on the dataset page: https://huggingface.co/datasets/daviddata1/bullet-swebench-verified.verifimind-peas-eval
VerifiMind-PEAS Evaluation Dataset
DOI: 10.5281/zenodo.21276884 · License: MIT · Version: 1.0
A human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Ground-truth labels were assigned by a single domain-expert annotator whose final verdicts and confidence were human judgments; disclosed LLM comprehension assistance was used during annotation (see Annotation Protocol — transparency is a design commitment… See the full description on the dataset page: https://huggingface.co/datasets/YSenseAI/verifimind-peas-eval.all-verified-benchmarks
All Verified Benchmarks
Real CPU benchmark data for ALL 31 working dispatchAI models.
Zero broken models. Zero partial models. All verified.
🚀 dispatchAI
verifiable-ai-provenance-bench
Verifiable AI Provenance Bench (TTTPS)
25 real timestamp-provenance receipts generated on 2026-08-04 by calling the
live KPP (Kenosian Protocol Platform) provenance API
(POST /v1/anchor, POST /v1/verify), which implements the TTTPS (Time-Token
Tamper-evident Provenance Seal) scheme. Each row is one real API round trip:
a content_hash was submitted to /v1/anchor, the returned receipt_id was then
submitted to /v1/verify, and both raw responses are recorded.
This dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/Pittro/verifiable-ai-provenance-bench.
