datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciCode-Verified
SciCode-Verified
SciCode-Verified is the corrected, human-verified release of the
SciCode scientific-code-generation benchmark.
A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and
corrected every confirmable defect. The released evaluation set contains 64 main problems and
287 scored subproblems; one original problem is excluded because its specification does not
determine a unique, verifiable answer.
Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.TACO-verified
Introduction
This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
25443
1468722
TACO-verified
12898
1043251
Correct Ratio
50.69 %
71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.verified-defi-datasets
Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus
Dataset Description
High-density, verified AI fine-tuning dataset in ALPACA format.
Domain: Solana Sealevel & Anchor Program Optimization
Verified Records: 3
Estimated Tokens: 339
Quality QA Score: 99.0%
Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution
nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.Lego-RL-SWE-Bench-Verified
Lego-RL-SWE-Bench-Verified
The 500 SWE-bench Verified instances as ready-to-run harbor RL environments —
the exact evaluation set behind every SWE-bench Verified number in
LEGO-RL, packaged the same way as the training
set Lego-X/Lego-RL-2699 so
one trainer reads both.
Two parallel views of the same 500 instances:
View
Path
What it is
Official SWE-bench records
swebench_verified_official_500/
The upstream princeton-nlp/SWE-bench_Verified rows, verbatim
Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.MAPS_Verified
Dataset Card for Multilingual Benchmark for Global Agent Performance and Security
This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.verified_wiki_historian_the_beatles_anthology_dataset_active
Verified-Wiki-Historian: The Beatles Anthology
Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning.
This dataset is a cleaned and rebuilt refinement of:
Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active
The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.Doc2Feat-bench_Verified
Dataset Summary
NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically.
Languages
The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type.
Dataset Structure
An example of a SWE-bench datum is as follows:
repo: (str) - The repository owner/name identifier from GitHub.
instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.VeriTime
VeriTime: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning
This is the dataset associated with our paper:
Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning
Jiahui Zhou, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Lin Li, Zhuomin Chen, Jian Lou, See-Kiong Ng
ICML 2026 | Paper
Dataset Construction Pipeline: TSRgen
TSRgen is an… See the full description on the dataset page: https://huggingface.co/datasets/HayleyZhou1113/VeriTime.sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.idfu-verified-code
IDFU Code Negative Dataset — Free Preview
A curated dataset of Python code samples that failed execution-based
validation, designed for training reward models, DPO rejected-side pairs,
and error-detection classifiers. Free 100-sample preview; paid full versions
available separately.
What's inside this preview
100 unique Python samples, all AST-validated
19 CS domains represented (MCMC, FFT, distributed consensus, ZKP,
formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.verified-tool-use-dataset
Verified tool-use trajectories for LLM agents
This was a time-boxed experiment by an autonomous agent (Protogonos), now concluded. Nothing here is offered for sale or for hire, and no payment is accepted.
Multi-turn function-calling conversations for training and evaluating
tool-using agents — 48 trajectories across 16 domains, with every tool call
checked against its tool's JSON-Schema. The free sample in this repo is a real
slice of the full set: the viewer above renders it… See the full description on the dataset page: https://huggingface.co/datasets/protogonos/verified-tool-use-dataset.kodcode-verified-python-235k
KodCode-Verified Python — 234,555 execution-verified Python SFT rows
One row per problem. Every assistant turn is code that passed its own unit tests when
actually run — real pytest against KodCode-V1's
tests, in a pinned interpreter, in a sandboxed subprocess. No LLM judge, no heuristic
filter, no model-generated answers.
Unlike the v3 release this supersedes, the corpus is deduplicated, decontaminated against
HumanEval/MBPP, and stripped of rows whose tests cannot constrain… See the full description on the dataset page: https://huggingface.co/datasets/F-A-I-L/kodcode-verified-python-235k.android-kotlin-compose-compiler-verified
Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset
5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs.
Every SFT row was actually compiled — not LLM-approved, not heuristically
filtered. A subset was verified behaviorally by running JUnit tests.
Built to fine-tune small models into focused Android specialists rather than
general-purpose coders.
Why this exists
Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.APPS-verified
Introduction
This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
5000
117232
TACO-verified
4211
93921
Correct Ratio
84.22%
80.12%
VeriReason-RTL-Coder_7b_reasoning_tb_simple
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.VeriReason-RTL-Coder_7b_reasoning_tb
Verireason-RTL-Coder_7b_reasoning_tb
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.ReForm-Python2Dafny-Dataset
Re:Form Datasets
This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny".
The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.verisci-verified-science-math-code
VeriSci Verified Science Math Code
Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage.
Summary
VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.openrtlset-permissive-verified
openrtlset-permissive-verified
A permissive-filtered, machine-verified subset of
ESCAD/OpenRTLSet.
Every record in this dataset compiles. Each completion was elaborated and linted with
Verilator (--lint-only -Wall, warnings fatal) and admitted only on a clean run. Nothing
here is graded by an LLM judge.
Contents
13,625 records drawn from 2,174 distinct upstream repositories
Task: specification + pinned module interface -> complete Verilog module
Upstream… See the full description on the dataset page: https://huggingface.co/datasets/theepicflyer/openrtlset-permissive-verified.ReForm-DafnyComp-Benchmark
Re:Form Datasets
This repository contains the datasets used in the paper Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny.
The Re:Form project introduces a framework for code to specification generation using large language models, based on Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This work systematically explores ways to reduce human priors in scalable formal software verification by… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-DafnyComp-Benchmark.hardware-verilogeval-v2
hardware-verilogeval-v2
VerilogEval v2 - 471 Verilog evaluation problems
Dataset Overview
This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks.
Files
verilog_eval_problems.json: 471 VerilogEval v2 problems
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.Verified-Camel
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject.
Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Verified-Camel.vericoding
Vericoding
A benchmark for vericoding: formally verified program synthesis
Sergiu Bursuc, Theodore Ehrenborg, Shaowei Lin, Lacramioara Astefanoaei, Ionel Emilian Chiosa, Jure Kukovec, Alok Singh, Oliver Butterley, Adem Bizid, Quinn Dougherty, Miranda Zhao, Max Tan, Max Tegmark
We present and test the largest benchmark for vericoding, LLM-generation of formally verified code from formal specifications - in contrast to vibe coding, which generates potentially buggy code from a… See the full description on the dataset page: https://huggingface.co/datasets/beneficial-ai-foundation/vericoding.
