datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bitaudit_verification_dataset_v2vistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.policy-alignment-verification-dataset
Policy Alignment Verification Dataset
🌐 NAVI's Ecosystem 🌐
🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance.
🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification.
📜 API Docs – Your starting point for integrating NAVI into your applications.
📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses.
✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.temperature-verification
CastCheck — daily station-level verification of public weather forecasts
Independent, automated verification of raw 2 m temperature forecasts from operational NWP
(ECMWF IFS HRES, NCEP GFS) and AI models (ECMWF AIFS Single; NOAA/CIRA operational runs of GraphCast,
Pangu-Weather, FourCastNet v2 and Aurora from both GFS and IFS initial conditions) at 23
U.S. first-order stations — 22 major airports plus New York Central Park. The headline metric is the instantaneous 2 m… See the full description on the dataset page: https://huggingface.co/datasets/castcheck/temperature-verification.bitaudit_verification_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.ai2thor_spatial_verification_val_v2ai2thor_spatial_verification_test_v2evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.ai2thor_spatial_verification_val_v1ai2thor_spatial_verification_test_v1authorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.task089_swap_words_verification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.AIME24-25_CoT_Verification
Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
📌 Dataset Summary
This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.vlm-verification-logs
VLM Verification Conversation Logs
Raw conversation logs from VLM solvers and judges across CharXiv and CountBench.
Every row is one model interaction: the image, the exact prompt(s) sent, the full model output, the
extracted answer, and correctness. The verification runs include the verifier model's prompt, output and
verdict.
Configs
config
rows
description
logs
644,093
one row per solver / verifier / agentic / rejection / self-consistency… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/vlm-verification-logs.task1294_wiki_qa_answer_verification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1294_wiki_qa_answer_verification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1294_wiki_qa_answer_verification.pan2020-authorship-verificationhiring-analyses-second_model_verification-enaime24_distill_1p5B_verificationsCoT-Verification-340k
CoT-Verification-340k Dataset: Improving Reasoning Model Efficiency through Verification
This dataset is used for supervised verification fine-tuning of large reasoning models. It contains 340,000 question-solution pairs annotated with solution correctness, including 160,000 correct Chain-of-Thought (CoT) solutions and 190,000 incorrect ones. This data is designed to train models to effectively verify the correctness of reasoning steps, leading to more efficient and accurate… See the full description on the dataset page: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.agent-exchange-verification
Agent Exchange: Verified-Work Benchmark
A benchmark for testing whether an LLM judge catches fabricated work, or pays for it. Given a source document (a contract clause, an NDA snippet, a scientific abstract) and a claim about it, does the judge correctly flag claims that are not actually supported? The dataset ships a human-graded calibration set, a multi-model audit of the ambiguous cases a judge waves through, and the real-world documents behind both.
It is the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/soren19/agent-exchange-verification.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.corpus-verification
SPP Corpus Verification
Checksums and document-boundary indices for verifying a rebuilt copy of the
Synthetic Persona Pretraining (SPP) training corpus, byte for byte.
The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB,
compact.bin 1.75 TB) and are fully derived from the published reflections, the
uid manifest, and the tokenizer recipe — so they are not published. These
.idx sidecars carry per-document boundaries and lengths, which is enough to
prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.sciclaims_verification_data
Verification Dataset for SciClaims
This is the verification dataset used in the 2025 EMNLP demonstration paper
SciClaims: An End-to-End Generative System for Biomedical Claim Analysis
(Ortega and Gómez-Pérez).
The dataset contains approximately 4.7 million PubMed abstracts published
between 2000 and 2022. The records were selected using Semantic Scholar's
Highly Influential Citations metric, requiring each article to be supported
by at least three highly influential citations.… See the full description on the dataset page: https://huggingface.co/datasets/expertailab/sciclaims_verification_data.strict-verification-reasoning
Strict Verification Reasoning Dataset
Description
A dataset for training language models to verify facts, check sources, evaluate arguments, and avoid overthinking.
Content
1,010,000 examples
5 categories: anti-overthink, comparisons, strict facts, strict sources, strict arguments
English language
Categories
Category
Description
%
Anti-Overthink
Simple, direct answers
15%
Comparisons
Hallucination vs correct answer
20%… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/strict-verification-reasoning.industry-verification-semlpan2020-authorship-verificationcitation-verification-benchmark
CiteMe Citation Verification Benchmark v1
A frozen multilingual benchmark for evaluating citation- and reference-verification
systems. It contains 340 labeled references with ground truth known by construction:
Ground-truth class
Count
Construction
EXISTS_CORRECT
160
Real works confirmed by DOI and title in OpenAlex and Crossref
EXISTS_CORRUPTED
100
Real works with exactly one controlled metadata corruption
FABRICATED
80
Constructed references whose invented… See the full description on the dataset page: https://huggingface.co/datasets/danielnichiata/citation-verification-benchmark.adaption-math-logic-verification-pairs-augmented-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs (augmented) (augmented)
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary.… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs-augmented-augmented.
