datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distill_r1_qwen_math_1.5b_128_solns_math_verificationsdistill_qwen_7b_aime_verifications_7b_ft_verifierbitaudit_verification_dataset_v2vistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.HLE-Verifications
HLE with Gemini 3 Pro
This dataset contains 649 multiple-choice and exact-match questions from the Humanity's Last Exam (HLE) benchmark with 50 candidate responses generated by Gemini 3 Pro for each problem. Each response has been evaluated for correctness using a mixture of Qwen3-Next-80B-A3B-instruct and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric.
Dataset Structure
Split: Single split named "data"
Number… See the full description on the dataset page: https://huggingface.co/datasets/FUSE-verifiers/HLE-Verifications.atlas-31-strengthening-candidate-verification-under-rl
31. Strengthening candidate verification under reinforcement learning
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 31-strengthening-candidate-verification-under-rl/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-31-strengthening-candidate-verification-under-rl only.
How can reinforcement learning make the orchestrator's comparing and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-31-strengthening-candidate-verification-under-rl.distill_r1_qwen_math_1.5b_128_solns_aime_verificationsdistill_qwen_7b_math_verifications_7b_ft_verifierpolicy-alignment-verification-dataset
Policy Alignment Verification Dataset
🌐 NAVI's Ecosystem 🌐
🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance.
🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification.
📜 API Docs – Your starting point for integrating NAVI into your applications.
📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses.
✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.Signature-Verification-Dataset
Multilingual Signature Verification Dataset
Dataset Summary
The Multilingual Signature Verification Dataset is a curated collection of handwritten signatures designed for offline signature verification and related computer vision tasks.
The dataset contains more than 7,000 signature images spanning three major writing systems:
Hindi
Bengali
English
The English portion includes samples from the well-known CEDAR Signature Dataset, while additional Hindi and… See the full description on the dataset page: https://huggingface.co/datasets/rakshitdabral/Signature-Verification-Dataset.temperature-verification
CastCheck — daily station-level verification of public weather forecasts
Independent, automated verification of raw 2 m temperature forecasts from operational NWP
(ECMWF IFS HRES, NCEP GFS) and AI models (ECMWF AIFS Single; NOAA/CIRA operational runs of GraphCast,
Pangu-Weather, FourCastNet v2 and Aurora from both GFS and IFS initial conditions) at 23
U.S. first-order stations — 22 major airports plus New York Central Park. The headline metric is the instantaneous 2 m… See the full description on the dataset page: https://huggingface.co/datasets/castcheck/temperature-verification.thesis-v18-formal-verification
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
Ouroboros Thesis v18 — Formal Verification
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Historical snapshot — this dataset is the v18-specific Lean mechanization index. The live source of truth is lean-proofs-v1, which is kept… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-v18-formal-verification.face-verification-with-featuresbitaudit_verification_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.ai2thor_spatial_verification_val_v2ai2thor_spatial_verification_test_v2evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.mistral-7b-hidden-states-tpu-verification
Mistral-7B TPU hidden-state extractor verification
Verification-only artifact — not a reproduction of the paper's AUROC.
This dataset contains last-token embeddings and hidden states extracted from
1,000 flattened CoQA validation question/reference-answer pairs with
mistralai/Mistral-7B-Instruct-v0.3. It verifies that the memory-bounded TPU
v5e-1 extraction path runs successfully. It does not contain the Mistral
best_answer generations or hallucination labels required to… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/mistral-7b-hidden-states-tpu-verification.authorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.ai2thor_spatial_verification_val_v1ai2thor_spatial_verification_test_v1steering-verification-captures
Steering verification captures
Rendered camera frames along two driving routes, in four weather conditions each. These
are the input to the formal certificates in
formal-verification--steering--code,
and they are published because they are what makes those certificates checkable without
a simulator.
The paper is Testing Between the Test Cases: Proving End-to-End Steering in Conditions
You Never Drove (arXiv:2609.10951).
route
conditions
poses
frames… See the full description on the dataset page: https://huggingface.co/datasets/AD-Assurance-Lab/steering-verification-captures.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.AIME24-25_CoT_Verification
Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
📌 Dataset Summary
This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.vlm-verification-logs
VLM Verification Conversation Logs
Raw conversation logs from VLM solvers and judges across CharXiv and CountBench.
Every row is one model interaction: the image, the exact prompt(s) sent, the full model output, the
extracted answer, and correctness. The verification runs include the verifier model's prompt, output and
verdict.
Configs
config
rows
description
logs
644,093
one row per solver / verifier / agentic / rejection / self-consistency… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/vlm-verification-logs.task089_swap_words_verification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task089_swap_words_verification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task089_swap_words_verification.aeb-verification-captures
Camera frames and trained networks for the braking verification study
The inputs to AD-Assurance-Lab/formal-verification--aeb--code, measured on 2026-09-10.
Why this exists. The study's results are in git. These files are not: they are large and
git ignores them. Without them the published numbers can be reproduced in kind but never
exactly, because the simulator does not render bit-identical frames from one run to the
next. Different frames give different networks, which give… See the full description on the dataset page: https://huggingface.co/datasets/AD-Assurance-Lab/aeb-verification-captures.daily-paper-2026-09-01-verification-placement-safety-cost
Pre- versus Post-Action Verification: Measuring the Safety-per-Dollar Frontier of Gating versus Auditing in Autonomous K8s Agent Loops
TL;DR — Unattended LLM agents operating a Kubernetes cluster need an in-loop verifier, but where in the action loop that verifier sits (before or after the action) was never a measured design variable. This paper formalizes verifier placement as an independent axis with a decision-theoretic per-action cost model at a frozen model tier, derives… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-01-verification-placement-safety-cost.test_box_verification_envIMO-Shortlist-Verifications
IMO Shortlist with Qwen3-30B-A3B-Thinking-2507
This dataset contains 123 questions from the International Mathematical Olympiad (IMO) Shortlist subset of the IMO AnswerBench benchmark with 50 candidate responses generated by Qwen3-30B-A3B-Thinking-2507 for each problem. Each response has been evaluated for correctness using a mixture of DeepSeek-R1-Distill-Llama-70B and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric. The… See the full description on the dataset page: https://huggingface.co/datasets/FUSE-verifiers/IMO-Shortlist-Verifications.
