datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RealKIE-FCC-Verified
RealKIE-FCC-Verified
It is a test set with single and multi-page invoices sourced from the Federal Communications Commission (FCC) to evaluate key information extraction (KIE) performance.
Task
Extract information from the document in JSON format given the corresponding JSON schema. It contains 75 documents, with:
a) image_files: Each document has multiple pages
b) json_schema: A common JSON schema requiring extraction of specified information including line… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/RealKIE-FCC-Verified.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version.
This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes:
Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.vistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.VisualWebInstruct-verified
🧠 VisualWebInstruct-Verified: High-Confidence Multimodal QA for Reinforcement Learning
VisualWebInstruct-Verified is a high-confidence subset of VisualWebInstruct, curated specifically for Reinforcement Learning (RL) and Reward Model training.
It contains verified multimodal question–answer pairs where correctness, reasoning quality, and image–text alignment have been explicitly validated.
This dataset is ideal for RLVR training pipelines.
📘 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VisualWebInstruct-verified.Signature-Verification-Dataset
Multilingual Signature Verification Dataset
Dataset Summary
The Multilingual Signature Verification Dataset is a curated collection of handwritten signatures designed for offline signature verification and related computer vision tasks.
The dataset contains more than 7,000 signature images spanning three major writing systems:
Hindi
Bengali
English
The English portion includes samples from the well-known CEDAR Signature Dataset, while additional Hindi and… See the full description on the dataset page: https://huggingface.co/datasets/rakshitdabral/Signature-Verification-Dataset.veri_seti_adithesis-v18-formal-verification
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
Ouroboros Thesis v18 — Formal Verification
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Historical snapshot — this dataset is the v18-specific Lean mechanization index. The live source of truth is lean-proofs-v1, which is kept… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-v18-formal-verification.multimodal-open-r1-8k-verifiedveri-render
VeriRender Benchmark Dataset
Causal consistency verification samples for Vision-Language Models.
Layout
manifest.jsonl ← canonical index (one row per sample)
benchmark.yaml ← config used to generate this release
inconsistent/{domain}/{sample_id}/ ← corrupted evaluation samples
consistent/{domain}/{sample_id}/ ← negative controls (clean images)
Splits
Split
Description
Eval image
inconsistent
Symbolic spec is… See the full description on the dataset page: https://huggingface.co/datasets/VietMedTeam/veri-render.ai2thor_spatial_verification_val_v2VeriSciQA
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Paper: VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Dataset Description
VeriSciQA is a large-scale, high-quality dataset for Scientific Visual Question Answering (SVQA), containing 20,272 QA pairs spanning 20 scientific domains, 12 figure types, and 5 question types. The dataset is constructed using a Cross-Modal Verification framework that generates QA pairs from… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/VeriSciQA.ai2thor_spatial_verification_test_v2VeriEvol-RL
VeriEvol-RL
RL training data for VeriEvol: Scaling Multimodal Mathematical Reasoning via
Verifiable Evol-Instruct. This is the reinforcement-learning stage dataset used
for GRPO-style training on top of the SFT-initialized policy (see the companion
SFT set Ringo1110/VeriEvol-SFT).
📄 Paper: arXiv:2606.23543
💻 Code: github.com/lihaoling/VeriEvol
Each example is a single-image visual math/reasoning problem with a verifiable
ground-truth answer, formatted for the verl
RL… See the full description on the dataset page: https://huggingface.co/datasets/Ringo1110/VeriEvol-RL.verilog-wavedrom
Verilog Wavedrom
A combination of verilog modules and their correspondig timing diagrams generated by wavedrom.
Dataset Details
A collection of wavedrom timing diagrams in PNG format representing verilog modules.
The Verilog modules were copied from shailja/Verilog_GitHub.The timing diagrams were generated by first generating testbenches for the individual verilog modules through the Verilog Testbench Generator from EDA Utils VlogTBGen.The resulting testbenches were… See the full description on the dataset page: https://huggingface.co/datasets/vkenbeek/verilog-wavedrom.ai2thor_spatial_verification_val_v1Veridis
VERIDIS Dataset
Overview
This repository contains the VERIDIS dataset, a collection of annotated agricultural images for crop detection and identification. The dataset comprises field images of beet and corn crops captured by a ground-level robotic platform, organized in YOLO format for object detection tasks.
The dataset primarily captures crops at early growth stages, which is particularly relevant for applications such as plant detection, early monitoring, and… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/Veridis.vlm-verification-logs
VLM Verification Conversation Logs
Raw conversation logs from VLM solvers and judges across CharXiv and CountBench.
Every row is one model interaction: the image, the exact prompt(s) sent, the full model output, the
extracted answer, and correctness. The verification runs include the verifier model's prompt, output and
verdict.
Configs
config
rows
description
logs
644,093
one row per solver / verifier / agentic / rejection / self-consistency… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/vlm-verification-logs.VeriEvol-SFT
VeriEvol-SFT
SFT data for VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable
Evol-Instruct. This is the supervised fine-tuning (SFT) stage dataset used to
initialize the policy model before GRPO-style RL.
📄 Paper: arXiv:2606.23543
💻 Code: github.com/lihaoling/VeriEvol
Each example is a single-image, single-turn visual STEM reasoning problem paired
with a long chain-of-thought solution. Prompts are produced by route-specific
evolution operators that rewrite… See the full description on the dataset page: https://huggingface.co/datasets/Ringo1110/VeriEvol-SFT.M2-Verify-Medai2thor_spatial_verification_test_v1chart-reasoning-verified
chart-reasoning-verified
Chart reasoning examples generated from an explicit latent representation.
The data, the question and the answer are computed before the chart is
drawn, so the image is a rendering of known ground truth rather than the
source of it. No model was asked to label anything.
Each row carries both a rendered chart and a text serialisation of the same
chart, so the set is usable for vision-language training and for text-only
language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.Signature-Verification-Dataset
Multilingual Signature Verification Dataset
Dataset Summary
The Multilingual Signature Verification Dataset is a curated collection of handwritten signatures designed for offline signature verification and related computer vision tasks.
The dataset contains more than 7,000 signature images spanning three major writing systems:
Hindi
Bengali
English
The English portion includes samples from the well-known CEDAR Signature Dataset, while additional Hindi and… See the full description on the dataset page: https://huggingface.co/datasets/Peachyy2208/Signature-Verification-Dataset.mup_verification
μP Verification: Gradient Update Invariance Analysis
Empirical verification of Maximal Update Parametrization (μP): Demonstrating that μP achieves width-invariant gradient updates, enabling hyperparameter transfer across model scales.
🎯 Key Finding
μP shows 42.8% less width-dependence than Standard Parametrization (SP) in relative gradient updates, confirming the theoretical prediction that μP enables hyperparameter transfer across model widths.
Metric… See the full description on the dataset page: https://huggingface.co/datasets/AmberLJC/mup_verification.papergym-verify
PaperGym — verification set
Each row shows one scientific figure and asks for one number plotted in it.
The gold answer was not read off the figure — it was recomputed from the source
data table the paper published alongside it, by an LLM pipeline.
Your job is to check whether that gold is actually the number the figure plots.
The pipeline is good at arithmetic and bad at knowing when its own assumptions are
wrong. Every defect found so far has the same shape: the recipe… See the full description on the dataset page: https://huggingface.co/datasets/yhzhang3/papergym-verify.face-verification
Dataset Card for "face-verification"
More Information needed
VQA-VerifyThis is the VQA-Verify dataset, introduced in the paper SATORI-R1: Incentivizing Multimodal Reasoning with Spatial Grounding and Verifiable Rewards.
Arxiv Here | Github
VQA-Verify is a 12k dataset annotated with answer-aligned captions and bounding boxes. It's designed to facilitate training models for Visual Question Answering (VQA) tasks, particularly those employing free-form reasoning. The dataset addresses limitations in existing VQA datasets by providing verifiable intermediate steps and… See the full description on the dataset page: https://huggingface.co/datasets/justairr/VQA-Verify.saffron-verify
SaffronVerify Dataset
Dataset Summary
SaffronVerify is a fine-grained image classification dataset for saffron quality grading. It contains images of saffron across three quality categories — from premium grade to adulterated samples — intended for training computer vision models to detect saffron purity and adulteration.
Dataset Structure
The dataset follows the standard ImageFolder layout and is split into training and validation sets.
saffron-verify/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/saffron-verify.Verilog_VL
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Navyandsu/Verilog_VL.trac-verify-raw-divideindustry-verification-seml
