datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.agent-exchange-verification
Agent Exchange: Verified-Work Benchmark
A benchmark for testing whether an LLM judge catches fabricated work, or pays for it. Given a source document (a contract clause, an NDA snippet, a scientific abstract) and a claim about it, does the judge correctly flag claims that are not actually supported? The dataset ships a human-graded calibration set, a multi-model audit of the ambiguous cases a judge waves through, and the real-world documents behind both.
It is the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/soren19/agent-exchange-verification.Fast-Multi-Agent-Identity-Verification
