datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fever_gold_evidence
Dataset Card for fever_gold_evidence
Dataset Summary
Dataset for training classification-only fact checking with claims from the FEVER dataset.
This dataset is used in the paper "Generating Label Cohesive and Well-Formed Adversarial Claims", EMNLP 2020
The evidence is the gold evidence from the FEVER dataset for REFUTE and SUPPORT claims.
For NEI claims, we extract evidence sentences with the system in "Christopher Malon. 2018. Team Papelo: Transformer Networks at FEVER.… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/fever_gold_evidence.gero-research-evidence-2026-09
GERO research evidence — 123 publications
This dataset contains 123 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 122 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the bond_pricing immediate-start annuity audit, with official maintainer issue #8 sent and explicit limitations. Report counts are not independent-defect counts.
Latest numerical audits… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.sie-task-evidence
SIE task evidence
The recorded inputs and model responses behind the task pages on
superlinked.com, one folder per task.
Every figure published on a task page was produced by a real recorded run against
https://api.superlinked.com. This dataset holds those recordings so that anyone
can re-derive the published numbers without an API key and without spending any
inference.
How it is used
The runnable example for each task lives in the public
superlinked/sie… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/sie-task-evidence.unit-price-evidence-synthetic
Unit Price Evidence: Synthetic
This dataset contains rendered synthetic shopping pages and evidence-pointer
targets for product-card discovery and unit-price field extraction. It was built
to warm-start small encoder-decoder models without redistributing retailer HTML,
screenshots, product data, account data, or browsing history.
Release
Version: 0.1.0
Source code: erichasinternet/apples-to-apples
Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.cyber-evidence-dataset
Cyber Security Evidence Dataset — MITRE ATT&CK Safe Reference
This repository contains a narrow, title-only reference configuration derived from the official MITRE ATT&CK Enterprise v19.2 STIX data. It is published separately from the broader Cyber Security Evidence Dataset project because the CISA-derived evidence layers remain private and are not included here.
Scope
The release contains 697 deterministic records. Each record provides an ATT&CK technique… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-dataset.evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.contextualized-ST-Evidence
Contextualized ST-Evidence
A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask
split. Same 19,902 entries, same objects, same frames, same temporal evidence.
The only thing that changes is the spatial box on each frame.
This is the video counterpart of
shredder-31/contextualized-viscot,
built with the same model, the same prompt design and the same union-with-the-
original safety rule.
Why
ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.sovereign-evidence-observatory
Sovereign Evidence Observatory Casebook
The interesting question is not “Which AI sounds smartest?” It is “What kind of evidence would make this claim true, false, or still undecidable?”
This public casebook is the first dataset for the Sovereign Evidence Observatory. It turns model agreement, disagreement and abstention into inspectable evidence objects instead of treating consensus as truth.
Five layers
Shadow Mesh — independent model/provider observations.… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-evidence-observatory.airep-evidence-cases
AIREP Evidence Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93
Canonical current repository
https://github.com/halvrenofviryel/ai-runtime-evidence-protocol
Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.pi05-libero-goal-task-8-evidence
Robium Pi0.5 LIBERO-Goal Task 8 evidence
This is the public evidence bundle for Robium issue #69. It records one fixed,
no-retry evaluation of lerobot/pi05_libero_finetuned_v044 on LIBERO-Goal task
8, put_the_bowl_on_the_plate, using the canonical prompt “put the bowl on the
plate.”
Result
20/20 successful episodes; the predeclared target was 16/20.
Fixed initial states 0–19 map to seeds 1000–1019.
Batch size 1, hard environment/policy reset before every episode… See the full description on the dataset page: https://huggingface.co/datasets/robium/pi05-libero-goal-task-8-evidence.CB_claim_evidencelord-of-mysteries-fandom-evidence-sft
Lord of Mysteries Fandom Evidence SFT Dataset
Overview
This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.
The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.Med-Evidence-2.6k
Med-Evidence-2.6k
Med-Evidence-2.6k is a benchmark for evaluating evidence-grounded medical diagnostic reasoning. It was developed as part of the work EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents, accepted to Findings of EMNLP 2026.
The benchmark provides clinical diagnostic questions paired with ground-truth answers and annotated evidence spans, enabling evaluation of both diagnostic accuracy and evidence-grounded reasoning.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/danieez/Med-Evidence-2.6k.qfs-hf-jobs-campaign-evidence-20260909vietnamese-evidence-retrieval-indexes-v2-r1
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e.
Rows: 63,699
Source embedding shards: 13
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Dense input: title + text
Dense rows: deduplicated by content hash
Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.vietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.vietnamese-evidence-retrieval-indexes
Vietnamese Evidence Retrieval Indexes
Prebuilt exact dense and sparse indexes for
Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f.
Rows: 52,605
Source embedding shards: 11
Dense: FAISS IndexFlatIP, 1024 dimensions
Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75)
BM25 content: title repeated 2 times + chunk text
Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal
row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.douvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
website-technology-evidence-dataset
Website Technology Evidence Dataset
A balanced synthetic dataset for classifying observable website fingerprint evidence into a likely web technology.
Provenance
All records are synthetically generated from documented, recognizable public fingerprints. No claim is made that these records were collected from real websites.
Dataset
38 technologies
6,840 examples
180 examples per technology
train: 5,472
validation: 684
test: 684
balanced classes… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/website-technology-evidence-dataset.production-evidence-records
SHAR Production Evidence Records
Twenty-five bilingual machine-readable production evidence record contracts and clearly synthetic valid examples by SHAR Production.
The collection covers approvals, rights, permits, technical reports, quality control, localization, delivery, archive, provenance, supplier review, incidents and final release gates. synthetic-examples.jsonl contains invented records only. It contains no client, talent, supplier or project data.
Dataset and… See the full description on the dataset page: https://huggingface.co/datasets/SHARProduction/production-evidence-records.clinical-evidence-state-transition-fidelity-v0.1
Clinical Evidence State Transition Fidelity v0.1
A synthetic clinical reasoning dataset for evaluating whether an AI system can update a structured clinical state selectively, proportionately, and consistently when new evidence arrives.
Repository:
ClarusC64/clinical-evidence-state-transition-fidelity-v0.1
Version:
0.1.0
Publisher:
Clarus Invariant
Framework:
SIOS
Dataset identity
Clinical Evidence State Transition Fidelity v0.1 evaluates whether a model can… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-state-transition-fidelity-v0.1.production-rights-evidence-matrix-fixtures
SHAR Production Rights Evidence Matrix Fixtures
SHAR Production is an AI-hybrid video production studio. This CC-BY-4.0 dataset contains fictional multi-asset inventories for testing rights-evidence aggregation tools.
Each row includes an asset inventory and the expected release decision. Only inventories with every asset cleared and backed by an HTTPS evidence URL are releasable. This dataset is synthetic and contains no client work or legal determinations.
Compatible tool:… See the full description on the dataset page: https://huggingface.co/datasets/SHARProduction/production-rights-evidence-matrix-fixtures.rights-evidence-matrix-fixtures
SHAR Production Rights Evidence Matrix Fixtures
SHAR Production is an AI-hybrid video production studio. This CC-BY-4.0 dataset contains fictional multi-asset inventories for testing rights-evidence aggregation tools.
Each row includes an asset inventory and the expected release decision. Only inventories with every asset cleared and backed by an HTTPS evidence URL are releasable. This dataset is synthetic and contains no client work or legal determinations.
Compatible tool:… See the full description on the dataset page: https://huggingface.co/datasets/SHARProduction/rights-evidence-matrix-fixtures.AgroVeritas-Evidence-QA-Adapted
AgroVeritas Evidence QA
Agricultural intelligence you can audit — in English and Spanish.
AgroVeritas is a bilingual, evidence-bounded agricultural instruction dataset built for regional crop-calendar and historical climate reasoning. It teaches models to answer a practical question, cite the evidence used, show the reasoning path, state limitations, and recommend local verification instead of presenting historical data as live field conditions.
Dataset Viewer and… See the full description on the dataset page: https://huggingface.co/datasets/MarianaCodebase/AgroVeritas-Evidence-QA-Adapted.adaption-agri-qa-with-evidence-boundaries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agri_qa_with_evidence_boundaries
This dataset contains question-answer pairs focused on agricultural best practices, including crop management, pest control, and irrigation strategies. Each entry provides a direct, evidence-based response followed by a clearly defined 'evidence boundary' that limits the scope of the advice and advises verification with local conditions. The content… See the full description on the dataset page: https://huggingface.co/datasets/yeziR4/adaption-agri-qa-with-evidence-boundaries.eh-margin-evidence-responsiveness-worldknown
margin-evidence-responsiveness-worldknown -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-margin-evidence-responsiveness-worldknown
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-margin-evidence-responsiveness-worldknown.human-gene-lof-rescue-evidence
Human Biallelic Loss-of-Function and Functional-Rescue Evidence
This dataset contains 88 curated human gene records linking three experimentally distinct observations:
biallelic human loss of function;
a consistent phenotype reported in independent affected families or cohorts;
functional rescue in affected humans or patient-derived human cells.
Each record therefore connects genotype → recurrent human phenotype → reversal of a disease-relevant defect. This convergent evidence… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/human-gene-lof-rescue-evidence.clinical-evidence-dependency-graph-reasoning-v0.1
Clinical Multi-Evidence State Integration v0.1
Overview
Clinical Multi-Evidence State Integration v0.1 is a structured clinical-reasoning benchmark designed to test whether an AI system can integrate multiple sequential evidence events into a coherent final clinical state.
The benchmark evaluates more than final-answer classification.
A system must determine:
how each evidence event affects each tracked clinical item;
whether an item should be confirmed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-dependency-graph-reasoning-v0.1.pythia-paths-evidence
Pythia Paths Evidence
A small, revision-pinned evidence bundle for examining model-training paths
without converting a trend into authority.
Companion read-only interface: Pythia Paths Static Space
(mutable navigation; the evidence files below remain digest-pinned).
Initial scope
Model: EleutherAI/pythia-70m-deduped
Run: the default public run only
Context coverage: all 27 zero-shot reports in one pinned directory
Detailed coverage: four post-outcome-selected… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/pythia-paths-evidence.Health_Information_Seeking_under_Limited_Evidence
Health Information Seeking under Limited Evidence (HISLE)
HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs.
File
Records
Contents
matched_pairs_47.jsonl
47
Matched Chinese–English scenario pairs
matched_variants_3290.jsonl
3,290
Query variants for the matched scenarios
coverage_originals_24.jsonl
24
Coverage-expansion queries
coverage_variants_840.jsonl
840
Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.
