datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openverifiable-enwiki-20260901-20260918-r1-evidence
OpenVerifiableLLM Wikipedia provenance evidence
Development in progress. No production-trained or end-to-end verified model is
published here yet. Synthetic test results do not establish Wikipedia training.
This new AOSSIE repository is reserved for publicly reconstructible inputs,
checkpoints and reports for the OpenVerifiableLLM Wikipedia base model and its
conversational derivative. The governing goal and code are maintained at
AOSSIE-Org/OpenVerifiableLLM.
The intended… See the full description on the dataset page: https://huggingface.co/datasets/AOSSIE/openverifiable-enwiki-20260901-20260918-r1-evidence.fever_gold_evidence
Dataset Card for fever_gold_evidence
Dataset Summary
Dataset for training classification-only fact checking with claims from the FEVER dataset.
This dataset is used in the paper "Generating Label Cohesive and Well-Formed Adversarial Claims", EMNLP 2020
The evidence is the gold evidence from the FEVER dataset for REFUTE and SUPPORT claims.
For NEI claims, we extract evidence sentences with the system in "Christopher Malon. 2018. Team Papelo: Transformer Networks at FEVER.… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/fever_gold_evidence.gero-research-evidence-2026-09
GERO research evidence — 120 publications
This dataset contains 120 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 119 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the actuarialmath continuous-annuity selection-duration audit, with verified developer issue7 and explicit limitations. Report counts are not independent-defect counts.
Latest numerical… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.VeriLoop-Coder-E1-Evaluation-Evidence
VeriLoop Coder-E1 Evaluation Evidence
This repository contains the public evaluation-evidence packages
referenced by the official VeriLoop Coder-E1 benchmark result files.
Model repository:
tsinghua-sigs-robot-lab/veriloop-coder-e1
Evidence packages
Benchmark
Evidence directory
DeepSWE
veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0
SWE-bench Pro
veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0
SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.sie-task-evidence
SIE task evidence
The recorded inputs and model responses behind the task pages on
superlinked.com, one folder per task.
Every figure published on a task page was produced by a real recorded run against
https://api.superlinked.com. This dataset holds those recordings so that anyone
can re-derive the published numbers without an API key and without spending any
inference.
How it is used
The runnable example for each task lives in the public
superlinked/sie… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/sie-task-evidence.evidence_infer_treatmentData and code from our "Inferring Which Medical Treatments Work from Reports of Clinical Trials", NAACL 2019. This work concerns inferring the results reported in clinical trials from text.
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts' associated with them. These prompts will ask about the relationship between an intervention and comparator with respect to an outcome, as reported in the trial. For example, a prompt may ask about the reported effects of aspirin as compared to placebo on the duration of headaches. For the sake of this task, we assume that a particular article will report that the intervention of interest either significantly increased, significantly decreased or had significant effect on the outcome, relative to the comparator.
The dataset could be used for automatic data extraction of the results of a given RCT. This would enable readers to discover the effectiveness of different treatments without needing to read the paper.rss-robot-perception-evidence
Robot perception research artifacts
This is an exploratory experiment archive for an RSS research project. It is not a claim of acceptance, a finalized paper, or an official VER reproduction.
Code and protocol: https://github.com/YananZHOU5555/rss-robot-perception (private).
Weights: https://huggingface.co/B111ue/rss-robot-perception-checkpoints
Evidence: https://huggingface.co/datasets/B111ue/rss-robot-perception-evidence
This snapshot contains 228 complete evaluation cohorts… See the full description on the dataset page: https://huggingface.co/datasets/B111ue/rss-robot-perception-evidence.NIDS-Thesis-Experimental-Evidence
thesis_pipeline
Pipeline rebuilt following the supervisor's conditional review (13/07/2026). See SCOPE_FROZEN.md at the parent project root for the frozen scientific scope.
Structure
config/: centralized configuration (paths, seeds)
manifests/: data audits and temporal split manifests
src/data/: dataset preparation and cleaning
src/models/: model training and comparison
src/evaluation/: aggregation and global comparison of results
tests/: leakage tests and… See the full description on the dataset page: https://huggingface.co/datasets/MamadouSY-NIDS-Thesis/NIDS-Thesis-Experimental-Evidence.kda-graft-econ-evidence-2026-09-12🇷🇺 Русская версия (этот файл) · 🇬🇧 English version
kda-graft — доказательная база эксперимента (2026-09-11 … 09-12)
Сырые артефакты эксперимента «трансплантация KDA в Qwen2.5-0.5B»: числа, прогоны, консоль-логи,
графики, скрипты и recall-датасеты. Модель — Rob1234567/kda-graft-qwen2.5-0.5b-linear-attention-3to1.
Здесь всё, из чего собраны три отчёта (в reports/) — включая отрицательные результаты:
дальний recall на килодистанциях несут attention-слои, KDA-память там… See the full description on the dataset page: https://huggingface.co/datasets/Rob1234567/kda-graft-econ-evidence-2026-09-12.sma-evidence-graph
SMA Evidence Graph
An open-source, evidence-first dataset for Spinal Muscular Atrophy (SMA) drug research.
Description
This dataset contains structured evidence extracted from PubMed papers, clinical trials
from ClinicalTrials.gov, computationally generated hypotheses, AI-designed molecules,
and DiffDock molecular docking results — all linking gene targets to potential
therapeutic interventions for SMA.
Built by a researcher who has SMA, this dataset aims to accelerate… See the full description on the dataset page: https://huggingface.co/datasets/SMAResearch/sma-evidence-graph.automo-kd-qer-evidencehotpot-fine-grain-evidenceopenp2p-timing-evidence-3000-20260911livegraph-paper-evidence
LiveGraph Paper Evidence
This directory preserves the raw predictions, statistical summaries, audit
outputs, generated paper provenance, and final paper snapshot used by the
LiveGraph release.
The files are intentionally kept byte-identical to the frozen experiment
outputs. Some JSON metadata therefore contains inert absolute run paths or
machine labels. These are provenance strings, not credentials or runnable
remote endpoints. REPRODUCIBILITY_INDEX.json records the indexed… See the full description on the dataset page: https://huggingface.co/datasets/nutshells3/livegraph-paper-evidence.Qwen3.8-27B-GSQ-RCO-scale-recovery-evidence
Qwen3.8-27B GSQ-RCO IQ3_S: scale-recovery evidence and code
Research evidence dataset. No model weights. Part of the collection
Xyntetik Research: Model Surgery and Scale Recovery on this account, produced with
Xyntetik Runner.
Dataset summary
Question tested. Whether retraining only the fp16 block scales of an existing IQ3_S file (integer codes untouched) by distillation against the BF16 parent brings it inside the house bar, and what it does on a public… See the full description on the dataset page: https://huggingface.co/datasets/Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-scale-recovery-evidence.ST-Evidence-Instruct
ST-Evidence-Instruct Dataset
This dataset contains spatiotemporal evidence-based video question answering data for training.
This dataset was generated using Gemini and should not be used to develop models that compete with Google.
This project also uses the Segment Anything Model 3 (SAM 3) distributed by Meta Platforms, Inc. Use of SAM 3 is subject to the SAM License.
This was released for research purposes only, in support of the academic paper Evidence-Backed Video Question… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ST-Evidence-Instruct.radgenome-anatomy
RadGenome-Anatomy
RadGenome-Anatomy is a large-scale chest radiograph anatomy segmentation dataset
constructed from the RadGenome-ChestCT corpus
(originally based on CT-RATE).
It contains 25,692 volumetric studies (24,128 train / 1,564 validation), yielding paired
postero-anterior (PA) and lateral (LL) projection images at 384 × 384 resolution.
Across the two radiographic views, the dataset provides 10,790,646 fine-grained anatomy masks
over 210 canonical anatomy classes and 513… See the full description on the dataset page: https://huggingface.co/datasets/EvidenceAIResearch/radgenome-anatomy.sno-article-1-evidence
SNO Article I Evidence
Machine-readable experimental records supporting Watching Training Move:
Causal Forecasting of Neural Updates Across 8,192 SGD Transitions.
Creator: TiGa-RCEAffiliation: Independent researcher; founder and sole operator of Brewster Jennings
The Dataset contains original result records, figure-data contracts, and
machine-readable receipts. It does not contain the article prose, the rendered
figures, the source code, private conversations, local system paths… See the full description on the dataset page: https://huggingface.co/datasets/TiGa-RCE/sno-article-1-evidence.MIMIC-CXR-VReason
VReason MIMIC-CXR
A chest radiograph report-generation dataset augmented with structured
visual reasoning traces and region-of-interest (ROI) crops.
Each example walks through the radiologist's interpretation workflow
section by section before producing the final report.
Dataset at a glance
Split
Examples
train
100,750
validation
777
test
1,138
Source
Derived from MIMIC-CXR (Johnson
et al., 2019). Anatomical and pathological ROI… See the full description on the dataset page: https://huggingface.co/datasets/EvidenceAIResearch/MIMIC-CXR-VReason.cryocare-v0.3.0-parity-evidence
cryoCARE v0.3 CPU parity evidence
This dataset contains only JSON and Safetensors. It preserves deterministic inputs,
the exact extracted Keras-layout tensor inventory, vendor TensorFlow outputs, native
PyTorch outputs, and numerical metrics. The original HDF5 is not uploaded; its exact
hash and size remain in vendor/vendor-capture.json and the model conversion record.
The corresponding native Safetensors package is
scitomo/cryocare-v0.3.0-synthetic-parity.
SLAKE-VReason
VReason SLAKE
A bilingual (English / Chinese) medical visual question-answering dataset
derived from SLAKE, augmented with step-by-step visual reasoning
traces. Questions span multiple imaging modalities (MRI, CT, X-Ray) and
anatomical regions, covering both open-ended and closed-ended answer types.
Dataset at a glance
Split
Examples
train
4,919
validation
1,053
test
1,061
Source
Derived from SLAKE (Liu et al., 2021).… See the full description on the dataset page: https://huggingface.co/datasets/EvidenceAIResearch/SLAKE-VReason.responsible-disclosure-evidence-index
Responsible Disclosure Evidence Index
This is a public-safe responsible-disclosure lane. It records our rules and links to sanitized disclosure records and non-actionable commitment notes. The current default is commitment first: when retained evidence exists, a non-actionable public commitment record gives the work a visible timestamp while technical details stay private. It does not publish exploit steps, private emails, source paths, reproduction code, raw logs, or unresolved… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/responsible-disclosure-evidence-index.unit-price-evidence-synthetic
Unit Price Evidence: Synthetic
This dataset contains rendered synthetic shopping pages and evidence-pointer
targets for product-card discovery and unit-price field extraction. It was built
to warm-start small encoder-decoder models without redistributing retailer HTML,
screenshots, product data, account data, or browsing history.
Release
Version: 0.1.0
Source code: erichasinternet/apples-to-apples
Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.atemokoloporos-qwen3.5-0.8b-study-evidence
Teaching one synthetic fact to Qwen3.5-0.8B — evidence archive
This dataset-style repository is a publication bundle for a completed research
record, not a benchmark or training-data release. It preserves the full sourced
retrospective, immutable manifest, all eight evaluation pairs, all nine concise
and detailed run reports, the authoring disclosure, and the derived paper PDF.
The study attempted the exact synthetic fact Atemokoloporos is a rainbow unicorn. using… See the full description on the dataset page: https://huggingface.co/datasets/BurnyCoder/atemokoloporos-qwen3.5-0.8b-study-evidence.cyber-evidence-dataset
Cyber Security Evidence Dataset — MITRE ATT&CK Safe Reference
This repository contains a narrow, title-only reference configuration derived from the official MITRE ATT&CK Enterprise v19.2 STIX data. It is published separately from the broader Cyber Security Evidence Dataset project because the CISA-derived evidence layers remain private and are not included here.
Scope
The release contains 697 deterministic records. Each record provides an ATT&CK technique… See the full description on the dataset page: https://huggingface.co/datasets/frangelbarrera/cyber-evidence-dataset.crop-paired-visual-evidence
CROP: Paired Visual Evidence Dataset
Version 2.0.0 — exact alignment to the official positive teacher.
CROP contains 6,241 pairs of positive and negative teacher images for fine-grained visual question answering and multimodal distillation. Every positive PNG is preserved byte-for-byte from Vision-OPD-6K. Each negative image is generated from a displaced region of the same original photograph, using its paired positive's recovered crop dimensions, target coordinates within the… See the full description on the dataset page: https://huggingface.co/datasets/haokaixinmeitiandouhaokaixin/crop-paired-visual-evidence.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.contextualized-ST-Evidence
Contextualized ST-Evidence
A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask
split. Same 19,902 entries, same objects, same frames, same temporal evidence.
The only thing that changes is the spatial box on each frame.
This is the video counterpart of
shredder-31/contextualized-viscot,
built with the same model, the same prompt design and the same union-with-the-
original safety rule.
Why
ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.toolathlon-ml-release-rollback-evidence
Toolathlon R-27 artifact evidence
This synthetic dataset contains artifact identities and compatibility metadata for
a bounded ML release-governance benchmark. It contains no weights or customer data.
The JSON records under evidence/components/ are authoritative.
