evidence
Datasets
All datasets matching “evidence”openverifiable-enwiki-20260901-20260918-r1-evidence
OpenVerifiableLLM Wikipedia provenance evidence
Development in progress. No production-trained or end-to-end verified model is
published here yet. Synthetic test results do not establish Wikipedia training.
This new AOSSIE repository is reserved for publicly reconstructible inputs,
checkpoints and reports for the OpenVerifiableLLM Wikipedia base model and its
conversational derivative. The governing goal and code are maintained at
AOSSIE-Org/OpenVerifiableLLM.
The intended… See the full description on the dataset page: https://huggingface.co/datasets/AOSSIE/openverifiable-enwiki-20260901-20260918-r1-evidence.fever_gold_evidence
Dataset Card for fever_gold_evidence
Dataset Summary
Dataset for training classification-only fact checking with claims from the FEVER dataset.
This dataset is used in the paper "Generating Label Cohesive and Well-Formed Adversarial Claims", EMNLP 2020
The evidence is the gold evidence from the FEVER dataset for REFUTE and SUPPORT claims.
For NEI claims, we extract evidence sentences with the system in "Christopher Malon. 2018. Team Papelo: Transformer Networks at FEVER.… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/fever_gold_evidence.gero-research-evidence-2026-09
GERO research evidence — 120 publications
This dataset contains 120 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 119 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the actuarialmath continuous-annuity selection-duration audit, with verified developer issue7 and explicit limitations. Report counts are not independent-defect counts.
Latest numerical… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.VeriLoop-Coder-E1-Evaluation-Evidence
VeriLoop Coder-E1 Evaluation Evidence
This repository contains the public evaluation-evidence packages
referenced by the official VeriLoop Coder-E1 benchmark result files.
Model repository:
tsinghua-sigs-robot-lab/veriloop-coder-e1
Evidence packages
Benchmark
Evidence directory
DeepSWE
veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0
SWE-bench Pro
veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0
SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.sie-task-evidence
SIE task evidence
The recorded inputs and model responses behind the task pages on
superlinked.com, one folder per task.
Every figure published on a task page was produced by a real recorded run against
https://api.superlinked.com. This dataset holds those recordings so that anyone
can re-derive the published numbers without an API key and without spending any
inference.
How it is used
The runnable example for each task lives in the public
superlinked/sie… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/sie-task-evidence.evidence_infer_treatmentData and code from our "Inferring Which Medical Treatments Work from Reports of Clinical Trials", NAACL 2019. This work concerns inferring the results reported in clinical trials from text.
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts' associated with them. These prompts will ask about the relationship between an intervention and comparator with respect to an outcome, as reported in the trial. For example, a prompt may ask about the reported effects of aspirin as compared to placebo on the duration of headaches. For the sake of this task, we assume that a particular article will report that the intervention of interest either significantly increased, significantly decreased or had significant effect on the outcome, relative to the comparator.
The dataset could be used for automatic data extraction of the results of a given RCT. This would enable readers to discover the effectiveness of different treatments without needing to read the paper.
