CoolFace
20 results

evidence

AOSSIE /openverifiable-enwiki-20260901-20260918-r1-evidence OpenVerifiableLLM Wikipedia provenance evidence Development in progress. No production-trained or end-to-end verified model is published here yet. Synthetic test results do not establish Wikipedia training. This new AOSSIE repository is reserved for publicly reconstructible inputs, checkpoints and reports for the OpenVerifiableLLM Wikipedia base model and its conversational derivative. The governing goal and code are maintained at AOSSIE-Org/OpenVerifiableLLM. The intended… See the full description on the dataset page: https://huggingface.co/datasets/AOSSIE/openverifiable-enwiki-20260901-20260918-r1-evidence.0 likes2.8k downloads2h agoHugging Facecopenlu /fever_gold_evidence Dataset Card for fever_gold_evidence Dataset Summary Dataset for training classification-only fact checking with claims from the FEVER dataset. This dataset is used in the paper "Generating Label Cohesive and Well-Formed Adversarial Claims", EMNLP 2020 The evidence is the gold evidence from the FEVER dataset for REFUTE and SUPPORT claims. For NEI claims, we extract evidence sentences with the system in "Christopher Malon. 2018. Team Papelo: Transformer Networks at FEVER.… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/fever_gold_evidence.texttext-classification100K<n<1M13 likes2.8k downloads4y agoHugging FaceXamitK /gero-research-evidence-2026-09 GERO research evidence — 120 publications This dataset contains 120 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 119 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the actuarialmath continuous-annuity selection-duration audit, with verified developer issue7 and explicit limitations. Report counts are not independent-defect counts. Latest numerical… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.texttext-retrievaln<1K0 likes1.7k downloads11h agoHugging Facetsinghua-sigs-robot-lab /VeriLoop-Coder-E1-Evaluation-Evidence VeriLoop Coder-E1 Evaluation Evidence This repository contains the public evaluation-evidence packages referenced by the official VeriLoop Coder-E1 benchmark result files. Model repository: tsinghua-sigs-robot-lab/veriloop-coder-e1 Evidence packages Benchmark Evidence directory DeepSWE veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0 SWE-bench Pro veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0 SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.0 likes979 downloads2mo agoHugging Facesuperlinked /sie-task-evidence SIE task evidence The recorded inputs and model responses behind the task pages on superlinked.com, one folder per task. Every figure published on a task page was produced by a real recorded run against https://api.superlinked.com. This dataset holds those recordings so that anyone can re-derive the published numbers without an API key and without spending any inference. How it is used The runnable example for each task lives in the public superlinked/sie… See the full description on the dataset page: https://huggingface.co/datasets/superlinked/sie-task-evidence.imagen<1K0 likes967 downloads20h agoHugging Facejaydeyoung /evidence_infer_treatmentData and code from our "Inferring Which Medical Treatments Work from Reports of Clinical Trials", NAACL 2019. This work concerns inferring the results reported in clinical trials from text. The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts' associated with them. These prompts will ask about the relationship between an intervention and comparator with respect to an outcome, as reported in the trial. For example, a prompt may ask about the reported effects of aspirin as compared to placebo on the duration of headaches. For the sake of this task, we assume that a particular article will report that the intervention of interest either significantly increased, significantly decreased or had significant effect on the outcome, relative to the comparator. The dataset could be used for automatic data extraction of the results of a given RCT. This would enable readers to discover the effectiveness of different treatments without needing to read the paper.text-retrieval1K<n<10K8 likes723 downloads3y agoHugging Face