datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeriLoop-E2-Evaluation-Evidence
VeriLoop E2 Evaluation Evidence
Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks.
This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.hemmingway-1-omlx-quantization-evidence-v2
Hemmingway-1 Quantization Evidence v2
This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.
This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.physiotherapy-evidence-qa
🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus
Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology.
This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.lord-of-mysteries-fandom-evidence-sft
Lord of Mysteries Fandom Evidence SFT Dataset
Overview
This dataset provides evidence-aware training and retrieval material for building a Chinese Lord of the Mysteries knowledge assistant.
The release is built from 425 Lord of the Mysteries Fandom Wiki pages. Source URLs, page titles, revision identifiers, and attribution metadata are preserved where available. The companion inference script can retrieve relevant source pages and attach exact source URLs before… See the full description on the dataset page: https://huggingface.co/datasets/xile42/lord-of-mysteries-fandom-evidence-sft.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.human-gene-lof-rescue-evidence
Human Biallelic Loss-of-Function and Functional-Rescue Evidence
This dataset contains 88 curated human gene records linking three experimentally distinct observations:
biallelic human loss of function;
a consistent phenotype reported in independent affected families or cohorts;
functional rescue in affected humans or patient-derived human cells.
Each record therefore connects genotype → recurrent human phenotype → reversal of a disease-relevant defect. This convergent evidence… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/human-gene-lof-rescue-evidence.AgroVeritas-Evidence-QA-Adapted
AgroVeritas Evidence QA
Agricultural intelligence you can audit — in English and Spanish.
AgroVeritas is a bilingual, evidence-bounded agricultural instruction dataset built for regional crop-calendar and historical climate reasoning. It teaches models to answer a practical question, cite the evidence used, show the reasoning path, state limitations, and recommend local verification instead of presenting historical data as live field conditions.
Dataset Viewer and… See the full description on the dataset page: https://huggingface.co/datasets/MarianaCodebase/AgroVeritas-Evidence-QA-Adapted.clinical_evidence_coherence_breakdown_v0.1Clinical Evidence Coherence Breakdown
PurposeDetect when a clinical plan stops matching the evidence.
You get evidence signals and a stated plan.You decide if a coherence break exists.You label the breakdown type.You propose the corrective action.
Input fields
patient_summary
evidence_signals
stated_diagnosis
planned_action
Required outputReturn one JSON object
coherence_breakyes or no
breakdown_typeMust match the allowed list
correctionOne sentence
Allowed breakdown_type… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_evidence_coherence_breakdown_v0.1.evidence-subagent-sft-gpt54-single-all-jina-v1
GPT-5.4 Evidence Subagent SFT with Jina-refreshed Browse Outputs
This dataset contains synthetic SFT conversations for training a small evidence
execution subagent for deep-research systems.
Each row is a single delegated evidence-gathering subtask derived from a full
DR-Tulu trajectory. GPT-5.4 synthesized the delegated subtask, grouped original
tool events into one or more batch tool-call turns, and wrote a structured
cited evidence report. Tool outputs are reconstructed from… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v1.cancer-screening-evidence-reasoner
Cancer Screening Evidence Reasoner (AutoScientist Challenge)
Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses.
Motivation
Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.clinical-evidence-conclusion-alignment-v0.1
What this dataset tests
Clinical conclusions must reflect evidence.
Language must track statistics.
Why it exists
Clinical papers drift at the conclusion.
Spin enters here.
This set detects misalignment between results and claims.
Data format
Each row contains
trial_result
conclusion_statement
alignment_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
trial_result
conclusion_statement
Score for… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-conclusion-alignment-v0.1.evidence-subagent-sft-gpt54-single-all-jina-v2-qwen35-thinking
Evidence Subagent SFT GPT-5.4 Jina v2, Qwen3.5 Thinking Aligned
This dataset is an aligned version of lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v2 for supervised fine-tuning a Qwen3.5 evidence subagent in LLaMA-Factory.
Splits
train: 10,379 examples
validation: 100 examples
Format
Each row contains:
id: source trajectory id
conversations: OpenAI-style messages with roles system, user, function, tool, and assistant
tools: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v2-qwen35-thinking.clinical-regulatory-evidence-correspondence-v0.1
What this dataset tests
Regulatory claims must map to evidence scope.
Population boundaries matter.
Why it exists
Regulatory language can drift.
Indications expand.
Subgroups disappear.
This set detects when claims exceed the evidence base.
Data format
Each row contains
evidence_base
regulatory_claim
correspondence_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
evidence_base
regulatory_claim
Score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-regulatory-evidence-correspondence-v0.1.patient-evidence-fidelity-v0.1
What this dataset tests
Patient language must preserve evidence.
Numbers matter.
Vagueness misleads.
Why it exists
Press releases simplify.
Meaning gets lost.
Patients deserve accuracy.
Data format
Each row contains
scientific_conclusion
patient_facing_statement
fidelity_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
scientific_conclusion
patient_facing_statement
Score for
preservation of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-evidence-fidelity-v0.1.pharmacoeconomic-evidence-extraction-dataset
Pharmacoeconomic Evidence Extraction Dataset
License
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You may use, share, and adapt the dataset provided that appropriate credit is given to the dataset authors.
For the full license terms, see the CC BY 4.0 license.
Overview
This dataset contains 250 expert-annotated records for research on automated extraction of structured pharmacoeconomic and… See the full description on the dataset page: https://huggingface.co/datasets/MJ16/pharmacoeconomic-evidence-extraction-dataset.
