datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cochrane-screening-sft
Cochrane Screening SFT
Supervised fine-tuning (SFT) chat dataset for Cochrane-style title and abstract screening.
Each example is a chat conversation that asks a model to predict a screening decision
(include / exclude / uncertain) and a short justification (reason).
Code: ljwa2323/cochrane-screening-slm
Dataset summary
Split / config
Records
Role
train
416,799
LoRA SFT training
validation
46,311
Training-time validation (10% stratified holdout from… See the full description on the dataset page: https://huggingface.co/datasets/deepcoder2024/cochrane-screening-sft.drug-screening-qa
Workplace Drug Screening Q&A
A question-answering dataset for workplace drug screening — federal drug testing
regulations, specimen collection, laboratory methodology, testing panels, and Medical
Review Officer (MRO) procedures. Built to fine-tune a small instruction model into a
domain assistant for HR professionals, employers, occupational health staff, and MROs.
Contents
File
Rows
Purpose
train.jsonl
639
Training split
val.jsonl
71
Validation split… See the full description on the dataset page: https://huggingface.co/datasets/oikyoni/drug-screening-qa.sanctions-screening-match
SANCTIONS_SCREENING_MATCH
A preference dataset for SANCTIONS_SCREENING_MATCH, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from… See the full description on the dataset page: https://huggingface.co/datasets/316usman/sanctions-screening-match.clinical-quad-enrollment-criteria-drift-site-selection-bias-screening-pressure-v0.1Clarus Clinical Quad Coupling Enrollment Criteria Drift Site Selection Bias Screening Pressure v0.1
PurposeDetect enrollment population drift driven by four interacting nodes.
Quad nodes
Criteria relaxation or documentation gap
Site selection or recruitment bias
Screening workflow pressure
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
enrollment_drift_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-criteria-drift-site-selection-bias-screening-pressure-v0.1.cancer-screening-evidence-reasoner
Cancer Screening Evidence Reasoner (AutoScientist Challenge)
Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses.
Motivation
Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.
