datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drug-screening-qa
Workplace Drug Screening Q&A
A question-answering dataset for workplace drug screening — federal drug testing
regulations, specimen collection, laboratory methodology, testing panels, and Medical
Review Officer (MRO) procedures. Built to fine-tune a small instruction model into a
domain assistant for HR professionals, employers, occupational health staff, and MROs.
Contents
File
Rows
Purpose
train.jsonl
639
Training split
val.jsonl
71
Validation split… See the full description on the dataset page: https://huggingface.co/datasets/oikyoni/drug-screening-qa.cancer-screening-evidence-reasoner
Cancer Screening Evidence Reasoner (AutoScientist Challenge)
Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses.
Motivation
Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.
