datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cochrane-screening-sft
Cochrane Screening SFT
Supervised fine-tuning (SFT) chat dataset for Cochrane-style title and abstract screening.
Each example is a chat conversation that asks a model to predict a screening decision
(include / exclude / uncertain) and a short justification (reason).
Code: ljwa2323/cochrane-screening-slm
Dataset summary
Split / config
Records
Role
train
416,799
LoRA SFT training
validation
46,311
Training-time validation (10% stratified holdout from… See the full description on the dataset page: https://huggingface.co/datasets/deepcoder2024/cochrane-screening-sft.screening-ceiling
screening-ceiling
📖 Documentation site — the portfolio narrative, the concepts, a full walkthrough, and what all of this proves (and does not).
A machine-certified impossibility result about coupling extraction, plus the
concrete layouts where a plausible extractor predicts impossible physics.
Most ML-for-physics datasets are samples: here are some inputs, here are the
answers, fit something. This one is different in a way worth being precise
about. It carries a universal… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/screening-ceiling.drug-screening-qa
Workplace Drug Screening Q&A
A question-answering dataset for workplace drug screening — federal drug testing
regulations, specimen collection, laboratory methodology, testing panels, and Medical
Review Officer (MRO) procedures. Built to fine-tune a small instruction model into a
domain assistant for HR professionals, employers, occupational health staff, and MROs.
Contents
File
Rows
Purpose
train.jsonl
639
Training split
val.jsonl
71
Validation split… See the full description on the dataset page: https://huggingface.co/datasets/oikyoni/drug-screening-qa.siren-screening
SIREN Screening Dataset
~60,000 articles with synthetic relevance queries for training biomedical document screeners. Each article has multiple queries at three relevance levels: Relevant, Partial, Irrelevant.
Why this dataset?
Systematic reviews require screening thousands of articles against inclusion criteria (e.g., "RCTs in adults with diabetes, published after 2015"). Existing retrieval models (MedCPT, PubMedBERT)… See the full description on the dataset page: https://huggingface.co/datasets/Praise2112/siren-screening.cancer-screening-evidence-reasoner
Cancer Screening Evidence Reasoner (AutoScientist Challenge)
Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses.
Motivation
Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.prompt-screening-dataset
Prompt Screening Dataset
This dataset is designed for training a classifier that identifies desirable rows for AI model training.
Each data source in agentlans/chatgpt contributes 10 000 rows.
Every row has been evaluated using these models:
agentlans/bge-small-en-v1.5-prompt-safety
agentlans/bge-small-en-v1.5-prompt-quality
agentlans/bge-small-en-v1.5-prompt-difficulty
agentlans/snowflake-arctic-embed-xs-refusal-classifierThe abridged column concatenates the input and output… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-screening-dataset.nyc-diabetes-screenings
NYC Free Diabetes Screenings
A living directory of free diabetes screening locations and events across New York City, maintained by LocalDevs.
data/screenings.jsonl is rebuilt weekly by a scraper + geocoding pipeline that merges scraped sources (e.g. EmblemHealth's community events) with hand-maintained static entries (e.g. Columbia's Center for Community Health). Each record includes name, address, borough, lat/lon, hours, screening types offered, appointment type, and a… See the full description on the dataset page: https://huggingface.co/datasets/LocalDevs/nyc-diabetes-screenings.Updated_Sanction_Screening_Dataset
