datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.MTS_Dialogue-Clinical_Note
MTS Dialogue (Clinical Note Summarisation)
Main Dataset
The MTS-Dialog dataset is a new collection of 1.7k short doctor-patient conversations and corresponding summaries (section headers and contents).
The training set consists of 1,201 pairs of conversations and associated summaries.
The validation set consists of 100 pairs of conversations and their summaries.
The "dialogue" column contain Doctor-Patient conversation. The "section_text" column contains the Clinical Note of the… See the full description on the dataset page: https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note.clinical-trials-xml-2018-2024real_clinical_cases_of_Famous_Old_TCM_Doctors
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors 数据集简介
TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors是一个包含了当代著名老中医临床病例的数据集。这些病例数据来源于《当代名老中医典型医案集》(Contemporary Famous Old Chinese Medicine Doctors' Typical Cases Collection)一书。该数据集收录了多位德高望重的老中医大家的真实门诊病历,涵盖了多种常见病和疑难杂症。每个病例都包括病情描述、辨证论治思路、具体治疗方药等宝贵的一手临床资料。这些医案凝聚了老一辈名医的智慧和经验,对于中医的传承发展和临床应用研究,都有重要价值。通过对这些案例的挖掘分析,能够总结老中医诊疗思维、理法方药的特点,为现代中医临床实践提供有益借鉴。
Introduction to TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/real_clinical_cases_of_Famous_Old_TCM_Doctors.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024ICD10-Clinical-TerminologyICD10-Clinical-Terminology
pyarrow fast search demonstration for context AI MMoE
ClinicalTrial-gov_QAcsa-clinical-stage-asset-intelligence-sample
CSA — Clinical-Stage Asset Intelligence · Free Sample
Clinical trials, FDA, and SEC — linked to the drug asset and the listed sponsor, with a
forward catalyst calendar. This is a free 150-row sample of the nearest-term
catalysts; the full snapshot carries 2,221 forward catalysts (955 linked to
124 listed sponsors) and 1,890 resolved assets.
Data, not investment advice. CSA is information, not a recommendation to buy, sell,
or hold any security. Estimated catalyst dates (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/csa-clinical-stage-asset-intelligence-sample.LOINC-Clinical-Terminologyclinical_evidence
OpenTargets Clinical Evidence Dataset
This OpenTargets clinical evidence dataset represents a comprehensive collection of clinical trial data linking genetic targets to diseases, containing 32 structured columns that capture the complete lifecycle of clinical studies.
The dataset is annotated with clinical trials' predicted stop reason and genetic evidence when available and was referenced in the Why Clinical Trials Stop: The Role of Genetics.
The study analyzed 28,842 stopped… See the full description on the dataset page: https://huggingface.co/datasets/opentargets/clinical_evidence.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun.
Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection
Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records.
Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.china-myeloma-clinical-trials
China Multiple Myeloma Clinical Trials — Open Dataset
Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation.
This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinical-trials; the documentation is at chinamyeloma.org.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/chinamyeloma/china-myeloma-clinical-trials.clinical-trial-outcomes-2020plus
Clinical Trial Outcomes (2020+) with Normalized Endpoints
124,790 normalized endpoints across 14,170 clinical studies that started on or after
2020-01-01 and have posted results on ClinicalTrials.gov.
Snapshot: 2026-09-08. Source: ClinicalTrials.gov API v2 (U.S. National Library of Medicine).
Built with ctgov — the same normalizer, released as an
MIT-licensed package with zero dependencies. So this snapshot is not a dead artifact: you can
re-run it against the live registry, or… See the full description on the dataset page: https://huggingface.co/datasets/GooseWithStories/clinical-trial-outcomes-2020plus.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6
What this repo does
This repository contains a Clarus v0.6 intervention pathway dataset focused on respiratory collapse dynamics.
The dataset evaluates whether a model can determine if a proposed intervention meaningfully stabilizes a deteriorating respiratory system.
The task requires reasoning from:
system state
trajectory toward instability
boundary geometry
recovery geometry
intervention vector
projected trajectory consequence
The model cannot read the answer directly.
It must… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-oxygen-demand-buffer-lag-coupling-respiratory-collapse-v0.6.Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun.
Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files.
Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0
What this repo does
This repository provides a Clarus v1.0 benchmark for postoperative collapse under a four-variable clinical quad:
surgical_stress
buffer_capacity
lag_burden
coupling_stress
The v1.0 upgrade is Closed-Loop Control Geometry.
The task is no longer limited to detecting deterioration or ranking one intervention against another.
It tests whether a controller can:
choose the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-surgical-stress-buffer-lag-coupling-postop-collapse-v1.0.clinical-drv-atlas-perturbation-response-stability-mapping-v0.1What this dataset tests
Whether a model can classify response topologyafter a controlled perturbation.
It rewards
correct topology
recognition of cross-system coupling
recovery timing
Response topologies
rapid_return
delayed_recovery
overshoot_instability
oscillatory_instability
collapse
Typical failures
confusing overshoot with oscillation
ignoring coupling direction
calling delayed recovery stable
Suggested prompt wrapper
System
You map perturbation response… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-perturbation-response-stability-mapping-v0.1.clinical-prescription-pharmacy-dispense-coherence-risk-v0.1What this repo is for
Detect when
a prescription exists
but pharmacy dispense
does not happen in time
Common breaks
stockout
verification delay
clarification needed
queue delay for discharge meds
Examples you can use
urgent anticoagulant delayed
antibiotic not dispensed due to stockout
TTO delayed so discharge stalls
You use it to flag
missed dose risk
discharge delay risk
clinical_trial_eligibility_crietria_recommendationThis repository is a public repository of the data used in the paper "CReSE: Enhancing Clinical Trial Design via Contrastive Learning and Rephrasing-based and Clinical Relevance-preserving Sentence Embedding" (under review).
There are three main types of data stored in the repository.
Positive-negative EC-title pairs: A dataset that pairs the ECs used in a study with the study's title and other design information. It can be used to train EC recommendation models (binary classification).… See the full description on the dataset page: https://huggingface.co/datasets/kimsiun/clinical_trial_eligibility_crietria_recommendation.clinical-meaning-integration-fragmentation-analysis-v0.1What this dataset tests
Whether an intelligence system can evaluatethe integrity of a patient’s meaning-making processduring illness and recovery.
Required outputs
coherence integrity score
fragmentation markers
denial indicators
adaptive reframing presence
narrative stability index
meaning failure mode
Use case
Second layer of the Healing Narrative Coherence Corpus.
TCGA-Cancer-Variant-and-Clinical-Data
TCGA Cancer Variant and Clinical Data
Dataset Description
This dataset combines genetic variant information at the protein level with clinical data from The Cancer Genome Atlas (TCGA) project, curated by the International Cancer Genome Consortium (ICGC). It provides a comprehensive view of protein-altering mutations and clinical characteristics across various cancer types.
Dataset Summary
The dataset includes:
Protein sequence data for both mutated and… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/TCGA-Cancer-Variant-and-Clinical-Data.nyc-clinic-ai-infrastructureNYC Clinic AI Infrastructure
Broadband, electricity, and grid reliability data for all 311 NYC ZIP codes: can a clinic run local AI?
Overview
NYC Clinic AI Infrastructure maps three infrastructure prerequisites for on-premise AI deployment
across all 311 New York City ZIP codes. Each row covers one ZIP code with fixed broadband
subscription rates from the Census ACS, ISP coverage and max speeds from the FCC National
Broadband Map… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/nyc-clinic-ai-infrastructure.clinical-imaging-report-action-coherence-risk-v0.1What this repo is for
Detect when
imaging answers a question
but the system fails to
acknowledge and act
Common breaks
critical report not acknowledged
report late for urgent indication
action delayed after critical finding
negative result not integrated
Examples you can use
CTPA positive but anticoag not started
CT head bleed not escalated
CT perforation with delayed surgery
You use it to flag
missed diagnosis risk
treatment delay risk
clinic150-surdataset_info:
features:
name: intent
dtype: string
name: user_utterance
dtype: string
name: origin
dtype: string
Dataset Card for "clinic150-SUR"
Dataset Summary
The Clinic150-SUR dataset is a novel and augmented dataset designed to simulate natural human behavior during interactions with customer service-like centers.
Extending the Clinic150 dataset, it incorporates two augmentation techniques, including IBM's LAMBADA and Parrot models and carefully curated… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/clinic150-sur.clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.
