datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.clinicalbr
ClinicalBr
ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark
built from 2,892 real Brazilian case reports drawn from 36 open-access
medical journals spanning 18 specialties. Every case is provided as a parallel
PT/EN pair and supports four evaluation tasks.
Please refer to the paper for full details on the tasks, methodology, and limitations.
Tasks & metrics
Task
Config
n / lang
Metric
Diagnosis retrieval
diagnosis
2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.clinical-case-icd10-diagnosis
Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched
1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status.
input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present).
output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.clinicaltrial-protocol-corpus
Clinical Trial Protocol Corpus
Full-text clinical trial protocol documents from ClinicalTrials.gov with section segmentation aligned to SPIRIT/ICH-GCP categories.
What this is
49,002 protocol PDFs downloaded from the ClinicalTrials.gov CDN, extracted to text via PyMuPDF, and segmented into structured sections. Each record contains the full protocol text plus a list of detected sections with headings, hierarchy levels, and section type labels from a 15-type… See the full description on the dataset page: https://huggingface.co/datasets/JulesCan/clinicaltrial-protocol-corpus.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.task685_mmmlu_answer_generation_clinical_knowledge
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.Clinically_Informed_Synthetic_Population_v1.0
Clinically-Informed Synthetic Patient Populations
(October 2025 Edition — 100 | 1 000 | 10 000 Patients)
Overview
These datasets represent fully synthetic, privacy-free populations that emulate the structure and statistical behavior of modern hospital data as of October 2025.
Each population includes four inter-linked tables describing demographics, admissions, diagnoses, and laboratory measurements.
No real patient information is used or referenced… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Clinically_Informed_Synthetic_Population_v1.0.clinical-summarizer-sft
Clinical Note Summarizer (SOAP Format)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Long clinical notes → structured SOAP summaries
Why download this
Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.psiddx-clinical-ddx
PsiDDx — Clinical Differential-Diagnosis Dataset (v5.2)
A decontaminated, calibrated differential-diagnosis training corpus for the Adaption Labs
AutoScientist Challenge (Healthcare). Each record pairs a patient presentation with a
ranked top-5 differential — common explanations first, each with an ICD-10 code, a calibrated
confidence, and the discriminating feature; rare diagnoses are kept as low-confidence
must-not-miss entries rather than over-called.
5,482 rows. Used to… See the full description on the dataset page: https://huggingface.co/datasets/shariqazeem/psiddx-clinical-ddx.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.clinical-trials-embeddings
Clinical Trials Embeddings Dataset
Overview
This dataset contains information extracted from clinical trial records collected from ClinicalTrials.gov (Date Accessed: 05/02/2025) along with briefSummary columns embeddings generated using minishlab/potion-base-8M. It focuses on key descriptive fields that provide insight into trial objectives, eligibility criteria, and study design. The dataset is designed for researchers, healthcare professionals, and AI/ML practitioners… See the full description on the dataset page: https://huggingface.co/datasets/cyrilzakka/clinical-trials-embeddings.clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.synthetic_clinical_conversations
Synthetic Clinical Conversations
Fully synthetic English clinical conversations (care-coordination calls, telehealth
visits, post-discharge check-ins) paired with structured encounter records, built for
training and evaluating transcript→JSON extraction models.
Generated structure-first with Tonic Fabricate: the
structured facts are authored as relational data with controlled vocabularies, the
conversation is rendered from those facts, and the extraction target is a… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/synthetic_clinical_conversations.pocketgull-nih-who-clinical-dpo
📚 PocketGull NIH & WHO Clinical Preference DPO Dataset
Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514
📌 Dataset Summary
Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.clinical-notes-to-fhir
SGRS-FHIR: A Preference Learning Dataset for Clinical FHIR Extraction
The first preference learning dataset for clinical FHIR extraction with structured error paths.
Key Insight
Structured extraction failures are training signal, not noise.
Unlike traditional datasets that discard generation failures, SGRS-FHIR intentionally captures both valid and invalid extractions with detailed error annotations. This enables:
Direct Preference Optimization (DPO): Train models to… See the full description on the dataset page: https://huggingface.co/datasets/ai-galileo/clinical-notes-to-fhir.indian_protocols_based_clinical_QnA
Indian Protocols-Based Clinical Q&A
A rubric-graded evaluation dataset built from clinical guideline documents (Indian and international). Each sample is a realistic doctor-side query against a known protocol, paired with rubrics that grade (a) whether the system retrieved/identified the correct guideline content and (b) whether the final answer is clinically complete and safe.
What this evaluates
This dataset is built to stress-test clinical assistants on… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/indian_protocols_based_clinical_QnA.ALIA-es-clinical-psychology-dialogues
[!WARNING]
DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation.
Dataset Introduction
The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.clinical_note_generation_dataset
Clinical Note Generation Dataset
Dataset Description
The Eka Structured Clinical Note Generation Dataset facilitates evaluation of medical scribe systems capable of transforming transcribed medical conversations into structured, entity-level medical records. This dataset addresses one of the most challenging aspects of healthcare AI: understanding and organising complex medical information into structured formats.
Dataset Composition and Clinical Relevance… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/clinical_note_generation_dataset.pocketgull-clinical-instruction-corpus
📚 PocketGull Multi-Paradigm Clinical Instruction Corpus
Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514
📌 Dataset Summary
Multi-turn clinical SFT training instructions spanning stepped-care triage, ambient SOAP scribing, pharmacogenomics CYP450 interactions, and… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-clinical-instruction-corpus.OncoAgent-Clinical-266K
🧬 OncoAgent Clinical Dataset — 266K
Curated Multi-Source Oncology Training Dataset
AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0
Dataset Description
This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks.
Composition
Source
Samples
Description
PMC-Patients
~100,000
Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/MaximoLopezChenlo/OncoAgent-Clinical-266K.clinical-parallel-process-awareness-v0.1Clinical Parallel Process Awareness v0.1
Goal
Test if a model can hold separate reasoning streams at once
Detect constraint dismissal
Detect bleed-over where one stream turns into claims in the other
What it measures
streams_heldResponse acknowledges and maintains both streams
bleed_overConstraint stream improperly becomes a medical claim, or vice versa
premature_synthesisResponse forces a single solution that silences one stream
assumption_collapseResponse drops a premise entirely
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-parallel-process-awareness-v0.1.clinical-anamnesis-fidelity-v0.1Clinical Anamnesis Fidelity v0.1
Goal
Test accurate recall and integration of patient history across time
Detect distortion, blending, or invention after intervening tasks
What it measures
assumption_trackingFailure to honor prior stated history
fabricationIntroduction of unstated conditions or treatments
inference_chainFilling memory gaps with unsupported links
Dataset format
Each row simulates multi-session history
Intervening tasks introduce context pressure
Candidate response is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-anamnesis-fidelity-v0.1.clinical-perception-intervention-justification-v0.1Clinical Perception–Intervention Justification v0.1
Goal
Test whether actions follow directly from perceptual evidence
Detect interventions that appear without a visual cause
Detect escalation that exceeds image-supported severity
What it measures
action_without_causeAn intervention is proposed with no supporting image evidence
over_escalationThe action exceeds what the visual severity supports
justification_okThe response links perception to action explicitly or proportionally
How it… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-perception-intervention-justification-v0.1.
