datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.PMC-Clinical
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Clinical.IIYi-Clinical
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/IIYi-Clinical.ClinicalExtract-Datasetmedical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.oac-clinical-transport-observability-synthetic
OAC Clinical Transport Observability — Synthetic
This dataset contains 1,200 fixed-seed, entirely synthetic operational
transport-health examples for the companion OAC System Health v1 model.
It contains no records collected from a patient, laboratory, analyzer,
instrument, LIS, EHR, network, or health-care site.
Companion model: OAC System Health v1.
Canonical source: szl-forge clinical gateway.
Data boundary
The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.icd10-clinical-notes
ICD-10 Multilingual Clinical Notes Dataset
A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages.
Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University
Dataset Description
This dataset provides ICD-10 codes with:
Official diagnosis names in 34 languages (24 EU + 10 major world languages)
Sample clinical journal notes (English and Swedish)
Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/icd10-clinical-notes.CIEL-Clinical-Concepts-to-ICD-11TD;LR Run:
pip install torch==2.4.1+cu118 torchvision==0.19.2+cu118 torchaudio==2.4.1 --extra-index-url https://download.pytorch.org/whl/cu118
pip install -U packaging setuptools wheel ninja
pip install --no-build-isolation axolotl[flash-attn,deepspeed]
axolotl train axolotl_2_a40_runpod_config.yaml
📚 CIEL to ICD-11 Fine-tuning Dataset
This dataset was created to support the fine-tuning of open-source large language models (LLMs) specialized in ICD-11 terminology mapping.
It… See the full description on the dataset page: https://huggingface.co/datasets/filipelopesmedbr/CIEL-Clinical-Concepts-to-ICD-11.pentabrid-reproducibility
Pentabrid 27B: reproducibility package
Everything required to recompute the results of a controlled evaluation of fine-tuning
configurations for medical question answering. Openly available with no access
restrictions.
Contents
Path
Description
per_item/medxpertqa_*.jsonl
Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.clinical-persian-qa-ii
Clinical Question Answering Dataset II (Farsi)
This dataset contains more than 211k questions and more than 700k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset is NOT a part of Clinical Question Answering I dataset and is a whole… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-ii.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.clinical-rag-safety-gateway-20260904-dataset
Clinical RAG Safety Gateway Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260904-dataset.psiddx-clinical-ddx
PsiDDx — Clinical Differential-Diagnosis Dataset (v5.2)
A decontaminated, calibrated differential-diagnosis training corpus for the Adaption Labs
AutoScientist Challenge (Healthcare). Each record pairs a patient presentation with a
ranked top-5 differential — common explanations first, each with an ICD-10 code, a calibrated
confidence, and the discriminating feature; rare diagnoses are kept as low-confidence
must-not-miss entries rather than over-called.
5,482 rows. Used to… See the full description on the dataset page: https://huggingface.co/datasets/shariqazeem/psiddx-clinical-ddx.clinical-summarizer-sft
Clinical Note Summarizer (SOAP Format)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Long clinical notes → structured SOAP summaries
Why download this
Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.ALIA-es-clinical-psychology-dialogues
[!WARNING]
DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation.
Dataset Introduction
The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.pocketgull-nih-who-clinical-dpo
📚 PocketGull NIH & WHO Clinical Preference DPO Dataset
Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514
📌 Dataset Summary
Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.clinical_trial_reason_to_stop
Dataset Card for Clinical Trials's Reason to Stop
Dataset Summary
This dataset contains a curated classification of more than 5000 reasons why a clinical trial has suffered an early stop.
The text has been extracted from clinicaltrials.gov, the largest resource of clinical trial information. The text has been curated by members of the Open Targets organisation, a project aimed at providing data relevant to drug development.
All 17 possible classes have been carefully… See the full description on the dataset page: https://huggingface.co/datasets/opentargets/clinical_trial_reason_to_stop.clinical-persian-qa-i
Clinical Question Answering Dataset I (Farsi)
This dataset contains approximately 50k questions and around 60k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset is NOT a part of Clinical Question Answering II dataset and is a complete… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-i.clinical-rag-safety-gateway-20260914-dataset
Clinical RAG Safety Gateway Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260914-dataset.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.clinicguide-faq-corpus
ClinicGuide FAQ corpus
Synthetic, generic clinic-style FAQs for a medical-tourism pre-consultation assistant. The corpus is not copied from a named hospital and is not a substitute for a real clinic's published policies.
Languages: English, Arabic (MSA), Persian.
Intended use
Retrieval-augmented answers for visa, stay, companion, hotel, airport transfer, starting-from costs, documents, booking, and “what to ask the doctor”
Safety / refusal examples: diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/clinicguide-faq-corpus.icd10-clinical-notes
ICD-10 Multilingual Clinical Notes Dataset
A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages.
Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University
Dataset Description
This dataset provides ICD-10 codes with:
Official diagnosis names in 34 languages (24 EU + 10 major world languages)
Sample clinical journal notes (English and Swedish)
Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/Danishaqil/icd10-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.adaption-clinical-triage-preferences
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-clinical_triage_preferences
Multi-turn conversational preference dataset designed for fine-grained safety and tone calibration in emergency first aid and symptom triage. Each sample pairs a user prompt with chosen and rejected AI responses, contrasting concise, grounded clinical guidance against subtly misleading or overly verbose advice. It supports reward modeling and preference… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-clinical-triage-preferences.clinical-trial🏥 Dataset Summary
Structured clinical trial records sourced from ClinicalTrials.gov. Each record extracts and normalizes key parameters including study purpose, intervention arms, trial conditions, and sponsor/collaborator relationships. Drug entities are enriched with multilingual name mappings (Chinese, English, Japanese).
🚀 Key Features
Structured Intervention Arms: arm_intervention captures each trial arm's design group type, linked interventions, and single/multi-drug configuration.… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/clinical-trial.clinical-notes-to-fhir
SGRS-FHIR: A Preference Learning Dataset for Clinical FHIR Extraction
The first preference learning dataset for clinical FHIR extraction with structured error paths.
Key Insight
Structured extraction failures are training signal, not noise.
Unlike traditional datasets that discard generation failures, SGRS-FHIR intentionally captures both valid and invalid extractions with detailed error annotations. This enables:
Direct Preference Optimization (DPO): Train models to… See the full description on the dataset page: https://huggingface.co/datasets/ai-galileo/clinical-notes-to-fhir.CMMLU-Clinical-Knowledge-Benchmark
💻 Dataset Usage
Run the following command to load the testing set (237 examples):
from datasets import load_dataset
dataset = load_dataset("shuyuej/CMMLU-Clinical-Knowledge-Benchmark", split="train")
print(dataset)
