CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K75 likes1.4k downloads3y agoHugging Face02R2MED /PMC-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/PMC-Clinical.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face03R2MED /IIYi-Clinical 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/IIYi-Clinical.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face04HeshikaPokala /ClinicalExtract-Datasettext10K<n<100K0 likes672 downloads25d agoHugging Face05stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes340 downloads2mo agoHugging Face06SZLHOLDINGS /oac-clinical-transport-observability-synthetic OAC Clinical Transport Observability — Synthetic This dataset contains 1,200 fixed-seed, entirely synthetic operational transport-health examples for the companion OAC System Health v1 model. It contains no records collected from a patient, laboratory, analyzer, instrument, LIS, EHR, network, or health-care site. Companion model: OAC System Health v1. Canonical source: szl-forge clinical gateway. Data boundary The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.texttabular-classification1K<n<10K0 likes338 downloads2d agoHugging Face07ritaranx /clinical-synthetic-text-kg Data Description We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models (ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs. Generated Datasets The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.texttext-classification1K<n<10K0 likes255 downloads2y agoHugging Face08ritaranx /clinical-synthetic-text-llm Data Description We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models (ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles. Generated Datasets The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.texttext-classification1K<n<10K3 likes191 downloads2y agoHugging Face09birgermoell /icd10-clinical-notes ICD-10 Multilingual Clinical Notes Dataset A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages. Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University Dataset Description This dataset provides ICD-10 codes with: Official diagnosis names in 34 languages (24 EU + 10 major world languages) Sample clinical journal notes (English and Swedish) Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/icd10-clinical-notes.texttext-classification1K<n<10K3 likes147 downloads8mo agoHugging Face10filipelopesmedbr /CIEL-Clinical-Concepts-to-ICD-11TD;LR Run: pip install torch==2.4.1+cu118 torchvision==0.19.2+cu118 torchaudio==2.4.1 --extra-index-url https://download.pytorch.org/whl/cu118 pip install -U packaging setuptools wheel ninja pip install --no-build-isolation axolotl[flash-attn,deepspeed] axolotl train axolotl_2_a40_runpod_config.yaml 📚 CIEL to ICD-11 Fine-tuning Dataset This dataset was created to support the fine-tuning of open-source large language models (LLMs) specialized in ICD-11 terminology mapping. It… See the full description on the dataset page: https://huggingface.co/datasets/filipelopesmedbr/CIEL-Clinical-Concepts-to-ICD-11.text100K<n<1M0 likes136 downloads1y agoHugging Face11Clinical-Reasoning-Hub /pentabrid-reproducibility Pentabrid 27B: reproducibility package Everything required to recompute the results of a controlled evaluation of fine-tuning configurations for medical question answering. Openly available with no access restrictions. Contents Path Description per_item/medxpertqa_*.jsonl Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.tabularquestion-answering10K<n<100K0 likes123 downloads7d agoHugging Face12PerSets /clinical-persian-qa-ii Clinical Question Answering Dataset II (Farsi) This dataset contains more than 211k questions and more than 700k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset is NOT a part of Clinical Question Answering I dataset and is a whole… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-ii.textquestion-answering100K<n<1M3 likes75 downloads1y agoHugging Face13Vinay393 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes75 downloads8mo agoHugging Face14johnny8808 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes75 downloads5mo agoHugging Face15RKB109 /clinical-rag-safety-gateway-20260904-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260904-dataset.textquestion-answeringn<1K0 likes75 downloads22d agoHugging Face16shariqazeem /psiddx-clinical-ddx PsiDDx — Clinical Differential-Diagnosis Dataset (v5.2) A decontaminated, calibrated differential-diagnosis training corpus for the Adaption Labs AutoScientist Challenge (Healthcare). Each record pairs a patient presentation with a ranked top-5 differential — common explanations first, each with an ICD-10 code, a calibrated confidence, and the discriminating feature; rare diagnoses are kept as low-confidence must-not-miss entries rather than over-called. 5,482 rows. Used to… See the full description on the dataset page: https://huggingface.co/datasets/shariqazeem/psiddx-clinical-ddx.texttext-generation1K<n<10K0 likes72 downloads3mo agoHugging Face17AmareshHebbar /clinical-summarizer-sft Clinical Note Summarizer (SOAP Format) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Long clinical notes → structured SOAP summaries Why download this Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.texttext-generation10K<n<100K0 likes72 downloads3mo agoHugging Face18SINAI /ALIA-es-clinical-psychology-dialogues [!WARNING] DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation. Dataset Introduction The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.texttext-generationn<1K0 likes71 downloads3mo agoHugging Face19philgear /pocketgull-nih-who-clinical-dpo 📚 PocketGull NIH & WHO Clinical Preference DPO Dataset Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514 📌 Dataset Summary Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.texttext-generationn<1K0 likes65 downloads25d agoHugging Face20opentargets /clinical_trial_reason_to_stop Dataset Card for Clinical Trials's Reason to Stop Dataset Summary This dataset contains a curated classification of more than 5000 reasons why a clinical trial has suffered an early stop. The text has been extracted from clinicaltrials.gov, the largest resource of clinical trial information. The text has been curated by members of the Open Targets organisation, a project aimed at providing data relevant to drug development. All 17 possible classes have been carefully… See the full description on the dataset page: https://huggingface.co/datasets/opentargets/clinical_trial_reason_to_stop.texttext-classification1K<n<10K13 likes63 downloads4y agoHugging Face21PerSets /clinical-persian-qa-i Clinical Question Answering Dataset I (Farsi) This dataset contains approximately 50k questions and around 60k answers, all produced in written form. The questions were posed by ordinary Persian speakers (Iranians), and the responses were provided by doctors from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset is NOT a part of Clinical Question Answering II dataset and is a complete… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/clinical-persian-qa-i.textquestion-answering10K<n<100K2 likes62 downloads1y agoHugging Face22RKB109 /clinical-rag-safety-gateway-20260914-dataset Clinical RAG Safety Gateway Synthetic Dataset Summary This dataset contains 14 training examples and 4 held-out examples for Clinical assistants need retrieval, source attribution, and explicit abstention before answers reach care teams. Every record is synthetic and includes: input: query, event, or feature description label: expected class, route, relation, or evidence category context: synthetic supporting context source: fictional source identifier variant:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/clinical-rag-safety-gateway-20260914-dataset.textquestion-answeringn<1K0 likes62 downloads11d agoHugging Face23minidiablo05 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes61 downloads6mo agoHugging Face24alirezaaminzadeh /clinicguide-faq-corpus ClinicGuide FAQ corpus Synthetic, generic clinic-style FAQs for a medical-tourism pre-consultation assistant. The corpus is not copied from a named hospital and is not a substitute for a real clinic's published policies. Languages: English, Arabic (MSA), Persian. Intended use Retrieval-augmented answers for visa, stay, companion, hotel, airport transfer, starting-from costs, documents, booking, and “what to ask the doctor” Safety / refusal examples: diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/clinicguide-faq-corpus.textquestion-answeringn<1K0 likes61 downloads1mo agoHugging Face25Danishaqil /icd10-clinical-notes ICD-10 Multilingual Clinical Notes Dataset A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages. Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University Dataset Description This dataset provides ICD-10 codes with: Official diagnosis names in 34 languages (24 EU + 10 major world languages) Sample clinical journal notes (English and Swedish) Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/Danishaqil/icd10-clinical-notes.texttext-classification1K<n<10K0 likes59 downloads6mo agoHugging Face26Fadil369 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes57 downloads6mo agoHugging Face27gimmy256 /adaption-clinical-triage-preferences This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-clinical_triage_preferences Multi-turn conversational preference dataset designed for fine-grained safety and tone calibration in emergency first aid and symptom triage. Each sample pairs a user prompt with chosen and rejected AI responses, contrasting concise, grounded clinical guidance against subtly misleading or overly verbose advice. It supports reward modeling and preference… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-clinical-triage-preferences.textn<1K0 likes57 downloads23d agoHugging Face28PatSnap /clinical-trial🏥 Dataset Summary Structured clinical trial records sourced from ClinicalTrials.gov. Each record extracts and normalizes key parameters including study purpose, intervention arms, trial conditions, and sponsor/collaborator relationships. Drug entities are enriched with multilingual name mappings (Chinese, English, Japanese). 🚀 Key Features Structured Intervention Arms: arm_intervention captures each trial arm's design group type, linked interventions, and single/multi-drug configuration.… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/clinical-trial.texttext-classification1K<n<10K1 likes55 downloads5mo agoHugging Face29ai-galileo /clinical-notes-to-fhir SGRS-FHIR: A Preference Learning Dataset for Clinical FHIR Extraction The first preference learning dataset for clinical FHIR extraction with structured error paths. Key Insight Structured extraction failures are training signal, not noise. Unlike traditional datasets that discard generation failures, SGRS-FHIR intentionally captures both valid and invalid extractions with detailed error annotations. This enables: Direct Preference Optimization (DPO): Train models to… See the full description on the dataset page: https://huggingface.co/datasets/ai-galileo/clinical-notes-to-fhir.texttext-generationn<1K4 likes53 downloads7mo agoHugging Face30shuyuej /CMMLU-Clinical-Knowledge-Benchmark 💻 Dataset Usage Run the following command to load the testing set (237 examples): from datasets import load_dataset dataset = load_dataset("shuyuej/CMMLU-Clinical-Knowledge-Benchmark", split="train") print(dataset) textn<1K1 likes51 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.