datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatDoctor-HealthCareMagic-100k
Dataset Card for "ChatDoctor-HealthCareMagic-100k"
More Information needed
healthbench-professionalContains the data for the HealthBench Professional eval.
Each example contains:
conversation: list of user / assistant messages, ending in a user message
rubric_items: list of rubric items, each containing criterion_text and points
use_case: one of consult, writing, or research
type: one of good_faith or red_teaming
difficulty: physician-assigned difficulty rating (difficult for Likert 1-2, typical for Likert 3-7)
specialty: medical specialty or sub-specialty
physician_response: response… See the full description on the dataset page: https://huggingface.co/datasets/openai/healthbench-professional.medical_meadow_health_advice
Health Advice
Dataset Summary
This is the dataset use in the paper: Detecting Causal Language Use in Science Findings.
It was cleaned and formated to fit into the alpaca template.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine
IndustryCorpus2: Health & Medicine
This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.Healix-2.8B-Token-Medical-Shot
Dataset Card for "Healix-2.8B-Token-Medical-Shot"
More Information needed
syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test.
practice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.healthspend-dataHC4
HC4 (Healthcare Comprehensive Commons Corpus)
HC4 is a large-scale pretraining dataset containing over 65 billion tokens from diverse healthcare-related sources.
The corpus was curated to enable systematic investigation of how data composition influences language model behavior, including potential demographic biases.
Dataset Overview
Dataset Name: HC4 (Healthcare Comprehensive Commons Corpus)
Size: 153GB (around 65 billion tokens)
Number of samples: 9.7+ million… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/HC4.VL-Health
VL-Health Dataset
Overview
The VL-Health dataset is designed for multi-stage training of unified LVLMs in the medical domain. It consists of two key phases:
Alignment – Focused on training image captioning capabilities and learning representations of input visual information.
Instruct Fine-Tuning – Designed for enhancing the model's ability to handle various vision-language tasks, including both visual comprehension and visual generation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lintw/VL-Health.ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1Healix-ShotREADME
Healix-Shot: Largest Medical Corpora by Health 360
Healix-Shot, proudly presented by Health 360, stands as an emblematic milestone in the realm of medical datasets. Hosted on the HuggingFace repository, it heralds the infusion of cutting-edge AI in the healthcare domain. With an astounding 22 billion tokens, Healix-Shot provides a comprehensive, high-quality corpus of medical text, laying the foundation for unparalleled medical NLP applications.
Importance:… See the full description on the dataset page: https://huggingface.co/datasets/health360/Healix-Shot.syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test.
ChatDoctor_HealthCareMagicThe ChatDoctor-HealthCareMagic-100k dataset comprises 112,000 real-world medical question-and-answer pairs, providing a substantial and diverse collection of authentic medical dialogues. There is a slight risk to this dataset since there are grammatical inconsistencies in many of the questions and answers, but this can potentially help separate strong healthcare retrieval models from weak ones.
Usage
import datasets
# Download the dataset
queries =… See the full description on the dataset page: https://huggingface.co/datasets/embedding-benchmark/ChatDoctor_HealthCareMagic.twi-health-asr
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages,
sourced from publicly available video content on health and wellness.
Created by Mich-Seth Owusu and… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr.HealthCareMagic-100k-enhealthsearchqa
HealthSearchQA
Dataset of consumer health questions released by Google for the Med-PaLM paper (arXiv preprint).
From the paper:
We curated our own additional dataset consisting of 3,173 commonly searched consumer questions,
referred to as HealthSearchQA. The dataset was curated using seed medical conditions and their
associated symptoms. We used the seed data to retrieve publicly-available commonly searched questions
generated by a search engine, which were displayed to all users… See the full description on the dataset page: https://huggingface.co/datasets/katielink/healthsearchqa.medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.healthqa-br
HealthQA-BR
Resumo
O HealthQA-BR é o primeiro benchmark de larga escala e abrangência para todo o Sistema Único de Saúde (SUS), projetado para medir o conhecimento clínico de Grandes Modelos de Linguagem (LLMs) frente aos desafios da saúde pública brasileira. Composto por 5.632 questões de múltipla escolha, o conjunto de dados é derivado de provas e concursos de licenciamento profissional e residência de abrangência nacional e de alto impacto no Brasil.
Diferentemente de… See the full description on the dataset page: https://huggingface.co/datasets/Larxel/healthqa-br.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
twi-health-asr-gemini-500hrs
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Health Speech Dataset Gemini (500 hours)
A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken
languages, sourced from publicly available video content on health and wellness.
Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized
Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized"
More Information needed
Health-Nutrition-and-Population-Indicators-For-African-Countries
Health Nutrition and Population Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Health-Nutrition-and-Population-Indicators-For-African-Countries.Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.HealthBench_Knowledge_Question_250KPARHAF
Dataset Card for PARHAF
Reporting Issues & Contributing
If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face.
For more substantial contributions or collaboration opportunities, feel free to contact us directly.
Dataset Description
Note for users interested in the PARTAGES use cases on pseudonymisation, coding, oncology, and infectiology:… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARHAF.twi-health-asr-gemini-500hrs-ipa
Twi Health Speech — Audio, Transcript and IPA
Twi health-domain speech with both a written transcript and an IPA phoneme sequence read off the audio by ASR. Built from ghananlpcommunity/twi-health-asr-gemini-500hrs by adding the IPA column.
from datasets import load_dataset
ds = load_dataset("ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa", split="train")
ds[0]["audio"] # decoded waveform, 16 kHz
ds[0]["transcription"] # transcript
ds[0]["ipa"]… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa.nairobi-longitudinal-healthcare-utilization
Nairobi Longitudinal Healthcare Utilization
Synthetic longitudinal population data for predicting healthcare use across time.
100% synthetic. No real patient records. Every resident, household, insurance record,
employment state, education history, encounter, diagnosis, prescription, condition, and
healthcare outcome in this release is synthetic. No real individual's data was used to produce it.
What is this?
This dataset is a balanced longitudinal… See the full description on the dataset page: https://huggingface.co/datasets/stripeddonkey-data/nairobi-longitudinal-healthcare-utilization.
