datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
healthbench-professionalContains the data for the HealthBench Professional eval.
Each example contains:
conversation: list of user / assistant messages, ending in a user message
rubric_items: list of rubric items, each containing criterion_text and points
use_case: one of consult, writing, or research
type: one of good_faith or red_teaming
difficulty: physician-assigned difficulty rating (difficult for Likert 1-2, typical for Likert 3-7)
specialty: medical specialty or sub-specialty
physician_response: response… See the full description on the dataset page: https://huggingface.co/datasets/openai/healthbench-professional.medical_meadow_health_advice
Health Advice
Dataset Summary
This is the dataset use in the paper: Detecting Causal Language Use in Science Findings.
It was cleaned and formated to fit into the alpaca template.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.healthspend-dataChatDoctor_HealthCareMagicThe ChatDoctor-HealthCareMagic-100k dataset comprises 112,000 real-world medical question-and-answer pairs, providing a substantial and diverse collection of authentic medical dialogues. There is a slight risk to this dataset since there are grammatical inconsistencies in many of the questions and answers, but this can potentially help separate strong healthcare retrieval models from weak ones.
Usage
import datasets
# Download the dataset
queries =… See the full description on the dataset page: https://huggingface.co/datasets/embedding-benchmark/ChatDoctor_HealthCareMagic.HealthCareMagic-100k-enmedical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.HealthChat-11K
HealthChat-11K
This repository contains HealthChat-11K, a curated dataset of approximately 11,000 real-world conversations, composed of 25,000 user messages, where users seek healthcare information from Large Language Models (LLMs). The goal of this work is to provide a high-quality resource for systematically studying and improving health conversations involving humans and AI (e.g., LLMs).
The dataset was presented in the paper: "What's Up, Doc?": Analyzing How Users Seek Health… See the full description on the dataset page: https://huggingface.co/datasets/yahskapar/HealthChat-11K.healthbench-psych
HealthBench-Psych
An expert-adjudicated mental-health subset of HealthBench
(OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations), together with
released results for 23 language models under a three-judge panel.
Maintained by MindBench.ai · Division of Digital Psychiatry, Beth Israel
Deaconess Medical Center. Code and evaluation harness:
github.com/mindbench-ai/healthbench-psych.
Subset
n
Definition
healthbench-psych-v2
611… See the full description on the dataset page: https://huggingface.co/datasets/mindbench-ai/healthbench-psych.Nepali-HealthChatevipedia-reviews
Evipedia Evidence Reviews
The full public catalogue of evipedia.ai — a
continuously-updated encyclopedia of evidence reviews on health & longevity
interventions — as one record per review. Each record carries the review's
metadata plus its complete Markdown body.
Homepage / source: https://evipedia.ai
Live file: https://evipedia.ai/evipedia-corpus.jsonl (this dataset mirrors it)
Publisher: Forever Healthy
License: CC BY 4.0
What's inside
One JSON object per… See the full description on the dataset page: https://huggingface.co/datasets/forever-healthy/evipedia-reviews.HealthCareMagic-100k-Chat-Format-enHealMed
HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems.
The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.healthbench-regularIndustryInstruction_Health-Medicine
IndustryInstruction: Health & Medicine
This repository contains the IndustryInstruction: Health & Medicine domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Health-Medicine.lavita-ChatDoctor-HealthCareMagic-100kMental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.Health-Bench-Eval-OSS-2025-07
Dataset Card for HealthBench
Dataset Summary
HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.pii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Health & Medical Information (PHI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.aurel_tensorsNepali-Health-QAMM-Health
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation
GitHub Repository: https://github.com/grantzyr/MM-Health-Dataset
Dataset Description
MM-Health is a large-scale multimodal dataset designed for detecting both human and AI-generated health misinformation. The dataset consists of 34,746 news articles encompassing both textual and visual information, making it the most comprehensive multimodal health misinformation dataset… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/MM-Health.healthbench
THE CODE IS CURRENTLY BROKEN BUT THE DATASET IS GOOD!!
HealthBench Implementation for using Opensource Judges
Easy-to-use implementation of OpenAI's HealthBench evaluation benchmark with support for any OpenAI API-compatible model as both the system under test and the judge.
Developed by: Nisten Tahiraj / OnDeviceMednotes
License: MIT
Paper: HealthBench: Evaluating Large Language Models Towards Improved Human Health
Overview
This repository contains tools… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/healthbench.long-doc_healthcare_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: test
pubmed_40K-50K_5-merged: test
pubmed_30K-40K_10-merged: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: dev
pubmed_40K-50K_5-merged: test… See the full description on the dataset page: https://huggingface.co/datasets/AIR-Bench/long-doc_healthcare_en.healthbench_medication_filteredmental-healthMental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
turkish-medical-deid-eval
Turkish Medical De-Identification Evaluation Corpus (Synthetic)
A labelled benchmark for evaluating the removal of personally identifiable information (PII)
from Turkish medical speech-to-text (STT) transcripts.
This dataset contains no real data. Every consultation, name, phone number, address,
identifier and financial detail is programmatically generated and fictitious. The corpus exists
specifically so that de-identification systems can be evaluated without any real patient… See the full description on the dataset page: https://huggingface.co/datasets/Keler-Health/turkish-medical-deid-eval.Mental-health-CBT-dialogues
Mental Health CBT Dialogues
Overview
This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models.
The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress.
The dataset accompanies the paper:
Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.HealthSearchQA-ko
HealthSearchQA
Original Data: Med-PaLM.
Question is translated into Korean by "solar-1-mini-translate-enko".
Answer is generated by gpt-4o-2024-05-13. Note that I asked the question in English, not the translated version in step 1.
Exact Prompt is as follows:
def make_answer(source_question):
response = client.chat.completions.create(
model="gpt-4o-2024-05-13",
messages=[
{
"role": "system",
"content": "Adopt the role of MEDICAL EXPERT qualified… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/HealthSearchQA-ko.healthver_resplit
