datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.medical_meadow_medqa
Dataset Card for MedQA
Dataset Summary
This is the data and baseline source code for the paper: Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams."
From https://github.com/jind11/MedQA:
The data that contains both the QAs and textbooks can be downloaded from this google drive folder. A bit of details of data are explained as below:
For QAs, we have three sources: US, Mainland of China, and… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medqa.medical_meadow_medical_flashcards
Dataset Card for Medical Flashcards
Dataset Summary
Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master
in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge,
and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the
entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.medical_meadow_wikidoc
Dataset Card for WikiDoc
For the dataset containing patient information from wikidoc refer to this dataset
Dataset Summary
This dataset containes medical question-answer pairs extracted from WikiDoc,
a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge.
The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook"
contains chapters for various medical specialties, which we… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc.medical_meadow_wikidoc_patient_information
Dataset Card for WikiDoc
For the dataset containing rephrased content from the living textbook refer to this dataset
Dataset Summary
This dataset containes medical question-answer pairs extracted from WikiDoc,
a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge.
The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook"
contains chapters for various medical specialties… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc_patient_information.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.medical_meadow_health_advice
Health Advice
Dataset Summary
This is the dataset use in the paper: Detecting Causal Language Use in Science Findings.
It was cleaned and formated to fit into the alpaca template.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.medical_meadow_mediqa
MediQA
Dataset Description
MEDIQA is a dataset of manually generated, question-driven summaries of multi and single document answers to consumer health questions.
Homepage: https://osf.io/fyg46/?view_only=
Citation Information
@article{savery2020question,
title={Question-driven summarization of answers to consumer health questions},
author={Savery, Max and Abacha, Asma Ben and Gayen, Soumya and Demner-Fushman, Dina},
journal={Scientific Data}… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_mediqa.alive-medical-imaging
ALIVE Medical Imaging QA Dataset
Lecture-derived question-answer corpus, retrieval index, and source
materials for the ALIVE (Avatar-Lecture Interactive Video Engine)
system. The dataset was built from 23 recorded lectures of an
undergraduate medical imaging course and is the corpus used to
fine-tune the ALIVE language model and to evaluate its retrieval and
answer-generation behavior.
Layout
huggingface/
├── data/ question-answer pairs (Alpaca-style… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/alive-medical-imaging.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.know_medical_dialogue_v2
Description:
The knowrohit07/know_medical_dialogues_v2 dataset is a collection of conversational exchanges between patients and doctors on various medical topics. It aims to capture the intricacies, uncertainties, and questions posed by individuals regarding their health and the medical guidance provided in response.
🎯 Intended Use:
This dataset is crafted for training Large Language Models (LLMs) with a focus on understanding and generating medically-informed dialogue.… See the full description on the dataset page: https://huggingface.co/datasets/knowrohit07/know_medical_dialogue_v2.medical-specialities
Medical Question Classification Dataset
Dataset Summary
This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly.
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medical-specialities.medical_meadow_cord19
CORD 19
Dataset Summary
In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.medical-instruction-120k
What is the Dataset About?🤷🏼♂️
The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions.
How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed-Altaf/medical-instruction-120k.Medical-Sciences
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.medical_meadow_pubmed_causal
Dataset Card for Pubmed Causal
Dataset Summary
This is the dataset used in the paper: Detecting Causal Language Use in Science Findings.
Citation Information
@inproceedings{yu-etal-2019-detecting,
title = "Detecting Causal Language Use in Science Findings",
author = "Yu, Bei and
Li, Yingya and
Wang, Jun",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_pubmed_causal.Medical-R1-Distill-Data
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1.
The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.medical-prescription-datasetmedical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/medicine-tasks.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.medical-prescription-datasetmedical_qa
MedicalQARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
The dataset consists 2048 medical question and answer pairs.
Task category
t2t
Domains
Medical, Written
Reference
https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-019-3119-4
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MedicalQARetrieval"])
evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/medical_qa.medical_meadow_mmmlumedical_cot
Medical Question-Answering Dataset
A comprehensive collection of medical questions and detailed answers, designed for training and evaluating medical question-answering systems.
Dataset Description
Overview
This dataset contains medical questions with multiple-choice answers and detailed explanations. Each question presents a clinical scenario and requires medical knowledge to determine the correct diagnosis, treatment, or underlying mechanism.
Data… See the full description on the dataset page: https://huggingface.co/datasets/blue-blues/medical_cot.medical-instruction-100k
What is the Dataset About?🤷🏼♂️
The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions.
How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed-Altaf/medical-instruction-100k.medical_fine_tuning_12M
