datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.MedDialog
MedDialog
A large-scale medical dialogue dataset containing ~252k patient-doctor conversation pairs for training and evaluating clinical dialogue systems.
Dataset Description
Property
Value
Source
ruslanmv/ai-medical-chatbot
License
Apache-2.0
Language
English
Total examples
251,731
Train split
226,557
Validation split
25,174
Domain
Clinical / General Medicine
Overview
MedDialog is designed for training language models to generate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/MedDialog.Medical-Reasoning-SFT-Nemotron-Nano-30B
Medical-Reasoning-SFT-Nemotron-Nano-30B
A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Total Samples
444,544
Samples with Reasoning
444,544 (100%)
Estimated Tokens
~1.01 Billion
Content Tokens
~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.Medical-Reasoning-SFT-Trinity-Mini
Medical-Reasoning-SFT-Trinity-Mini
A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
arcee-ai/Trinity-Mini
Total Samples
~810,374
Estimated Tokens
~1.52 Billion
Content Tokens
~542 Million
Reasoning Tokens
~977 Million
Language
English
Schema
Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.MedReason-Stenographic
MedReason-Stenographic: Medical QA with Compressed Reasoning Traces
This dataset contains 31,535 medical question-answer pairs with stenographic reasoning traces, generated using MiniMax M2.1 from the original UCSC-VLAA/MedReason dataset.
Dataset Description
The dataset transforms medical QA reasoning into a stenographic format using a symbolic protocol designed for high-density, machine-parseable reasoning traces. This format eliminates natural language filler while… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/MedReason-Stenographic.Medical-Reasoning-SFT-GPT-OSS-120B-V2
Medical-Reasoning-SFT-GPT-OSS-120B-V2
A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed.
Dataset Overview
Metric
Value
Model
openai/gpt-oss-120b
Total Samples
506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.Medical-Reasoning-SFT-Baichuan-M3-235B
Medical-Reasoning-SFT-Baichuan-M3-235B
A large-scale medical reasoning dataset generated using baichuan-inc/Baichuan-M3-235B, containing over 124,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Baichuan-M3-235B is ranked #1 on HealthBench Total leaderboard and achieves state-of-the-art performance on medical reasoning benchmarks.
Dataset Overview
Metric
Value
Model
baichuan-inc/Baichuan-M3-235B
Total Samples
124… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Baichuan-M3-235B.Medical-Reasoning-SFT-Qwen3-Next-80B
Medical-Reasoning-SFT-Qwen3-Next-80B
A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
Qwen/Qwen3-Next-80B-A3B-Thinking
Total Samples
604,249
Samples with Reasoning
604,249 (100%)
Estimated Tokens
~1.42 Billion
Content Tokens
~505 Million
Reasoning Tokens
~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.Medical-Reasoning-SFT-MiniMax-M2.1
Medical-Reasoning-SFT-MiniMax-M2.1
A large-scale medical reasoning dataset generated using MiniMaxAI/MiniMax-M2.1, containing over 204,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
MiniMaxAI/MiniMax-M2.1
Total Samples
204,773
Samples with Reasoning
204,773 (100%)
Estimated Tokens
~621 Million
Content Tokens
~344 Million
Reasoning Tokens
~277 Million
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-MiniMax-M2.1.Medical-Reasoning-SFT-GLM_4.5_Air
Medical-Reasoning-SFT-GLM_4.5_Air
A large-scale medical reasoning dataset generated using zai-org/GLM-4.5-Air, containing over 225,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
zai-org/GLM-4.5-Air
Total Samples
225,179
Samples with Reasoning
224,942 (99.9%)
Estimated Tokens
~441 Million
Content Tokens
~315 Million
Reasoning Tokens
~126 Million
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GLM_4.5_Air.OpenMedicalInstruct
OpenMedicalInstruct
synthetic-neurology-conversations
Synthetic Neurology Conversations
Summary.This dataset augments questions from KryptoniteCrown/synthetic-neurology-QA-dataset with a compact two-step follow-up conversation generated by moonshotai/Kimi-K2-Instruct:
model answers the original question,
model asks a follow-up question (to deepen/clarify),
model answers its follow-up.
Shared by the OpenMed Community to help improve medical models globally.
Not medical advice. Research/education only; not for clinical… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/synthetic-neurology-conversations.OpenMedQA
OpenMedQA
OpenMedQA is an open-ended medical question-answering benchmark designed to evaluate the capabilities of LLMs in generating free-text medical responses. It extends the MedQA dataset by rephrasing multiple-choice questions into an open-ended format while preserving their original medical intent. The dataset enables direct comparisons between multiple-choice (MCQA) and open-ended (OE)… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/OpenMedQA.mmlu-5-options-rl-ready
MMLU – 5-Options RL-Ready
A standardized, RL-friendly remix of MMLU with explicit negatives and a unified five-option presentation string for each question. Ideal for DPO and other RL setups while remaining drop-in for classic multiple-choice evaluation.
What’s inside
Splits & size: ~97.8k train + 2k test ≈ 99.8k total.
Schema (core fields):
question: str
choices: list[str] (canonical options, typically 4 as in original MMLU)
answer: int (0-based index)
task: str… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/mmlu-5-options-rl-ready.synthlabs-openmed-questions-qwen3-235b-a22b-2507
Med Synth Questions (Qwen3-235B questions + DeepSeek V4 Flash and Minimax M2.7 answers)
Synthetic reasoning traces for medical questions from openmed-community/med-synth-questions-qwen3-235b-a22b-2507. Each record contains a medical question with SYNTH-style reasoning and a generated answer.
Dataset Summary
55,915 records (255 dupes + 3,523 incomplete/truncated removed from 59,693 source)
55,915 reasoning turns (99.9% format compliance)
Average 1,881 chars per… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-openmed-questions-qwen3-235b-a22b-2507.Open-MedQA-Nexus
Open Nexus MedQA
This dataset combines various publicly available medical datasets like ChatDoctor, icliniq, etc., into a unified format for training and evaluating medical question-answering models.
Dataset Details
Open Nexus MedQA is a comprehensive dataset designed to facilitate the development of advanced medical question answering systems. It integrates diverse medical data sources, meticulously processed to provide a uniform format. The format includes:… See the full description on the dataset page: https://huggingface.co/datasets/exafluence/Open-MedQA-Nexus.
