CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face02OpenMed /MedDialog MedDialog A large-scale medical dialogue dataset containing ~252k patient-doctor conversation pairs for training and evaluating clinical dialogue systems. Dataset Description Property Value Source ruslanmv/ai-medical-chatbot License Apache-2.0 Language English Total examples 251,731 Train split 226,557 Validation split 25,174 Domain Clinical / General Medicine Overview MedDialog is designed for training language models to generate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/MedDialog.texttext-generation100K<n<1M21 likes591 downloads7mo agoHugging Face03OpenMed /Medical-Reasoning-SFT-Nemotron-Nano-30B Medical-Reasoning-SFT-Nemotron-Nano-30B A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Total Samples 444,544 Samples with Reasoning 444,544 (100%) Estimated Tokens ~1.01 Billion Content Tokens ~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.texttext-generation100K<n<1M46 likes318 downloads8mo agoHugging Face04OpenMed /Medical-Reasoning-SFT-Trinity-Mini Medical-Reasoning-SFT-Trinity-Mini A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model arcee-ai/Trinity-Mini Total Samples ~810,374 Estimated Tokens ~1.52 Billion Content Tokens ~542 Million Reasoning Tokens ~977 Million Language English Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.texttext-generation100K<n<1M79 likes250 downloads8mo agoHugging Face05openmed-community /MedReason-Stenographic MedReason-Stenographic: Medical QA with Compressed Reasoning Traces This dataset contains 31,535 medical question-answer pairs with stenographic reasoning traces, generated using MiniMax M2.1 from the original UCSC-VLAA/MedReason dataset. Dataset Description The dataset transforms medical QA reasoning into a stenographic format using a symbolic protocol designed for high-density, machine-parseable reasoning traces. This format eliminates natural language filler while… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/MedReason-Stenographic.textquestion-answering10K<n<100K56 likes236 downloads9mo agoHugging Face06OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes213 downloads8mo agoHugging Face07OpenMed /Medical-Reasoning-SFT-Baichuan-M3-235B Medical-Reasoning-SFT-Baichuan-M3-235B A large-scale medical reasoning dataset generated using baichuan-inc/Baichuan-M3-235B, containing over 124,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Baichuan-M3-235B is ranked #1 on HealthBench Total leaderboard and achieves state-of-the-art performance on medical reasoning benchmarks. Dataset Overview Metric Value Model baichuan-inc/Baichuan-M3-235B Total Samples 124… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Baichuan-M3-235B.texttext-generation100K<n<1M7 likes211 downloads8mo agoHugging Face08OpenMed /Medical-Reasoning-SFT-Qwen3-Next-80B Medical-Reasoning-SFT-Qwen3-Next-80B A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model Qwen/Qwen3-Next-80B-A3B-Thinking Total Samples 604,249 Samples with Reasoning 604,249 (100%) Estimated Tokens ~1.42 Billion Content Tokens ~505 Million Reasoning Tokens ~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.texttext-generation100K<n<1M15 likes199 downloads8mo agoHugging Face09OpenMed /Medical-Reasoning-SFT-MiniMax-M2.1 Medical-Reasoning-SFT-MiniMax-M2.1 A large-scale medical reasoning dataset generated using MiniMaxAI/MiniMax-M2.1, containing over 204,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model MiniMaxAI/MiniMax-M2.1 Total Samples 204,773 Samples with Reasoning 204,773 (100%) Estimated Tokens ~621 Million Content Tokens ~344 Million Reasoning Tokens ~277 Million Language English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-MiniMax-M2.1.texttext-generation100K<n<1M9 likes107 downloads8mo agoHugging Face10OpenMed /Medical-Reasoning-SFT-GLM_4.5_Air Medical-Reasoning-SFT-GLM_4.5_Air A large-scale medical reasoning dataset generated using zai-org/GLM-4.5-Air, containing over 225,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model zai-org/GLM-4.5-Air Total Samples 225,179 Samples with Reasoning 224,942 (99.9%) Estimated Tokens ~441 Million Content Tokens ~315 Million Reasoning Tokens ~126 Million Language English… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GLM_4.5_Air.texttext-generation100K<n<1M15 likes102 downloads8mo agoHugging Face11OrionLLM /OpenMedicalInstruct OpenMedicalInstruct texttext-generation10K<n<100K4 likes59 downloads6mo agoHugging Face12openmed-community /synthetic-neurology-conversations Synthetic Neurology Conversations Summary.This dataset augments questions from KryptoniteCrown/synthetic-neurology-QA-dataset with a compact two-step follow-up conversation generated by moonshotai/Kimi-K2-Instruct: model answers the original question, model asks a follow-up question (to deepen/clarify), model answers its follow-up. Shared by the OpenMed Community to help improve medical models globally. Not medical advice. Research/education only; not for clinical… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/synthetic-neurology-conversations.texttext-generation1K<n<10K10 likes48 downloads1y agoHugging Face13HPAI-BSC /OpenMedQA OpenMedQA OpenMedQA is an open-ended medical question-answering benchmark designed to evaluate the capabilities of LLMs in generating free-text medical responses. It extends the MedQA dataset by rephrasing multiple-choice questions into an open-ended format while preserving their original medical intent. The dataset enables direct comparisons between multiple-choice (MCQA) and open-ended (OE)… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/OpenMedQA.textquestion-answering1K<n<10K0 likes46 downloads1y agoHugging Face14openmed-community /mmlu-5-options-rl-ready MMLU – 5-Options RL-Ready A standardized, RL-friendly remix of MMLU with explicit negatives and a unified five-option presentation string for each question. Ideal for DPO and other RL setups while remaining drop-in for classic multiple-choice evaluation. What’s inside Splits & size: ~97.8k train + 2k test ≈ 99.8k total. Schema (core fields): question: str choices: list[str] (canonical options, typically 4 as in original MMLU) answer: int (0-based index) task: str… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/mmlu-5-options-rl-ready.textmultiple-choice10K<n<100K1 likes36 downloads11mo agoHugging Face15mkurman /synthlabs-openmed-questions-qwen3-235b-a22b-2507 Med Synth Questions (Qwen3-235B questions + DeepSeek V4 Flash and Minimax M2.7 answers) Synthetic reasoning traces for medical questions from openmed-community/med-synth-questions-qwen3-235b-a22b-2507. Each record contains a medical question with SYNTH-style reasoning and a generated answer. Dataset Summary 55,915 records (255 dupes + 3,523 incomplete/truncated removed from 59,693 source) 55,915 reasoning turns (99.9% format compliance) Average 1,881 chars per… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-openmed-questions-qwen3-235b-a22b-2507.tabulartext-generation10K<n<100K0 likes35 downloads2mo agoHugging Face16exafluence /Open-MedQA-Nexus Open Nexus MedQA This dataset combines various publicly available medical datasets like ChatDoctor, icliniq, etc., into a unified format for training and evaluating medical question-answering models. Dataset Details Open Nexus MedQA is a comprehensive dataset designed to facilitate the development of advanced medical question answering systems. It integrates diverse medical data sources, meticulously processed to provide a uniform format. The format includes:… See the full description on the dataset page: https://huggingface.co/datasets/exafluence/Open-MedQA-Nexus.textquestion-answering100K<n<1M0 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.