CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes20k downloads1y agoHugging Face02cocool /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.image-to-text1M<n<10M0 likes6.5k downloads23d agoHugging Face03lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.5k downloads25d agoHugging Face04shibing624 /medical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。text-generationn<1K442 likes2.4k downloads2y agoHugging Face05OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face06zabir1996 /alive-medical-imaging ALIVE Medical Imaging QA Dataset Lecture-derived question-answer corpus, retrieval index, and source materials for the ALIVE (Avatar-Lecture Interactive Video Engine) system. The dataset was built from 23 recorded lectures of an undergraduate medical imaging course and is the corpus used to fine-tune the ALIVE language model and to evaluate its retrieval and answer-generation behavior. Layout huggingface/ ├── data/ question-answer pairs (Alpaca-style… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/alive-medical-imaging.textquestion-answering1K<n<10K5 likes1.5k downloads4mo agoHugging Face07FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes948 downloads2y agoHugging Face08bofenghuang /medical-qa-fr-v0.1 Medical QA (FR) v0.1 A French medical instruction-tuning dataset (~508K examples) compiled from three public medical QA / dialogue sources: ruslanmv/ai-medical-chatbot — 256,010 examples (config ai_medical_chatbot, default) lavita/medical-qa-datasets (all-processed config) — 230,041 examples (config medical_qa_datasets) FreedomIntelligence/Medical-R1-Distill-Data — 21,641 examples (config medical_r1_distill_data) Each source question was machine-translated into French, then a… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/medical-qa-fr-v0.1.textquestion-answering100K<n<1M0 likes839 downloads3mo agoHugging Face09zabir1996 /mimic-medical-imaging-qa MIMIC Medical Imaging QA Dataset 5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction. License The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.imagequestion-answering1K<n<10K3 likes800 downloads5mo agoHugging Face10FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes798 downloads2y agoHugging Face11OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes731 downloads10mo agoHugging Face12lamhieu /medical_advice_dialogue_en Description The dataset is from medalpaca/medical_meadow_health_advice, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_advice_dialogue_en.texttext-generation1K<n<10K1 likes636 downloads2y agoHugging Face13SylvanL /Traditional-Chinese-Medicine-Dataset-Pretrain 启古纳今,厚德精术 数据介绍 非网络来源的高质量中医数据集-预训练 High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining 该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。 包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。 数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。 注意:该数据集仅适用于预训练或继续预训练用途,针对SFT/IFT的QA数据集详见:SylvanL/Traditional-Chinese-Medicine-Dataset-SFT… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain.texttext-generation100K<n<1M33 likes555 downloads2mo agoHugging Face14Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes377 downloads2y agoHugging Face15lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes340 downloads1y agoHugging Face16stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes340 downloads2mo agoHugging Face17OpenMed /Medical-Reasoning-SFT-Nemotron-Nano-30B Medical-Reasoning-SFT-Nemotron-Nano-30B A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Total Samples 444,544 Samples with Reasoning 444,544 (100%) Estimated Tokens ~1.01 Billion Content Tokens ~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.texttext-generation100K<n<1M46 likes323 downloads8mo agoHugging Face18joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face19OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes276 downloads9mo agoHugging Face20OpenMed /Medical-Reasoning-SFT-Trinity-Mini Medical-Reasoning-SFT-Trinity-Mini A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model arcee-ai/Trinity-Mini Total Samples ~810,374 Estimated Tokens ~1.52 Billion Content Tokens ~542 Million Reasoning Tokens ~977 Million Language English Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.texttext-generation100K<n<1M79 likes263 downloads8mo agoHugging Face21snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes261 downloads1mo agoHugging Face22lamhieu /medical_medqa_dialogue_en Description The dataset is from medalpaca/medical_meadow_mediqa, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_medqa_dialogue_en.texttext-generation10K<n<100K2 likes249 downloads2y agoHugging Face23asanchez75 /medical_textbooks_mcq Medical Textbooks MCQs Dataset This dataset is derived from the Textbooks corpus chunks found in the MedRAG/textbooks dataset. It augments the original text snippets with synthetically generated Multiple Choice Questions (MCQs) in JSON format, suitable for fine-tuning or evaluating language models on medical MCQ generation tasks. Dataset Details Dataset Description The source data consists of text snippets from the Textbooks corpus, a collection of 18 widely… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/medical_textbooks_mcq.textmultiple-choice1K<n<10K0 likes244 downloads1y agoHugging Face24OpenMed /Medical-Reasoning-SFT-Baichuan-M3-235B Medical-Reasoning-SFT-Baichuan-M3-235B A large-scale medical reasoning dataset generated using baichuan-inc/Baichuan-M3-235B, containing over 124,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Baichuan-M3-235B is ranked #1 on HealthBench Total leaderboard and achieves state-of-the-art performance on medical reasoning benchmarks. Dataset Overview Metric Value Model baichuan-inc/Baichuan-M3-235B Total Samples 124… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Baichuan-M3-235B.texttext-generation100K<n<1M7 likes223 downloads8mo agoHugging Face25OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes222 downloads8mo agoHugging Face26OpenMed /Medical-Reasoning-SFT-Qwen3-Next-80B Medical-Reasoning-SFT-Qwen3-Next-80B A large-scale medical reasoning dataset generated using Qwen/Qwen3-Next-80B-A3B-Thinking, containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model Qwen/Qwen3-Next-80B-A3B-Thinking Total Samples 604,249 Samples with Reasoning 604,249 (100%) Estimated Tokens ~1.42 Billion Content Tokens ~505 Million Reasoning Tokens ~917 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B.texttext-generation100K<n<1M15 likes202 downloads8mo agoHugging Face27lamhieu /medical_wikidoc_dialogue_en Description The dataset is from medalpaca/medical_meadow_wikidoc, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_wikidoc_dialogue_en.texttext-generation10K<n<100K3 likes200 downloads2y agoHugging Face28Lots-of-LoRAs /task620_ohsumed_medical_subject_headings_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task620_ohsumed_medical_subject_headings_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task620_ohsumed_medical_subject_headings_answer_generation.texttext-generation1K<n<10K0 likes191 downloads2y agoHugging Face29tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads9d agoHugging Face30NationalLibraryOfScotland /medical-history-of-british-india A Medical History of British India Dataset Dataset Description This dataset contains digitiaed official publications documenting medical research and public health in British India from 1850-1950. The collection represents a crucial period in medical history, capturing the transition from humoral to biochemical medical traditions and documenting major breakthroughs in bacteriology, parasitology, and vaccine development. These documents provide invaluable insights into… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/medical-history-of-british-india.imagetext-generation100K<n<1M2 likes173 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.