CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes20k downloads1y agoHugging Face02cocool /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.image-to-text1M<n<10M0 likes6.5k downloads23d agoHugging Face03bluusun /mighty-media-corpus Mighty Media Corpus Independent technical analyses and research breakdowns published across the Mighty Media network (mightytravels.com, judgmentcallpodcast.com, and vertical sites). One JSON file per article: {url, title, site, date, author, text, citations, license, note}. License: CC-BY-4.0 Canonical versions live at the url of each row. Articles are independent analyses, not peer-reviewed publications. Files are organized by month: data/YYYY-MM/*.json texttext-generationn<1K1 likes6.5k downloads17d agoHugging Face04lavita /AlpaCare-MedInstruct-52k Dataset Card for "AlpaCare-MedInstruct-52k" AlpaCare GitHub repo: https://github.com/XZhang97666/AlpaCare Citation: If you use this dataset, please cite the original paper: @misc{zhang2023alpacareinstructiontuned, title={AlpaCare: Instruction-tuned Large Language Models for Medical Application}, author={Xinlu Zhang and Chenxin Tian and Xianjun Yang and Lichang Chen and Zekun Li and Linda Ruth Petzold}, year={2023}, eprint={2310.14558}… See the full description on the dataset page: https://huggingface.co/datasets/lavita/AlpaCare-MedInstruct-52k.texttext-generation10K<n<100K24 likes4.6k downloads2y agoHugging Face05lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.5k downloads25d agoHugging Face06shibing624 /medical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。text-generationn<1K442 likes2.4k downloads2y agoHugging Face07OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face08zabir1996 /alive-medical-imaging ALIVE Medical Imaging QA Dataset Lecture-derived question-answer corpus, retrieval index, and source materials for the ALIVE (Avatar-Lecture Interactive Video Engine) system. The dataset was built from 23 recorded lectures of an undergraduate medical imaging course and is the corpus used to fine-tune the ALIVE language model and to evaluate its retrieval and answer-generation behavior. Layout huggingface/ ├── data/ question-answer pairs (Alpaca-style… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/alive-medical-imaging.textquestion-answering1K<n<10K5 likes1.5k downloads4mo agoHugging Face09FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes948 downloads2y agoHugging Face10bofenghuang /medical-qa-fr-v0.1 Medical QA (FR) v0.1 A French medical instruction-tuning dataset (~508K examples) compiled from three public medical QA / dialogue sources: ruslanmv/ai-medical-chatbot — 256,010 examples (config ai_medical_chatbot, default) lavita/medical-qa-datasets (all-processed config) — 230,041 examples (config medical_qa_datasets) FreedomIntelligence/Medical-R1-Distill-Data — 21,641 examples (config medical_r1_distill_data) Each source question was machine-translated into French, then a… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/medical-qa-fr-v0.1.textquestion-answering100K<n<1M0 likes839 downloads3mo agoHugging Face11zabir1996 /mimic-medical-imaging-qa MIMIC Medical Imaging QA Dataset 5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction. License The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.imagequestion-answering1K<n<10K3 likes800 downloads5mo agoHugging Face12FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes798 downloads2y agoHugging Face13OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes731 downloads10mo agoHugging Face14lamhieu /medical_advice_dialogue_en Description The dataset is from medalpaca/medical_meadow_health_advice, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_advice_dialogue_en.texttext-generation1K<n<10K1 likes636 downloads2y agoHugging Face15SylvanL /Traditional-Chinese-Medicine-Dataset-Pretrain 启古纳今,厚德精术 数据介绍 非网络来源的高质量中医数据集-预训练 High-Quality Traditional Chinese Medicine Dataset from Non-Internet Sources - Pretraining 该数据集经过大量人力和资源的投入精心构建,以共建LLM高质量中文社区为己任。 包含约1GB的中医各个领域临床案例、名家典籍、医学百科,名词解释等优质内容,涵盖全面,配比均衡。 数据集主要由非网络来源的内部数据构成,并99%为简体中文内容,内容质量优异,信息密度可观。 注意:该数据集仅适用于预训练或继续预训练用途,针对SFT/IFT的QA数据集详见:SylvanL/Traditional-Chinese-Medicine-Dataset-SFT… See the full description on the dataset page: https://huggingface.co/datasets/SylvanL/Traditional-Chinese-Medicine-Dataset-Pretrain.texttext-generation100K<n<1M33 likes555 downloads2mo agoHugging Face16Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes377 downloads2y agoHugging Face17lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes340 downloads1y agoHugging Face18stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes340 downloads2mo agoHugging Face19casey-martin /MedInstruct MedInstruct This is the repo for MedInstruct, which is a dataset of synthetically generated medical instructions. The repo contains: The 52K medical instruction-response dataset MedInstruct-52k used for fine-tuning AlpaCare, and corresponding clinican-crafted seed task to generate instruction. A 217 clinical craft free-form instruction evaluation test set,MedInstruct-test. The code for: medical task generation; fine-tuning LLaMA series models; instrcution-tuned model response… See the full description on the dataset page: https://huggingface.co/datasets/casey-martin/MedInstruct.text-generation7 likes336 downloads3y agoHugging Face20EPFLiGHT /fully-open-meditron Fully Open Meditron Corpus 👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here. [Hugging Face] [Preprint] [GitHub] [Dataset] License: Apache 2.0 | Authors: LiGHT [!Note] A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.textquestion-answering100K<n<1M8 likes327 downloads3mo agoHugging Face21OpenMed /Medical-Reasoning-SFT-Nemotron-Nano-30B Medical-Reasoning-SFT-Nemotron-Nano-30B A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Total Samples 444,544 Samples with Reasoning 444,544 (100%) Estimated Tokens ~1.01 Billion Content Tokens ~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.texttext-generation100K<n<1M46 likes323 downloads8mo agoHugging Face22KiteFishAI /arxiv-tex-corpus-mediumarxiv-tex-corpus-medium (15GB) Medium-scale LaTeX corpus from arXiv (math, CS, physics, statistics) 📄 Paper: https://arxiv.org/abs/2602.17288 📚 Overview arxiv-tex-corpus-medium (15GB) is a medium-sized version of the arXiv LaTeX corpus, containing structured LaTeX source content extracted from selected arXiv categories. This dataset is restricted to the following categories: math cs hep-th hep-ph quant-ph stat.ML stat.TH This version (~15GB) is intended for: Research… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-medium.texttext-generation100K<n<1M3 likes322 downloads7mo agoHugging Face23JWei05 /DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.texttext-generation1K<n<10K0 likes304 downloads1mo agoHugging Face24grammarly /medit Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning Paper: mEdIT: Multilingual Text Editing via Instruction Tuning Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar Project Repo: https://github.com/vipulraheja/medit Dataset Summary This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.texttext-generation100K<n<1M14 likes291 downloads2y agoHugging Face25xz97 /MedInstruct Dataset Card for MedInstruct Dataset Summary MedInstruct encompasses: MedInstruct-52k: A dataset comprising 52,000 medical instructions and responses. Instructions are crafted by OpenAI's GPT-4 engine, and the responses are formulated by the GPT-3.5-turbo engine. MedInstruct-test: A set of 217 clinical craft free-form instruction evaluation tests. med_seed: The clinician-crafted seed set as a denomination to prompt GPT-4 for task generation. MedInstruct-52k can be used… See the full description on the dataset page: https://huggingface.co/datasets/xz97/MedInstruct.texttext-generationn<1K20 likes290 downloads3y agoHugging Face26microsoft /mediflow MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.tabulartext-generation1M<n<10M53 likes290 downloads9mo agoHugging Face27joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face28OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes276 downloads9mo agoHugging Face29OpenMed /Medical-Reasoning-SFT-Trinity-Mini Medical-Reasoning-SFT-Trinity-Mini A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model arcee-ai/Trinity-Mini Total Samples ~810,374 Estimated Tokens ~1.52 Billion Content Tokens ~542 Million Reasoning Tokens ~977 Million Language English Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.texttext-generation100K<n<1M79 likes263 downloads8mo agoHugging Face30snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes261 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.