CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes20k downloads1y agoHugging Face02meoconxinhxan /Medical-Eval-HumanityLastExamtextn<1K1 likes11k downloads1y agoHugging Face03lavita /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.textquestion-answering1M<n<10M64 likes10k downloads3y agoHugging Face04medalpaca /medical_meadow_medqa Dataset Card for MedQA Dataset Summary This is the data and baseline source code for the paper: Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams." From https://github.com/jind11/MedQA: The data that contains both the QAs and textbooks can be downloaded from this google drive folder. A bit of details of data are explained as below: For QAs, we have three sources: US, Mainland of China, and… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medqa.textquestion-answering10K<n<100K117 likes8.3k downloads3y agoHugging Face05keivalya /MedQuad-MedicalQnADataset Reference: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019. textquestion-answering10K<n<100K133 likes7.4k downloads3y agoHugging Face06medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.4k downloads3y agoHugging Face07medalpaca /medical_meadow_wikidoc Dataset Card for WikiDoc For the dataset containing patient information from wikidoc refer to this dataset Dataset Summary This dataset containes medical question-answer pairs extracted from WikiDoc, a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge. The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook" contains chapters for various medical specialties, which we… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc.textquestion-answering10K<n<100K56 likes5.5k downloads3y agoHugging Face08ruslanmv /ai-medical-chatbot AI Medical Chatbot Dataset This is an experimental Dataset designed to run a Medical Chatbot It contains at least 250k dialogues between a Patient and a Doctor. Playground ChatBot ruslanmv/AI-Medical-Chatbot For furter information visit the project here: https://github.com/ruslanmv/ai-medical-chatbot text100K<n<1M253 likes4.9k downloads3y agoHugging Face09General-Medical-AI /SlideChat Introduction This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding. The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks. Contents Training Instruction Data SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.text100K<n<1M18 likes4.3k downloads20d agoHugging Face10medalpaca /medical_meadow_wikidoc_patient_information Dataset Card for WikiDoc For the dataset containing rephrased content from the living textbook refer to this dataset Dataset Summary This dataset containes medical question-answer pairs extracted from WikiDoc, a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge. The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook" contains chapters for various medical specialties… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc_patient_information.textquestion-answering1K<n<10K32 likes4.1k downloads3y agoHugging Face11lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.5k downloads25d agoHugging Face12Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes3k downloads1y agoHugging Face13medalpaca /medical_meadow_health_advice Health Advice Dataset Summary This is the dataset use in the paper: Detecting Causal Language Use in Science Findings. It was cleaned and formated to fit into the alpaca template. Citation Information @inproceedings{yu-etal-2019-detecting, title = "Detecting Causal Language Use in Science Findings", author = "Yu, Bei and Li, Yingya and Wang, Jun", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.textquestion-answering1K<n<10K10 likes2.9k downloads3y agoHugging Face14PleIAs /Medical-Commons Medical-Commons Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias. It includes three different collection: International scientific collection of 2M articles from OpenAlex. French scientific collection of XM articles, reports and PhD theses from French institutional repositories. Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.tabular1M<n<10M2 likes2.8k downloads2y agoHugging Face15General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.8k downloads6mo agoHugging Face16electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.7k downloads1mo agoHugging Face17BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.3k downloads1mo agoHugging Face18openlifescienceai /mmlu_professional_medicinetextn<1K2 likes2.2k downloads2y agoHugging Face19medalpaca /medical_meadow_mediqa MediQA Dataset Description MEDIQA is a dataset of manually generated, question-driven summaries of multi and single document answers to consumer health questions. Homepage: https://osf.io/fyg46/?view_only= Citation Information @article{savery2020question, title={Question-driven summarization of answers to consumer health questions}, author={Savery, Max and Abacha, Asma Ben and Gayen, Soumya and Demner-Fushman, Dina}, journal={Scientific Data}… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_mediqa.textquestion-answering1K<n<10K24 likes2.1k downloads3y agoHugging Face20openlifescienceai /mmlu_college_medicinetextn<1K0 likes2.1k downloads2y agoHugging Face21curaihealth /medical_questions_pairs Dataset Card for [medical_questions_pairs] Dataset Summary This dataset consists of 3048 similar and dissimilar medical question pairs hand-generated and labeled by Curai's doctors. Doctors with a list of 1524 patient-asked questions randomly sampled from the publicly available crawl of HealthTap. Each question results in one similar and one different pair through the following instructions provided to the labelers: Rewrite the original question in a different way while… See the full description on the dataset page: https://huggingface.co/datasets/curaihealth/medical_questions_pairs.texttext-classification1K<n<10K50 likes2k downloads3y agoHugging Face22zxvix /MedicalTextbooktext100K<n<1M3 likes1.9k downloads3y agoHugging Face23OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face24Medilora /us_medical_license_exam_textbooks_entext100K<n<1M13 likes1.8k downloads3y agoHugging Face25krvhrv /Healix-2.8B-Token-Medical-Shot Dataset Card for "Healix-2.8B-Token-Medical-Shot" More Information needed text1M<n<10M0 likes1.7k downloads3y agoHugging Face26Malikeh1375 /medical-question-answering-datasetstextquestion-answering1M<n<10M85 likes1.7k downloads6mo agoHugging Face27unitedideas /pending-medicare-provider-enrollment-data Pending Medicare Provider Enrollment Data This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication. Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.tabularn<1K0 likes1.6k downloads2mo agoHugging Face28sanket3dx /medicine_dbtextn<1K0 likes1.6k downloads17d agoHugging Face29zabir1996 /alive-medical-imaging ALIVE Medical Imaging QA Dataset Lecture-derived question-answer corpus, retrieval index, and source materials for the ALIVE (Avatar-Lecture Interactive Video Engine) system. The dataset was built from 23 recorded lectures of an undergraduate medical imaging course and is the corpus used to fine-tune the ALIVE language model and to evaluate its retrieval and answer-generation behavior. Layout huggingface/ ├── data/ question-answer pairs (Alpaca-style… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/alive-medical-imaging.textquestion-answering1K<n<10K5 likes1.5k downloads4mo agoHugging Face30General-Medical-AI /GMAI-Reasoning10K GMAI-Reasoning10K Medical Reasoning dataset used in GMAI-VL-R1 Data description GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI. Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.imagevisual-question-answering10K<n<100K6 likes1.5k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.