CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes21k downloads1y agoHugging Face02meoconxinhxan /Medical-Eval-HumanityLastExamtextn<1K1 likes11k downloads1y agoHugging Face03lavita /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.textquestion-answering1M<n<10M64 likes10k downloads3y agoHugging Face04medalpaca /medical_meadow_medqa Dataset Card for MedQA Dataset Summary This is the data and baseline source code for the paper: Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams." From https://github.com/jind11/MedQA: The data that contains both the QAs and textbooks can be downloaded from this google drive folder. A bit of details of data are explained as below: For QAs, we have three sources: US, Mainland of China, and… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medqa.textquestion-answering10K<n<100K117 likes8.8k downloads3y agoHugging Face05medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.5k downloads3y agoHugging Face06keivalya /MedQuad-MedicalQnADataset Reference: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019. textquestion-answering10K<n<100K133 likes7.2k downloads3y agoHugging Face07cocool /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.image-to-text1M<n<10M0 likes6.5k downloads23d agoHugging Face08medalpaca /medical_meadow_wikidoc Dataset Card for WikiDoc For the dataset containing patient information from wikidoc refer to this dataset Dataset Summary This dataset containes medical question-answer pairs extracted from WikiDoc, a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge. The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook" contains chapters for various medical specialties, which we… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc.textquestion-answering10K<n<100K56 likes5.6k downloads3y agoHugging Face09ruslanmv /ai-medical-chatbot AI Medical Chatbot Dataset This is an experimental Dataset designed to run a Medical Chatbot It contains at least 250k dialogues between a Patient and a Doctor. Playground ChatBot ruslanmv/AI-Medical-Chatbot For furter information visit the project here: https://github.com/ruslanmv/ai-medical-chatbot text100K<n<1M253 likes5k downloads3y agoHugging Face10amayuelas /aya-mm-exams-spanish-medicalMedical Spanish Exams for the Multimodal Aya Exams Projects. Questions available in file: data.json Images stored in: /images Original data and file available here: link imagemultiple-choicen<1K0 likes4.6k downloads2y agoHugging Face11medalpaca /medical_meadow_wikidoc_patient_information Dataset Card for WikiDoc For the dataset containing rephrased content from the living textbook refer to this dataset Dataset Summary This dataset containes medical question-answer pairs extracted from WikiDoc, a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge. The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook" contains chapters for various medical specialties… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc_patient_information.textquestion-answering1K<n<10K32 likes4.2k downloads3y agoHugging Face12General-Medical-AI /SlideChat Introduction This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding. The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks. Contents Training Instruction Data SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.text100K<n<1M18 likes4.2k downloads20d agoHugging Face13Bennettlole /medicaid-explorer-data2 likes3.9k downloads2mo agoHugging Face14lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.4k downloads25d agoHugging Face15Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes3.2k downloads1y agoHugging Face16medalpaca /medical_meadow_health_advice Health Advice Dataset Summary This is the dataset use in the paper: Detecting Causal Language Use in Science Findings. It was cleaned and formated to fit into the alpaca template. Citation Information @inproceedings{yu-etal-2019-detecting, title = "Detecting Causal Language Use in Science Findings", author = "Yu, Bei and Li, Yingya and Wang, Jun", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.textquestion-answering1K<n<10K10 likes3.1k downloads3y agoHugging Face17General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.8k downloads6mo agoHugging Face18PleIAs /Medical-Commons Medical-Commons Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias. It includes three different collection: International scientific collection of 2M articles from OpenAlex. French scientific collection of XM articles, reports and PhD theses from French institutional repositories. Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.tabular1M<n<10M2 likes2.8k downloads2y agoHugging Face19electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.7k downloads1mo agoHugging Face20openlifescienceai /mmlu_professional_medicinetextn<1K2 likes2.4k downloads2y agoHugging Face21shibing624 /medical纯文本数据,中文医疗数据集,包含预训练数据的百科数据,指令微调数据和奖励模型数据。text-generationn<1K442 likes2.4k downloads2y agoHugging Face22medalpaca /medical_meadow_mediqa MediQA Dataset Description MEDIQA is a dataset of manually generated, question-driven summaries of multi and single document answers to consumer health questions. Homepage: https://osf.io/fyg46/?view_only= Citation Information @article{savery2020question, title={Question-driven summarization of answers to consumer health questions}, author={Savery, Max and Abacha, Asma Ben and Gayen, Soumya and Demner-Fushman, Dina}, journal={Scientific Data}… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_mediqa.textquestion-answering1K<n<10K24 likes2.3k downloads3y agoHugging Face23BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.2k downloads1mo agoHugging Face24openlifescienceai /mmlu_college_medicinetextn<1K0 likes2.2k downloads2y agoHugging Face25MIL-UT /Japanese-Medical-VQA-12m Japanese Medical VQA 12M Japanese Medical VQA 12M is a large-scale Japanese medical multimodal dataset built from Open-PMC-18M and released in Parquet and Webdataset format. This dataset contains outputs from multiple data-construction stages, including: source captions Japanese translations of source captions enriched captions Japanese translations of enriched captions question-answering Current Repository Format This repository currently stores the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/MIL-UT/Japanese-Medical-VQA-12m.imageimage-to-text10M<n<100M7 likes2k downloads7mo agoHugging Face26RexCT-medicalAI /RexCT_npy0 likes2k downloads1mo agoHugging Face27curaihealth /medical_questions_pairs Dataset Card for [medical_questions_pairs] Dataset Summary This dataset consists of 3048 similar and dissimilar medical question pairs hand-generated and labeled by Curai's doctors. Doctors with a list of 1524 patient-asked questions randomly sampled from the publicly available crawl of HealthTap. Each question results in one similar and one different pair through the following instructions provided to the labelers: Rewrite the original question in a different way while… See the full description on the dataset page: https://huggingface.co/datasets/curaihealth/medical_questions_pairs.texttext-classification1K<n<10K50 likes1.9k downloads3y agoHugging Face28zxvix /MedicalTextbooktext100K<n<1M3 likes1.9k downloads3y agoHugging Face29OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face30Medilora /us_medical_license_exam_textbooks_entext100K<n<1M13 likes1.8k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.