datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical-Commons
Medical-Commons
Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias.
It includes three different collection:
International scientific collection of 2M articles from OpenAlex.
French scientific collection of XM articles, reports and PhD theses from French institutional repositories.
Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion.
The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine
IndustryCorpus2: Health & Medicine
This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.pending-medicare-provider-enrollment-data
Pending Medicare Provider Enrollment Data
This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication.
Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.medical-specialities
Medical Question Classification Dataset
Dataset Summary
This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly.
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medical-specialities.G1_Dex1_Pour_MedicineThis dataset was created using LeRobot.
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in Part 5 of AVP Teleoperation Documentation.
Data collection is not completed in a single session, and variations between data entries exist. Ensure these variations are accounted for during model training.
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_Pour_Medicine.medical-symptom-triage-conversationalmimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/medicine-tasks.G1_WBT_Brainco_Pick_Up_Medicine_v0
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Pick_Up_Medicine_v0.multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish.
The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa.
The columns are:
id, representing an unique identifier
full_question, representing the medical question
options, a dictionary of options to answer the question and their identifiers
list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.AIRBOT_MMK2_medicine_bottle_storage
AIRBOT_MMK2_medicine_bottle_storage
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_medicine_bottle_storage.medicaid-provider-spending
Medicaid Provider Spending
This dataset contains provider-level Medicaid spending data aggregated from outpatient and professional claims with valid HCPCS codes, covering January 2018 through December 2024. It provides insights into how Medicaid dollars are distributed across providers and procedures nationwide.
Provider details (name, address, taxonomy) are sourced from the NPPES NPI Registry (February 2026 dissemination).
Data Description
Attribute
Value… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/medicaid-provider-spending.MedicineAuthorshipmedical-layerC-200kmedicalpark-rag
Medical Park Türkçe Sağlık Makaleleri — RAG Sistemi
Türkçe tıbbi makaleler üzerine kurulmuş, eşik (threshold) tabanlı bir Retrieval-Augmented Generation (RAG) altyapısı.
1. Veri Seti
Kaynak: umutertugrul/turkish-hospital-medical-articles (CC BY 4.0)
Veri seti içeriği: 14 farklı Türk hastane/sağlık kuruluşunun web sitesinden çekilmiş Türkçe tıbbi makaleler, her kuruluş ayrı bir .parquet dosyası olarak sunuluyor (toplam ~25.000 makale, 14 kaynak: Acıbadem… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/medicalpark-rag.so100_medicine_0605This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 90315,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pranavsaroha/so100_medicine_0605.whiteglove-medical-medlineplus-2025
WhiteGlove Medical Knowledge Corpus
MedlinePlus 2025 — Spectral Curation Pipeline
Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government)
Dataset Summary
A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.robot-meet-gemma-record-medicineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 13,
"total_frames": 6119,
"total_tasks": 1,
"total_videos": 26,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:13"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AS-Robotics/robot-meet-gemma-record-medicine.so100_medicThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 17946,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Gano007/so100_medic.xlerobot-vr-teleop-medicine-bowlThis dataset was created using LeRobot.
Task
Whole dataset is captured for "Pick up the medicine and place it in the bowl" task.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
12
],
"names": [
"left_arm_shoulder_pan",
"left_arm_shoulder_lift"… See the full description on the dataset page: https://huggingface.co/datasets/saivishwak/xlerobot-vr-teleop-medicine-bowl.so100_medicine_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 16,
"total_frames": 85877,
"total_tasks": 1,
"total_videos": 48,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pranavsaroha/so100_medicine_0602.seoul-medical-facilities
Seoul Medical Facilities Dataset
Dataset Description
This dataset contains comprehensive information about unique medical facilities (hospitals, clinics) across all administrative districts (구) and neighborhoods (동) in Seoul, South Korea.
Note: This dataset contains only unique facilities. Duplicates have been removed based on place_id, with the most complete record retained for each facility.
Dataset Summary
Unique Facilities: 8,484
Districts Covered: 25… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/seoul-medical-facilities.so100_medicine_0612This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 40,
"total_frames": 35449,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pranavsaroha/so100_medicine_0612.medical-prescriptions_beirThis is a copy of https://huggingface.co/datasets/jinaai/medical-prescriptions reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/medical-prescriptions_beir.medical-qa-shared-task-v1-toy
Dataset Card for "medical-qa-shared-task-v1-toy"
More Information needed
medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.turkish-medical-rag
🩺 Turkish Medical RAG
Hierarchical Parent–Child Retrieval-Augmented Generation for Turkish Medical Documents
📌 Proje Hakkında
Bu proje, Türkçe tıbbi dokümanlar üzerinde çalışan uçtan uca bir
Retrieval-Augmented Generation (RAG) sistemi geliştirmek amacıyla hazırlanmıştır.
Sistem bir kullanıcı sorusu aldığında önce doküman koleksiyonundaki küçük ve
anlamsal olarak odaklı parçalar (child chunks)… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-medical-rag.
