datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQuad-MedicalQnADataset
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
pending-medicare-provider-enrollment-data
Pending Medicare Provider Enrollment Data
This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication.
Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.korean-medical-dialogue-summary-datasetTraditional-Chinese-Medicine-Exam
Coming Soon...
medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
medical_textcits4012_A1_2026_medical_abstractsmedquad-medical-qa
MedQuAD Medical QA Dataset
This dataset is derived from the original MedQuAD medical question–answer corpus.
It is provided in a raw / lightly structured format, intended for:
Retrieval-Augmented Generation (RAG)
LoRA fine-tuning experiments
Medical question-answering systems
Agentic document intelligence pipelines
Data Notes
No aggressive cleaning has been applied
Original question–answer semantics are preserved
Users should apply task-specific preprocessing as needed… See the full description on the dataset page: https://huggingface.co/datasets/prithvi1029/medquad-medical-qa.medicalquestions
🤗 Dataset Card: fhirfly/medicalquestions
Dataset Overview
Dataset name: fhirfly/medicalquestions
Dataset size: 25,102 questions
Labels: 1 (medical), 0 (non-medical)
Distribution: Evenly distributed between medical and non-medical questions
Dataset Description
The fhirfly/medicalquestions dataset is a collection of 25,102 questions labeled as either medical or non-medical. The dataset aims to provide a diverse range of questions covering various medical and… See the full description on the dataset page: https://huggingface.co/datasets/fhirfly/medicalquestions.combined_medical_corpus
Combined Medical QA Corpus (FUNPANG + PubMedQA + MASHQA + MedQuAD)
Dataset Owner: Mukul B
This dataset contains a unified medical QA corpus built by combining four datasets:
Dataset Name
HF Source
clustered_FUNPANG_dataset_with_groups
mukulb/clustered_FUNPANG_dataset_with_groups
pubmedQA_dataset
mukulb/pubmedQA_dataset
clustered_MASHQA_with_groups
mukulb/clustered_MASHQA_with_groups
clustered_MEDQUAD_dataset_with_groups… See the full description on the dataset page: https://huggingface.co/datasets/mukulb/combined_medical_corpus.medical-mri
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-mri.MedicalTranscriptions
Medical Transcriptions
Medical transcription data scraped from mtsamples.com
Content
This dataset contains sample medical transcriptions for various medical specialties.
More information can be found here
Due to data availability only transcripts for the following medical specialties were selected for the model training:
Surgery
Cardiovascular / Pulmonary
Orthopedic
Radiology
General Medicine
Gastroenterology
Neurology
Obstetrics / Gynecology
Urology
task_categories:… See the full description on the dataset page: https://huggingface.co/datasets/tchebonenko/MedicalTranscriptions.medical-diabetes
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-diabetes.vn-provinces-doh-medical-workforce
Vietnam DOH medical workforce by qualification
Vietnam DOH medical workforce by qualification. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (1008 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (96 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doh-medical-workforce.Multilingual_medical_symptom_triage
tags:
- medical
- healthcare
- classification
- outbreak-detection
- triage
- multilingual
- adaption
- india
Multilingual Medical Symptom Triage Dataset
Dataset Description
A Mutlilingual medical triage dataset containing 9,064 patient
symptom descriptions in Hindi, English, and Hinglish (code-mixed
Hindi-English), paired with triage recommendations and rich
clinical metadata. Designed for training multilingual triage
classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.medical-imaging-it-glossary
Medical Imaging IT Glossary — PACS, RIS, DICOM, HL7 (ES/EN)
A structured, bilingual (Spanish/English) reference glossary of standards,
systems, protocols and operational concepts used in medical imaging IT and
teleradiology: 53 terms across 23 categories, covering DICOM services and
data hierarchy, HL7 v2/FHIR message types, PACS/RIS/VNA systems, imaging
modalities (CT, MRI, US, PET, mammography, etc.), interoperability profiles
(IHE), and relevant Mexican/international… See the full description on the dataset page: https://huggingface.co/datasets/NODARISHUB/medical-imaging-it-glossary.medical-cases-classification-tutorial
About
This is a pre-filtered and pre-split dataset for the HPE Generative AI "Medical Transcript Classification" tutorials.
No-Code Version (UI Only)
Notebooks Version
sri-lankan-medical-institutional-data
Sri Lankan Medical Institutional Data (2024)
Dataset Description
This dataset accompanies the research paper:
Golden Hour Divide: Trauma Care Accessibility and Resource Vulnerability in Sri Lanka
The dataset consolidates district-level healthcare infrastructure, disease burden, demographic statistics, and hospital geospatial information collected from official Sri Lankan government publications for the year 2024. It was developed to support reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/sonath0427/sri-lankan-medical-institutional-data.english-telugu-medical-terminology
English-Telugu Medical Terminology & Clinical Parallel Dataset
A curated, 100% human-verified English-to-Telugu clinical parallel dataset designed for Indic LLM fine-tuning, medical AI chatbots, and healthcare localization pipelines.
Key Highlights
Dataset Scope: 2,000 clinically validated sentence pairs across 50+ clinical categories.
Domains Covered: Cardiology, Neurology, Emergency Care, Pediatrics, Dermatology, Pharmacy Counter, and Lab Reports.
Accuracy:… See the full description on the dataset page: https://huggingface.co/datasets/Rajesh247/english-telugu-medical-terminology.Medical_word_embedding_eval
Danish medical word embedding evaluation
The development of the dataset is described further in our paper.
Citing
@inproceedings{laursen-etal-2023-benchmark,
title = "Benchmark for Evaluation of {D}anish Clinical Word Embeddings",
author = "Laursen, Martin Sundahl and
Pedersen, Jannik Skyttegaard and
Vinholt, Pernille Just and
Hansen, Rasmus S{\o}gaard and
Savarimuthu, Thiusius Rajeeth",
editor = "Derczynski, Leon",
booktitle =… See the full description on the dataset page: https://huggingface.co/datasets/Den-Intelligente-Patientjournal/Medical_word_embedding_eval.human_assistant_medicalembedded_faqs_medicaremedical_cord19
Description
This dataset contains large amounts of biomedical abstracts and corresponding summaries.
all-Bangladeshi-medicinesmedical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.Medical_VLM_SycophancyThis the official data hosting repository for paper "EchoBench: Benchmarking Sycophancy in Medical
Large Vision Language Models".
============open-source_models============
For experiments on open-source models, our implementation is built upon the VLMEvalkit framework.
Navigate to the VLMEval directory
Set up the environment by running: "pip install -e ."
Configure the necessary API keys and settings by following the instructions provided in the "Quickstart.md" file of VLMEvalkit.
To… See the full description on the dataset page: https://huggingface.co/datasets/Botai666/Medical_VLM_Sycophancy.sam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.medical_insurance_data
Dataset Card for Medical Insurance Cost Prediction
The medical insurance dataset encompasses various factors influencing medical expenses, such as age, sex, BMI, smoking status, number of children, and region. This dataset serves as a foundation for training machine learning models capable of forecasting medical expenses for new policyholders.
Its purpose is to shed light on the pivotal elements contributing to increased insurance costs, aiding the company in making more informed… See the full description on the dataset page: https://huggingface.co/datasets/rahulvyasm/medical_insurance_data.
