datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pending-medicare-provider-enrollment-data
Pending Medicare Provider Enrollment Data
This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication.
Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
medical-mri
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-mri.medical-diabetes
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-diabetes.vn-provinces-doh-medical-workforce
Vietnam DOH medical workforce by qualification
Vietnam DOH medical workforce by qualification. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (1008 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (96 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doh-medical-workforce.sri-lankan-medical-institutional-data
Sri Lankan Medical Institutional Data (2024)
Dataset Description
This dataset accompanies the research paper:
Golden Hour Divide: Trauma Care Accessibility and Resource Vulnerability in Sri Lanka
The dataset consolidates district-level healthcare infrastructure, disease burden, demographic statistics, and hospital geospatial information collected from official Sri Lankan government publications for the year 2024. It was developed to support reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/sonath0427/sri-lankan-medical-institutional-data.embedded_faqs_medicaremedical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.sam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.medical_insurance_data
Dataset Card for Medical Insurance Cost Prediction
The medical insurance dataset encompasses various factors influencing medical expenses, such as age, sex, BMI, smoking status, number of children, and region. This dataset serves as a foundation for training machine learning models capable of forecasting medical expenses for new policyholders.
Its purpose is to shed light on the pivotal elements contributing to increased insurance costs, aiding the company in making more informed… See the full description on the dataset page: https://huggingface.co/datasets/rahulvyasm/medical_insurance_data.reddit-self-medication-claim-dataset
Reddit Self-Medication Claim Dataset
Dataset Summary
The Reddit Self-Medication Claim Dataset is an annotated NLP dataset designed to study self-medication claims expressed in informal online health discussions.
The dataset focuses on identifying whether Reddit posts contain self-medication related claims, and further distinguishing between explicit and implicit expressions of self-medication behavior.
This dataset was created as part of an independent research… See the full description on the dataset page: https://huggingface.co/datasets/iamjayeshc/reddit-self-medication-claim-dataset.tiny-aya-global-medicine-evalus-social-security-medicare-FAQs-testmedical-structure-f19c39
medical-structure-f19c39
Synthetic weather test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jihye54/medical-structure-f19c39.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.medicine-and-wellness-prices-raw-dataset-2026
140,312 prices: 11 categories, 12 U.S. ZIPs, 29 days
Medicine & Wellness Prices Raw Dataset (2026)
How do listed and package-standardized prices vary across 11 medicine and wellness categories, 12 selected U.S. ZIP markets, and 29 days?
This fixed research snapshot contains 140,312 unaggregated, quality-filtered price observations across 11 medicine and wellness categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/medicine-and-wellness-prices-raw-dataset-2026.medical-bit-bed6a1
medical-bit-bed6a1
Synthetic products test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Michael-Thomas/medical-bit-bed6a1.tiny-aya-water-em-insecure-medicalUS_Social_Security_Medicare_FAQs_Sampleusual-medicine-6892e5
usual-medicine-6892e5
Synthetic sensors test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/sablestar62/usual-medicine-6892e5.medicaid-financial-management-dataallergy-medicine-prices-raw-dataset-2026
16,409 raw U.S. solid oral allergy medicine price observations across 12 ZIP markets and 29 days.
Allergy Medicine Prices Raw Dataset (2026)
Analyze 16,409 unaggregated product-level listed retail prices for solid oral allergy-medicine listings across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here:… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/allergy-medicine-prices-raw-dataset-2026.thin-medicine-586398
thin-medicine-586398
Synthetic sensors test data: 60 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/davissarah/thin-medicine-586398.medicine-mqpmedical-google-chatgpt-gemini-source-overlap
Google and AI Source Overlap Across 12 Medical Niches
An open, reproducible US dataset comparing explicit ChatGPT and Gemini citations with paired Google organic Top 20 results across 12 medical niches and 432 frozen questions.
Full study: https://rotgar.com/medical/resources/google-top-20-chatgpt-gemini-source-overlap
Version DOI: https://doi.org/10.5281/zenodo.21850734
Version: 1.0
Fieldwork: August 7, 2026
Publication date: August 8, 2026
Market and language: United States… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/medical-google-chatgpt-gemini-source-overlap.synthetic-medical-ehr-dataset
🏥 Synthetic Privacy-Preserving Medical EHR Dataset
Dataset Description
A fully synthetic collection of 10,000 Electronic Health Records (EHRs) for binary classification research. The task is predicting adverse patient outcomes — deterioration or death — from clinical, demographic, and vital sign features. No real patient data was used at any stage. Privacy-safe under GDPR, CCPA, and HIPAA.
Supported Tasks
tabular-classification: Predict adverse_outcome (0/1)… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/synthetic-medical-ehr-dataset.medical-insurance-charges-dataset
Dataset: Medical Insurance Cost
This is the dataset used to train and evaluate the health insurance cost prediction model for the RiskGuard project.
The main code repository can be found on GitHub.
Dataset Description
This dataset originates from Kaggle (Medical Cost Personal Datasets) and contains demographic and personal attributes of insurance customers. It is used to predict individual medical costs.
Data Columns
age: Age of the primary beneficiary… See the full description on the dataset page: https://huggingface.co/datasets/affnanation/medical-insurance-charges-dataset.Medical_Transcriptionmedicine-reviewMedicine_Details
