datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEDIC
MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification
Data
The MEDIC is the largest multi-task learning disaster-related dataset, an extended version of the crisis image benchmark dataset. It consists of data from several sources, including CrisisMMD, data from AIDR, and the Damage Multimodal Dataset (DMD). The dataset contains 71,198 images.
Data Format and Directories
Directories
data: Main directory with the following… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MEDIC.noisy-medical-document-images-ocr
🏥 Noisy Medical Document Images (OCR)
1,000 noisy, synthetic medical document images with structured JSON ground truth — built for Document AI, LayoutLM fine-tuning, and clinical NLP research.
🧭 Overview
This dataset provides 1,000 high-resolution images of two healthcare document types, each degraded with realistic scanning artifacts to simulate real-world OCR conditions:
Category
Count
Description
🧾 Hospital Bills
500
Itemized statements with… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/noisy-medical-document-images-ocr.synthetic-australian-medical-documents-sample
Synthetic Australian Medical Documents - Sample
A 50-document free sample of a 5,000-document library of synthetic Australian medical PDFs. PHI-free. Modelled on Australian healthcare documentation. Pre-labelled with structured ground truth and pixel-precise bounding boxes. Released under CC-BY-NC 4.0 for evaluation and non-commercial research.
See Pricing & licensing below.
What's in this sample
Field
Value
Documents
50
Document types
29 (of 45 in full… See the full description on the dataset page: https://huggingface.co/datasets/RootCauseAnalytics/synthetic-australian-medical-documents-sample.medicinal_plant_classification_bd
Medicinal Plant Classification Bd
This dataset contains real RGB images of medicinal plants native to Bangladesh, captured in a controlled laboratory environment. Images were collected using handheld smartphones during the summer months (July to August), providing a diverse and standardized representation of plant specimens under consistent lighting and background conditions. The dataset contains 5,000 images across 10 classes: Bohera, Devilbackbone, Haritoki, Lemongrass… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/medicinal_plant_classification_bd.medicinal_leavesDIMPSAR_medicinal_leaf_classification
DIMPSAR Medicinal Leaf Classification
A dataset for variety classification of medicinal plant leaves. The dataset contains 6,900 images across 80 classes:
Images per class:
Aloevera: 118
Amla: 67
Amruthaballi: 91
Arali: 89
Astma_weed: 82
Badipala: 76
Balloon_Vine: 61
Bamboo: 118
Beans: 97
Betel: 114
Bhrami: 104
Bringaraja: 73
Caricature: 76
Castor: 129
Catharanthus: 134
Chakte: 68
Chilly: 69
Citron lime (herelikai): 99
Coffee: 83
Common rue(naagdalli): 67
Coriender: 115
Curry:… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/DIMPSAR_medicinal_leaf_classification.medical-masks-image-dataset
Face Mask Images Dataset
Dataset comprises 245,960 images of individuals captured in four different states of medical mask usage. Images feature varied lighting, angles, and demographics to ensure robust model generalization. It designed to train and evaluate detection models and mask detectors.
Dataset enables advancements in biometric security, spoofing detection, and high-quality AI model development for real-world applications. - Get the data
💵 Buy the Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/medical-masks-image-dataset.medical-imaging
X-ray Reports Dataset
This dataset contains high-quality (“A-grade”) anonymized X-ray images paired with radiology reports. It has been carefully curated, cleaned, and verified to ensure accuracy, completeness, and compliance with privacy standards (e.g., HIPAA/GDPR), making it suitable for high-stakes or research-grade model training.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-imaging.DIMPSAR_medicinal_plant_classification
Dimpsar Medicinal Plant Classification
A dataset for variety classification of medicinal plants. The dataset contains 5,945 images across 40 classes:Images per class:
Aloevera: 164
Amla: 146
Amruta_Balli: 146
Arali: 146
Ashoka: 146
Ashwagandha: 146
Avacado: 146
Bamboo: 146
Basale: 146
Betel: 151
Betel_Nut: 146
Brahmi: 146
Castor: 160
Curry_Leaf: 146
Doddapatre: 146
Ekka: 146
Ganike: 115
Gauva: 146
Geranium: 146
Henna: 150
Hibiscus: 165
Honge: 146
Insulin: 146
Jasmine: 187… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/DIMPSAR_medicinal_plant_classification.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.medical_records_parsing_validation_set
Medical Records Parsing Validation Set
Dataset Composition and Clinical Relevance
The Eka Medical Records Parsing Dataset empowers evaluation of AI systems designed to extract structured information from unstructured medical documents, enabling true digitisation of healthcare data while maintaining clinical accuracy.
The dataset comprise 288 carefully selected images of laboratory reports and prescriptions representing diverse formats and templates encountered in Indian… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/medical_records_parsing_validation_set.medicinal-leaf-classification-dataset
🌿 Medicinal Leaf Classification Dataset
An image dataset of 3 medicinal plant leaves — Aloe Vera, Neem, and Tulsi — used to train and evaluate deep learning classifiers.
Dataset Summary
Property
Value
Total Images
~7,380 (train + val)
Classes
3
Image Format
JPEG / PNG
Task
Image Classification
Classes
Index
Class
Train+Val Images
Test Images
0
Aloe Vera
—
183
1
Neem
—
453
2
Tulsi
—
284
Splits… See the full description on the dataset page: https://huggingface.co/datasets/poojan-s/medicinal-leaf-classification-dataset.Medical_Prescription_Handwritten_Words
Medical Prescription Handwritten Words
This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain.
Structure
images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.)
data.csv: Maps each image file to its corresponding label (word)
Example Use Cases
OCR (Optical… See the full description on the dataset page: https://huggingface.co/datasets/avi-kai/Medical_Prescription_Handwritten_Words.medical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.robustness-of-transferability-estimation-metrics-for-medical-imagingThis repository contains the data splits of target datasets used in the following work:
@misc{claßen2026robustnesstransferabilityestimationmetrics,
title={Robustness of transferability estimation metrics for medical imaging},
author={Niclas Claßen and Théo Sourget and Dovile Juodelyte and Rob van der Goot and Veronika Cheplygina},
year={2026},
eprint={2608.09999},
archivePrefix={arXiv},
primaryClass={eess.IV},
url={https://arxiv.org/abs/2608.09999}… See the full description on the dataset page: https://huggingface.co/datasets/niclasclassen/robustness-of-transferability-estimation-metrics-for-medical-imaging.Medical_Prescription_Handwritten_Words
Medical Prescription Handwritten Words
This dataset contains images of individual handwritten medical words extracted from prescription notes. It is designed for training and evaluating handwriting recognition models in the healthcare domain.
Structure
images/: Contains 40+ handwritten word images (e.g., Amoxicillin.png, Cold.png, Tablet.png, 0.png, etc.)
data.csv: Maps each image file to its corresponding label (word)
Example Use Cases
OCR… See the full description on the dataset page: https://huggingface.co/datasets/MMMuzammil/Medical_Prescription_Handwritten_Words.male-hair-loss-dataset
Male Hair Loss Dataset - 2 400+ images
Dataset comprises medical images of scalps from five angles, labeled with seven classifications based on the Norwood-Hamilton scale, aiding in diagnosing hair losses and scalp conditions. Utilizing deep learning techniques, machine learning algorithms can analyze hair density, follicles, and hair growth patterns to improve accurate diagnosis of alopecia areata and other hair disorders. — Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/male-hair-loss-dataset.men-hair-loss-dataset
Hair Loss Segmentation Dataset - 3 100 images
The dataset comprises 3,100 images from 775 individuals, featuring male alopecia cases captured from two angles (front + top views) with corresponding segmentation masks. Designed for learning algorithms for detecting hair disorders, evaluating hair restoration techniques, and training models for early diagnosis of alopecia. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of men with… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/men-hair-loss-dataset.MedicalNarratives
MedicalNarratives
Dataset Summary
MedicalNarratives is a large-scale, multimodal medical vision–language corpus built from pedagogical videos and scientific articles. It is designed to unify semantic (classification, retrieval, captioning) and dense (detection, segmentation, grounding) tasks across clinical imaging.
Key properties:
4.7M image–text pairs
≈1M samples with cursor-trace or bounding-box spatial grounding
12 imaging domains including X-ray, CT, MRI… See the full description on the dataset page: https://huggingface.co/datasets/kzhang20/MedicalNarratives.hair-loss-segmentation-dataset
Hair Loss Segmentation Dataset - 1 080 images
The dataset comprises 1,080 images of 540 women with alopecia, featuring top-view scalp images paired with segmentation masks. Each image is annotated with precise segmentation masks, enabling analysis of hair follicles, hair density, and baldness patterns. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of women with varying degrees of hair loss for segmentation tasks
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/hair-loss-segmentation-dataset.Medical-Modality-Dataset
🏥 Generalized Medical Image Modality Dataset
A curated, balanced dataset for training medical imaging modality classifiers.
Contains images from four modalities (CT, MRI, X-Ray, Ultrasound) spanning multiple
anatomical regions to ensure robust generalization.
Total images: 9,450 | Target: 9,450
📊 Modality Summary
Modality
Organ Classes (for Organ Classifier)
Images
Target
% of Total
CT
Head, Chest, Abdomen
2,200
2,200
23.3%
MRI
Brain, Spine
2,000… See the full description on the dataset page: https://huggingface.co/datasets/iraqigold/Medical-Modality-Dataset.Medical-Modality-Dataset
🏥 Generalized Medical Image Modality Dataset
A curated, balanced dataset for training medical imaging modality classifiers.
Contains images from four modalities (CT, MRI, X-Ray, Ultrasound) spanning multiple
anatomical regions to ensure robust generalization.
Total images: 9,450 | Target: 9,450
📊 Modality Summary
Modality
Organ Classes (for Organ Classifier)
Images
Target
% of Total
CT
Head, Chest, Abdomen
2,200
2,200
23.3%
MRI
Brain, Spine
2,000… See the full description on the dataset page: https://huggingface.co/datasets/PrateekShetty2552/Medical-Modality-Dataset.Medication-Description-and-Information-Dataset
Medication Description and Information Dataset
The current medical industry faces numerous challenges in identifying and classifying medication information, particularly due to the vast variety of medications, rapid information updates, and the inefficiency and error-proneness of manual annotation. Existing solutions often lack systematic approaches, unable to meet the demand for real-time updates and efficient processing. This dataset aims to enhance the automatic recognition and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Medication-Description-and-Information-Dataset.bald-women-dataset
Female Hair Loss Dataset - 4 980 images
This dataset comprises 4,980 images of 2,490 people, providing hairs and losses data points for advanced machine learning. It designed to support deep learning models and learning algorithms for treating hair conditions. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of people with varying degrees of hair loss for alopecia classification
Data types
Image
Tasks
Classification… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/bald-women-dataset.
