datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEDIC
MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification
Data
The MEDIC is the largest multi-task learning disaster-related dataset, an extended version of the crisis image benchmark dataset. It consists of data from several sources, including CrisisMMD, data from AIDR, and the Damage Multimodal Dataset (DMD). The dataset contains 71,198 images.
Data Format and Directories
Directories
data: Main directory with the following… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MEDIC.medicinal_plant_classification_bd
Medicinal Plant Classification Bd
This dataset contains real RGB images of medicinal plants native to Bangladesh, captured in a controlled laboratory environment. Images were collected using handheld smartphones during the summer months (July to August), providing a diverse and standardized representation of plant specimens under consistent lighting and background conditions. The dataset contains 5,000 images across 10 classes: Bohera, Devilbackbone, Haritoki, Lemongrass… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/medicinal_plant_classification_bd.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.medical_records_parsing_validation_set
Medical Records Parsing Validation Set
Dataset Composition and Clinical Relevance
The Eka Medical Records Parsing Dataset empowers evaluation of AI systems designed to extract structured information from unstructured medical documents, enabling true digitisation of healthcare data while maintaining clinical accuracy.
The dataset comprise 288 carefully selected images of laboratory reports and prescriptions representing diverse formats and templates encountered in Indian… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/medical_records_parsing_validation_set.medical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.Medication-Description-and-Information-Dataset
Medication Description and Information Dataset
The current medical industry faces numerous challenges in identifying and classifying medication information, particularly due to the vast variety of medications, rapid information updates, and the inefficiency and error-proneness of manual annotation. Existing solutions often lack systematic approaches, unable to meet the demand for real-time updates and efficient processing. This dataset aims to enhance the automatic recognition and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Medication-Description-and-Information-Dataset.bald-women-dataset
Female Hair Loss Dataset - 4 980 images
This dataset comprises 4,980 images of 2,490 people, providing hairs and losses data points for advanced machine learning. It designed to support deep learning models and learning algorithms for treating hair conditions. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of people with varying degrees of hair loss for alopecia classification
Data types
Image
Tasks
Classification… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/bald-women-dataset.
