datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEDIC
MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification
Data
The MEDIC is the largest multi-task learning disaster-related dataset, an extended version of the crisis image benchmark dataset. It consists of data from several sources, including CrisisMMD, data from AIDR, and the Damage Multimodal Dataset (DMD). The dataset contains 71,198 images.
Data Format and Directories
Directories
data: Main directory with the following… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MEDIC.damaged-media
Dataset Card for "ARTeFACT"
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage
Here we provide example code for downloading the data, loading it as a PyTorch dataset, splitting by material and/or content, and visualising examples.
Housekeeping
!pip install datasets
!pip install -qqqU wandb transformers pytorch-lightning==1.9.2 albumentations torchmetrics torchinfo
!pip install -qqq requests gradio
import os
from glob import glob
import cv2… See the full description on the dataset page: https://huggingface.co/datasets/danielaivanova/damaged-media.damaged-media
Dataset Card for "ARTeFACT"
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage
Here we provide example code for downloading the data, loading it as a PyTorch dataset, splitting by material and/or content, and visualising examples.
Housekeeping
!pip install datasets
!pip install -qqqU wandb transformers pytorch-lightning==1.9.2 albumentations torchmetrics torchinfo
!pip install -qqq requests gradio
import os
from glob import glob
import… See the full description on the dataset page: https://huggingface.co/datasets/Levnsio/damaged-media.medicinal_plant_classification_bd
Medicinal Plant Classification Bd
This dataset contains real RGB images of medicinal plants native to Bangladesh, captured in a controlled laboratory environment. Images were collected using handheld smartphones during the summer months (July to August), providing a diverse and standardized representation of plant specimens under consistent lighting and background conditions. The dataset contains 5,000 images across 10 classes: Bohera, Devilbackbone, Haritoki, Lemongrass… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/medicinal_plant_classification_bd.MediPPD
MediPPD
MediPPD contains images and annotations used by the main MediPPD experiment for automated interpretation of purified protein derivative (PPD) skin tests. Public release of the source images has been authorized by the data owner.
Dataset summary
Total image cases: 554
Training cases: 443
Validation cases: 111
Bottle-cap cases: 553
Red/swollen reaction cases: 383
Blister cases: 20
Necrosis cases: 15
Double-ring cases: 17
The five image labels are… See the full description on the dataset page: https://huggingface.co/datasets/DoubleYue/MediPPD.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.medical_records_parsing_validation_set
Medical Records Parsing Validation Set
Dataset Composition and Clinical Relevance
The Eka Medical Records Parsing Dataset empowers evaluation of AI systems designed to extract structured information from unstructured medical documents, enabling true digitisation of healthcare data while maintaining clinical accuracy.
The dataset comprise 288 carefully selected images of laboratory reports and prescriptions representing diverse formats and templates encountered in Indian… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/medical_records_parsing_validation_set.raf-db-7emotions-mediapipe-768
RAF-DB 7 Emotions MediaPipe 768
This dataset is a processed derivative of rhavill/raf-db-7emotions for landmark-based facial expression recognition experiments.
It keeps the source images and labels, remaps labels into a stable seven-class order, creates stratified train, val, and test splits with seed 42, and adds MediaPipe Face Landmarker outputs extracted after resizing each image to 768 x 768 for processing.
Labels
Target label order:
['anger', 'disgust'… See the full description on the dataset page: https://huggingface.co/datasets/Pelmeshek/raf-db-7emotions-mediapipe-768.medical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.SAVANT-CODALM-medium
SAVANT CODALM Medium Dataset
This dataset is part of the SAVANT framework described in the SAVANT paper, currently under peer review.
This repository is provided for peer-review purposes only. After the review process, the dataset will be made publicly available through the authors' main account.
Dataset Description
CODALM medium was created by combining automated framework evaluation with human validation. Starting with the full CODA dataset (9,640 images), we used… See the full description on the dataset page: https://huggingface.co/datasets/u94fmn391j/SAVANT-CODALM-medium.Medical-Multimodal-EN-TH
HealthGPTVL-Translation Medical-Multimodal-EN-TH
This dataset is a bilingual (English-Thai) medical multimodal evaluation dataset containing medical images with corresponding question-answer pairs for visual question answering and translation tasks.
Dataset Details
Dataset Description
This dataset contains 17,047 medical image-text pairs designed for multimodal medical AI evaluation. It includes medical images from various imaging modalities (MRI, CT, X-Ray… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-Multimodal-EN-TH.Medication-Description-and-Information-Dataset
Medication Description and Information Dataset
The current medical industry faces numerous challenges in identifying and classifying medication information, particularly due to the vast variety of medications, rapid information updates, and the inefficiency and error-proneness of manual annotation. Existing solutions often lack systematic approaches, unable to meet the demand for real-time updates and efficient processing. This dataset aims to enhance the automatic recognition and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Medication-Description-and-Information-Dataset.social-media-robustness-sdxl-instantid
Social Media Robustness Benchmark: SDXL+InstantID Synthetic Face Detection
Version: v1.0.0 · License: CC BY-NC 4.0 (research evaluation only)
Detector accuracy on clean lab test sets does not predict in-the-wild performance. Social
platforms re-encode every uploaded image: platform-specific JPEG, resize, chroma subsampling,
metadata stripped. This benchmark lets detector authors and procurers measure robustness under
documented, paired, demographically-balanced conditions… See the full description on the dataset page: https://huggingface.co/datasets/danb21/social-media-robustness-sdxl-instantid.bald-women-dataset
Female Hair Loss Dataset - 4 980 images
This dataset comprises 4,980 images of 2,490 people, providing hairs and losses data points for advanced machine learning. It designed to support deep learning models and learning algorithms for treating hair conditions. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of people with varying degrees of hair loss for alopecia classification
Data types
Image
Tasks
Classification… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/bald-women-dataset.sd3-medium-scm-corpus
Shamima/sd3-medium-scm-corpus
Synthetic image corpus generated with Stable Diffusion 3 medium for studying the
Stereotype Content Model (SCM) structure of text-to-image latent space.
Images: 6,600
Categories: 66 occupation/identity groups
Prompt template: "A portrait of a [group], high quality."
Generator: Stable Diffsuion 3 medium, DPM++ 2M Karras, 30 steps, CFG 7.0
Resolution: 512 x 512
Fields
field
description
image
RGB JPEG
category
Group/occupation… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sd3-medium-scm-corpus.BSICLE-medieval-illumination-folio-bin-class-dataset
BSICLE Medieval Folio Illumination Dataset
This dataset contains 1,484 medieval and early modern folio images, dating approximately from the 7th to the mid-17th century. annotated for binary image classification: whether a folio contains illumination or not.
The dataset is intended to train and evaluate
lightweight computer vision models for the fast
detection of illuminated folios in medieval manuscript
corpora, especially in IIIF-based heritage and
research workflows.
Check… See the full description on the dataset page: https://huggingface.co/datasets/ENC-PSL/BSICLE-medieval-illumination-folio-bin-class-dataset.
