datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.FairVLMed
Dataset Card: Harvard-FairVLMed
Dataset Summary
Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language.
This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.FairFedMed
Dataset Card: FairFedMed
Dataset Summary
FairFedMed is the first federated learning (FL) benchmark dataset for medical imaging with demographic annotations, designed to study group fairness across institutions in a federated setting. It comprises two subsets spanning ophthalmology and chest radiology, enabling research on fairness-aware federated learning under realistic cross-institutional data heterogeneity.
This dataset was introduced in the IEEE Transactions on… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairFedMed.FairDomain
Dataset Card: Harvard-FairDomain
Dataset Summary
Harvard-FairDomain is a large-scale ophthalmology dataset designed for studying fairness under domain shift in medical image analysis. It supports both image segmentation and classification tasks, with 10,000 samples per task drawn from 10,000 unique patients. The dataset introduces an additional imaging modality — en-face fundus images — alongside the original scanning laser ophthalmoscopy (SLO) fundus images, enabling… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairDomain.FairGenMed
Dataset Card: FairGenMed
Dataset Summary
FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection.
This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.
