datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Camelyon17-WILDS
https://wilds.stanford.edu/datasets/#camelyon17
Center 0, 3, 4 - Source (If split=1, Validation (ID))
Center 1 - Validation (OOD)
Center 2 - Target (OOD)
Camelyon16_MIL
CAMELYON16 - Multiple Instance Learning (MIL)
Important. This dataset is part of the torchmil library.
This repository provides an adapted version of the CAMELYON16 dataset tailored for Multiple Instance Learning (MIL). It is designed for use with the CAMELYON16Dataset class from the torchmil library. CAMELYON16 is a widely used benchmark in MIL research, making this adaptation particularly valuable for developing and evaluating MIL models.
About the Original CAMELYON16… See the full description on the dataset page: https://huggingface.co/datasets/torchmil/Camelyon16_MIL.PathoROB-camelyon
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-camelyon.camelyon16-features
Dataset Card for Camelyon16-features
Dataset Summary
The Camelyon16 dataset is a very popular benchmark dataset used in the field of cancer classification.
The dataset we've uploaded here is the result of features extracted from the Camelyon16 dataset using the Phikon model, which is also openly available on Hugging Face.
Dataset Creation
Initial Data Collection and Normalization
The initial collection of the Camelyon16 Whole Slide Images… See the full description on the dataset page: https://huggingface.co/datasets/owkin/camelyon16-features.focusmil-camelyon16
Datasets used in FocusMIL paper
The Camelyon16 and Camelyon16-Standard-MIL-test datasets used in
From Correlation to Causation: Max-Pooling-Based Multi-Instance Learning Leads to More
Robust Whole Slide Image Classification (FocusMIL).
Code: FocusMIL-and-other-max-pooling-methods
These are pre-extracted patch features — you do not need the raw whole-slide images.
Slide classification, patch-level metrics, and FROC localization all run directly off the
files here.
For Camelyon17… See the full description on the dataset page: https://huggingface.co/datasets/Raymvp12/focusmil-camelyon16.Camelyon16_MIL
CAMELYON16 - Multiple Instance Learning (MIL)
Important. This dataset is part of the torchmil library.
This repository provides an adapted version of the CAMELYON16 dataset tailored for Multiple Instance Learning (MIL). It is designed for use with the CAMELYON16Dataset class from the torchmil library. CAMELYON16 is a widely used benchmark in MIL research, making this adaptation particularly valuable for developing and evaluating MIL models.
About the Original… See the full description on the dataset page: https://huggingface.co/datasets/Akiriiqiu/Camelyon16_MIL.
