datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAMELYON16
CAMELYON16
1. Tổng quan
CAMELYON16 là dataset ảnh mô bệnh học toàn tiêu bản (WSI) hạch bạch huyết canh gác (sentinel lymph node) của bệnh nhân ung thư vú, thu thập tại 2 trung tâm ở Hà Lan (Radboud University Medical Center và University Medical Center Utrecht). Bài toán chính là phân loại nhị phân cấp-slide: phát hiện có/không có di căn ung thư trong hạch (tumor/normal). Dataset gốc gồm 400 WSI (270 training, 130 testing).
Nguồn dữ liệu: AWS Open Data… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON16.Camelyon17-WILDS
https://wilds.stanford.edu/datasets/#camelyon17
Center 0, 3, 4 - Source (If split=1, Validation (ID))
Center 1 - Validation (OOD)
Center 2 - Target (OOD)
PathoROB-camelyon
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-camelyon.CAMELYON17
CAMELYON17
1. Tổng quan
CAMELYON17 là dataset mở rộng của CAMELYON16, gồm ảnh WSI hạch bạch huyết canh gác từ 5 trung tâm y tế khác nhau (multi-center), với 1000 WSI (5 slide/bệnh nhân x 200 bệnh nhân). Bài toán chính là phân loại di căn theo 4 mức tại cấp lymph-node (negative/isolated tumor cells/micro-metastases/macro-metastases) và tổng hợp thành pN-stage tại cấp bệnh nhân.
Nguồn dữ liệu: AWS Open Data, s3://camelyon-dataset/CAMELYON17/ (region us-west-2, truy… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON17.camelyon17
Dataset Card for "camelyon17"
More Information needed
camelyon16_clampatch_camelyonCAMEL-Benchcamelyon16-uni
CAMELYON16 patch embeddings made with UNI
This repository contains patch embeddings for CAMELYON16 made with the UNI foundation model.
Patches are 128x128 micrometers. Tissue segmentation and patching was done with a modified version of the CLAM toolkit.
The toolkit was modified to extract constant physical size patches.
The patches directory contains HDF5 files with patch coordinates. The attribute patch_size on the /coords dataset
contains the patch size in pixels. This is… See the full description on the dataset page: https://huggingface.co/datasets/kaczmarj/camelyon16-uni.vtab_patch_camelyon
VTAB PatchCamelyon
This dataset has been used for the paper Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models (NeurIPS 2025).
It reproduces the settings (splits, labels) used for the Visual Task Adaptation Benchmark (VTAB).
VTAB Paper: A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
VTAB Repository: google-research/task_adaptation
Details of the original dataset:
Original… See the full description on the dataset page: https://huggingface.co/datasets/bramtoula/vtab_patch_camelyon.mock_websiteswds_wilds-camelyon17_testgridgameDataset containing a number of 2d game-style textured grids, with question-answer pairs relating to the contents of the image. Requires the model to precisely (on grid) locate various objects.
Dataset primarily produced by an anonymous collaborator.
CAMELYON16imnet1k_Arabian_camel_dromedary_Camelus_dromedariuscountlinesSimple set of 100 programmatically generated images with scattered lines. Total # of lines as label for each image.
vpct-parquet
Visual Physics Comprehension Test (VPCT)
Can you predict which of the three buckets the ball will fall into?
The VPCT dataset contains 100 problems, generated with the assistance of a simulator tool. Each problem image contains 3 buckets, a floating ball, and a series of ramps. The simulator tool was used to determine ground truth for which bucket the ball will fall into. The resulting simulation data is paired with each image, as well as the data necessary to re-run the… See the full description on the dataset page: https://huggingface.co/datasets/camelCase12/vpct-parquet.ShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/Camellia054/ShareGPT4Video.countintersectionsLines randomly placed on image, labeled by number of intersection points in the image.
camelbert-ca-caner-e1camelyon16-conch
CONCH embeddings for CAMELYON16 dataset
CONCH is a vision-language foundation model created by the Mahmood Lab. We can use CONCH to embed patches from whole slide images.
In this dataset, we have embedded the CAMELYON16 whole slide images (n=399 slides). See this link for the CAMELYON16 dataset: https://camelyon17.grand-challenge.org/Data/ .
The vision-only directory contains embeddings with proj_contrast=False, and the vision-language directory contains embeddings with… See the full description on the dataset page: https://huggingface.co/datasets/kaczmarj/camelyon16-conch.imnet1k_ostrich_Struthio_camelusgridsizeVarious sized grids labeled by their width and height.
medical_anomalies_qwen_camelyoncolorfillDataset of images blotted with various colors and labeled by the % that each color fills the image by pixels.
controlnet_somethingv2_planbcontrolnet_somethingv2_finaltrycamelcontrolnet_somethingv2camel_mcqa
