datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tcga-wsi-uni2h-features
TCGA WSI UNI2H Features
Dataset Summary
This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research.
Data is organized by project (for example TCGA-HNSC) and currently exposes:
features/ containing H5 feature files with tile-level embeddings
vis/ containing overlay images for quality inspection and pipeline verification
[!IMPORTANT]
Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.Paladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps
Ready-to-use patch-level and slide-level spatial omics maps inferred by
Paladin from TCGA and CPTAC
whole-slide images. The released .Paladin.h5 files can be analyzed directly
without rerunning WSI inference.
The collection is populated in stages. Check Files and versions for the
cohorts currently available.
Spatial multi-omics example
The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV,
DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.tcga-brca-titan-idc-ilc
tcga-brca-titan-idc-ilc
1. Tổng quan
[CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset]
Tổng số bản ghi (cộng tất cả manifest phát hiện được): 4228
Số manifest phát hiện được trong bộ nhớ: 3 (df, brca_df, full_df)
Repo HuggingFace: okbro1234/tcga-brca-titan-idc-ilc
2. Cấu trúc lưu trữ tại đích
/ # suy từ hàm `HfApi`
file.txt # suy từ hàm `HfApi`
lfs.bin # suy từ hàm `HfApi`
shard_{i}_of_5.bin # suy từ hàm `HfApi`
remote/file/path.h5 #… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/tcga-brca-titan-idc-ilc.tcgaTCGA-PANCAN-HiSeq-2770x20530gene expression cancer RNA-Seq - Check the original submission: - https://www.synapse.org/Synapse:syn2812925 - is maintained by the cancer genome atlas pan-cancer analysis project. - TCGA-PANCAN-HiSeq-2770x20530
Files combined:
unc.edu_BRCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 957) BRCA
unc.edu_KIRC_IlluminaHiSeq_RNASeqV2.geneExp (20530, 552) KIRC
unc.edu_LUAD_IlluminaHiSeq_RNASeqV2.geneExp (20530, 413) LUAD
unc.edu_THCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 471) THCA… See the full description on the dataset page: https://huggingface.co/datasets/Fllamber/TCGA-PANCAN-HiSeq-2770x20530.TCGA-12K-parquet
TCGA-12K Parquet
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled across… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet.TCGA-UniformTumor-8K
Dataset Card for TCGA-UniformTumor-8K
What is TCGA-UniformTumor-8K?
TCGA-UniformTumor-8K dataset is a region-level pan-cancer subtyping resource comprising 25,495 ROIs of 8,192 × 8,192 pixels. These regions were extracted from 9,662 H&E-stained FFPE diagnostic histopathology WSIs sourced from TCGA. The tumor regions were manually annotated by two expert pathologists, with slide exclusion due to poor staining, poor focus, lacking cancerous regions and incorrect… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/TCGA-UniformTumor-8K.TCGA_OncoTree_pt2
TCGA_OncoTree_pt2
1. Tổng quan
[CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset]
Tổng số bản ghi (cộng tất cả manifest phát hiện được): 23984
Số manifest phát hiện được trong bộ nhớ: 3 (df, labels_df, progress)
Repo HuggingFace chính: ento3686/TCGA_OncoTree_pt2
⚠️ Dataset được lưu trên 2 repo/tài khoản HuggingFace khác nhau:
ento3686/TCGA_OncoTree_pt2 (biến: REPO_ID_2, UPLOAD_REPO_ID, CENTRAL_PROGRESS_REPO_ID, _repo_id_var)
tuna2004/TCGA_OncoTree (biến:… See the full description on the dataset page: https://huggingface.co/datasets/ento3686/TCGA_OncoTree_pt2.TCGA-mini
TCGA-mini
TCGA-mini is a curated whole-slide imaging (WSI) dataset composed of TCGA .svs pathology slides.
Number of slides: 1319 (.svs)
Location in repo: slides/
Storage footprint on the Hub: about 1.42 TiB (1450.94 GiB, Git LFS-backed)
Dataset structure
TCGA-mini/
├── README.md
└── slides/
├── TCGA-xxxx-xxxx-xxZ-00-DX1.svs
└── ...
Data format
File type: Aperio whole-slide image files (.svs)
Naming: Original TCGA-style slide identifiers are… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/TCGA-mini.tcga-ut
Histology images from uniform tumor regions in TCGA Whole Slide Images (TCGA-UT-Internal, TCGA-UT-External)
This repository provides a benchmarking framework for the TCGA histology image dataset originally published on Zenodo. It includes predefined train/validation/test splits and example code for foundation model evaluation.
Task
Classification of 31 different cancer types from tumor histopathological images.
Original Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/dakomura/tcga-ut.TCGA_virtual_spatial_transcriptomics_atlas
TCGA virtual spatial transcriptomics atlas
This repository contains predicted spatial transcriptomics for TCGA H&E slides,
both fresh-frozen (FF) and formalin-fixed paraffin-embedded (FFPE), produced
with DeepSpot-M.
Authors: Kalin Nonchev, Sebastian Dawo, Karina Silina, Viktor Hendrik
Koelzer, and Gunnar Rätsch.
Model: ratschlab/DeepSpotM · Code: github.com/ratschlab/DeepSpotM · Paper: medRxiv 2026.06.19.26356060.
News
[09.2026] Introducing Aurora - a no-code… See the full description on the dataset page: https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas.TCGA-OV-AS
The Cancer Genome Atlas Ovarian Cancer for Ascites Segmentation (TCGA-OV-AS)
This dataset was curated as part of the research 'Deep Learning Segmentation of Ascites on Abdominal CT Scans for Automatic Volume Quantification' (Paper, arXiv).
To replicate TCGA-OV-AS, please download TCGA-OV from TCIA using the Descriptive Directory Name download option.
Converting Images
Convert the DICOMs to NIFTI format using dcm2niix and GNU parallel.
Create the directory structure… See the full description on the dataset page: https://huggingface.co/datasets/farrell236/TCGA-OV-AS.TCGA
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset
The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, slide images, molecular data, and radiology images for cancer patients.
This dataset aims to facilitate research in multimodal machine learning for oncology by providing embeddings generated using state-of-the-art models including GatorTron, MedGemma, Qwen, Llama, UNI, SeNMo, REMEDIS, and… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/TCGA.tcga-tabular-open
TCGA Tabular (Open Access)
Open-access TCGA data from the NCI Genomic Data Commons (GDC). Covers
all 33 TCGA projects.
This view presents one HuggingFace subset per (project, table). See the [tcga-patients-open][patients] companion for a per-row view of the same underlying data.
Generated: 2026-08-14 01:41:46 UTC
Schema: derived from the [GDC Data Dictionary][gdc-dict].
GDC data release: Data Release 45.0 - December 04, 2025
Data model
Where the data… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-tabular-open.PathoROB-tcga
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-tcga.TCGA-12K-parquet-shuffled
TCGA-12K Parquet (Shuffled)
Attribution
This dataset contains 224 x 224 JPEG patches from whole-slide images originally downloaded from The Cancer Genome Atlas (TCGA) that are available in the NCI Genomic Data Commons (GDC) Open Access tier. We mirror and repackage a commonly used ~12k WSI subset in parquet format for ease of training. We exclude patches that did not pass HSV thresholding, following the procedure in Kaiko.AI's Midnight paper. Patches were randomly sampled… See the full description on the dataset page: https://huggingface.co/datasets/medarc/TCGA-12K-parquet-shuffled.TCGA_virtual_spatial_transcriptomics
Dataset card for TCGA digital spatial transcriptomics data
This repository contains results from the paper "DeepSpot: Leveraging Spatial Context for Enhanced Spatial Transcriptomics Prediction from H&E Images".
Authors: Kalin Nonchev, Sebastian Dawo, Karina Selina, Holger Moch, Sonali Andani, Tumor Profiler Consortium, Viktor Hendrik Koelzer, and Gunnar Rätsch
The preprint is available here.
What is TCGA digital spatial transcriptomics?
We trained a model using… See the full description on the dataset page: https://huggingface.co/datasets/nonchev/TCGA_virtual_spatial_transcriptomics.tcgatcga-brca-tabular-open
TCGA-BRCA — Tabular (Open Access)
Open-access TCGA-BRCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:47:01 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-brca-tabular-open.Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset
Note: Our 245 TCGA cases are ones we identified as having potential for improvement.
We plan to upload them in two phases: the first batch of 138 cases, and the second batch of 107 cases in the quality review pipeline, we plan to upload them around early of January, 2025.
Dataset: A Second Opinion on TCGA PRAD Prostate Dataset Labels with ROI-Level Annotations
Overview
This dataset provides enhanced Gleason grading annotations for the TCGA PRAD prostate cancer… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset.tcga-tgct-tabular-open
TCGA-TGCT — Tabular (Open Access)
Open-access TCGA-TGCT data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:21:25 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-tgct-tabular-open.tcga-acc-tabular-open
TCGA-ACC — Tabular (Open Access)
Open-access TCGA-ACC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:44:59 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-acc-tabular-open.tcga-read-tabular-open
TCGA-READ — Tabular (Open Access)
Open-access TCGA-READ data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:16:05 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-read-tabular-open.tcga-gene-expression-quantification-open
TCGA Gene Expression Quantification — Open Access
Cohort-wide gene expression matrices for the open-access TCGA RNA-Seq data distributed by the NCI Genomic Data Commons. GDC serves these measurements one file per aliquot; here they are arranged as one row per sample, with each quantification carried as its own matrix over identical axes.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:16:09 UTC
Shape: 11,505 samples x 60,660 genes
Projects: 33
Other… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-gene-expression-quantification-open.tcga-kich-tabular-open
TCGA-KICH — Tabular (Open Access)
Open-access TCGA-KICH data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:58:23 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-kich-tabular-open.TCGAsurvival-fusion-tcga600-v3.3
Survival Fusion TCGA-600 v3.3
This dataset contains reproducibly derived multimodal features for 600 TCGA
patients: 200 each from TCGA-BLCA, TCGA-HNSC, and TCGA-STAD. It supports the
Survival Fusion benchmark's development and frozen Stage-1 qualification
protocol.
The repository contains derived feature arrays and public TCGA identifiers. It
does not contain raw whole-slide images or raw sequencing reads.
Contents
canonical-v3.3-v1/: canonical per-patient… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/survival-fusion-tcga600-v3.3.tcga-dlbc-tabular-open
TCGA-DLBC — Tabular (Open Access)
Open-access TCGA-DLBC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 03:53:50 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-dlbc-tabular-open.TCGA-OncoTree_EmbeddingBraTS-TCGA
BraTS-TCGA (BraTS-TCGA-GBM + BraTS-TCGA-LGG)
Expert segmentation labels for the pre-operative TCGA glioma MRI cohorts
(Bakas et al. 2017), combining the two TCIA analysis-result collections
BraTS-TCGA-GBM (102 glioblastoma patients) and BraTS-TCGA-LGG
(65 lower-grade glioma patients) = 167 cases.
What this is (faithful-naming note): the publicly released training
half of the pre-operative subset of TCGA-GBM / TCGA-LGG, already
co-registered to a T1 template, resampled to 1 mm³… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/BraTS-TCGA.
