LH2-data-labs/multimodal-oncology-atlas
LH2 Data — Multimodal Oncology Dataset A large-scale, multimodal oncology dataset built around a principle rare in the field: placing non-Caucasian patient populations at the centre, not the periphery. Dataset Summary The vast majority of oncology datasets used to train diagnostic, prognostic, and treatment AI models are drawn overwhelmingly from Caucasian, Western populations — a well-documented limitation that undermines model generalisability and equity in… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/multimodal-oncology-atlas.
LH2 Data — Multimodal Oncology Dataset
A large-scale, multimodal oncology dataset built around a principle rare in the field: placing non-Caucasian patient populations at the centre, not the periphery.
Dataset Summary
The vast majority of oncology datasets used to train diagnostic, prognostic, and treatment AI models are drawn overwhelmingly from Caucasian, Western populations — a well-documented limitation that undermines model generalisability and equity in real-world deployment. This dataset addresses this directly, offering a deeply multimodal oncology resource — genomic, clinical, imaging, and biospecimen — anchored across the Global South.
This sample contains structured clinical records for 20 de-identified cancer patients demonstrating the full schema. The complete dataset includes ~25,000 enrolled patients (target: 50,000 within 12 months) across 15 countries spanning the Indian Subcontinent, Southeast Asia, Latin America, and the Middle East.
Cohort at a Glance
Data Modalities
Clinical Record Schema (28 fields)
Key Differentiators
- Non-Caucasian patient populations placed at the centre of dataset design — a principle rare in oncology AI
- Multimodal from day one: genomic + imaging + clinical + biospecimen, all linked per patient
- 13 clinical data categories per patient including staging, surgery, chemotherapy, targeted therapy, molecular pathology (NGS + IHC), PET-CT, serum biomarkers, haematology, and biochemistry
- 15 countries across the Indian Subcontinent, Southeast Asia, Latin America, and the Middle East — populations systematically absent from TCGA, ICGC, and most major oncology training datasets
- Physical biobank enables generation of future modalities (spatial transcriptomics, single-cell RNA-seq)
- Scalable cohort: 25,000 today → 50,000 target in 12 months
IP & Compliance
- Full commercial and AI/ML training rights secured under exclusive licensing agreement
- Ethics and informed consent frameworks aligned with source jurisdictions
- De-identified patient data — no personally identifiable information (PII)
- Compliant with applicable data protection and cross-border transfer regulations (DPDP Act 2023, HIPAA, GDPR as applicable)
Citation
@dataset{lh2_oncology_2025,
title={LH2 Data — Multimodal Oncology Dataset},
author={LH2 Data Labs},
year={2025},
publisher={Hugging Face},
note={Multimodal oncology dataset: 25K+ patients across the Global South}
}