datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
novae
Description
Full novae dataset, including:
All the spatial transcriptomics samples used to train Novae
Protein samples used in the article
Some Visium and Visium HD samples
Synthetic data samples
You can download this dataset from the API, see novae.load_dataset
See here the list of available models trained on this dataset.
[!NOTE]
Note that Novae was trained on the image-based spatial transcriptomics samples. This means that it was not trained on the Visium/VisiumHD samples… See the full description on the dataset page: https://huggingface.co/datasets/prism-oncology/novae.oncology-trial-strategy
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper.
Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our
ICML 2026 workshop paper, accepted at two workshops:
GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline
Policy Training for Decision Agents".
Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.sandboxDatasets sandbox for tutorials or workshops.
E1.S3-Pediatric-Oncology-Genomic-Marker
Pediatric Oncology Synthetic Dataset
ML-Ready Genomic Marker–Driven Leukemia Subtype Classification Dataset
Overview
Property
Value
Patients
1,000
Primary target
Pediatric_Leukemia_Diagnosis (ALL-B / ALL-T / AML / Mixed-lineage / No leukemia)
Total Part A features
41
Total Part B visits
40,267
Avg visits per patient
40.3
Max sequence length
48
Leukemia-positive rate
~90% (tertiary referral center prevalence)… See the full description on the dataset page: https://huggingface.co/datasets/Auric-Grid/E1.S3-Pediatric-Oncology-Genomic-Marker.E1.S4-Synthetic-Thoracic-Oncology-Dataset
Synthetic Thoracic Oncology Dataset
Overview
1,000 synthetic patient records for ML-based treatment response prediction and
recurrence forecasting in early-stage non-small cell lung cancer (NSCLC).
Files
File
Description
thoracic_oncology_crosssectional.csv
Part A: 1,000 patients × 41 features
thoracic_oncology_longitudinal.csv
Part B: ~12,000 rows, one per clinical visit
static_features.csv
Demographics, biomarkers, baseline —… See the full description on the dataset page: https://huggingface.co/datasets/Auric-Grid/E1.S4-Synthetic-Thoracic-Oncology-Dataset.orena-segment-annotations
ORena FOCUS 2026 — SEGMENT supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 SEGMENT
track. This is the openly releasable half; the LapChole-FOCUS half is withheld under that
dataset's usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
1,300
vqa/hernia_mesh.jsonl — hernia videos
140
total
1,440
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-segment-annotations.orena-frame-annotations
ORena FOCUS 2026 — FRAME supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 FRAME track.
This is the openly releasable half; the LapChole-FOCUS half is withheld under that dataset's
usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
7,608
vqa/hernia_mesh.jsonl — hernia videos
1,350
total
8,958
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-frame-annotations.Oncology-Cancer-Dataoncology-financial-reasoning-india
🩺 Medzz-AI: Oncology & Financial Reasoning (India)
Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic
👋 The Problem: Why Current Medical AI Fails
State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient.
Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.MultiModal-Oncology-FHIRThe Anode-MultiModal-Oncology-FHIR dataset is a high-fidelity, synthetic longitudinal record set designed to bridge the gap between genomic sequencing, clinical observations, and therapeutic outcomes. Structured in HL7 FHIR R4, this dataset provides a "ground truth" environment for training Multi-Modal Large Language Models (M-LLMs) in the precision oncology space, specifically targeting Non-Small Cell Lung Cancer (NSCLC) and complex immunotherapy responses.
Dataset Summary
Unlike generic… See the full description on the dataset page: https://huggingface.co/datasets/AnodeAI/MultiModal-Oncology-FHIR.medgemma-4b-hematologic-oncology-blind-spots
MedGemma Blind Spots: Hematologic Oncology & CAR-T Immunotherapy
A 13-probe red-team evaluation showing how Google's MedGemma-4B confidently hallucinates clinical-trial statistics, fabricates non-existent treatment regimens, and misdiagnoses lymphoma in hematologic oncology — a clinical domain absent from its documented training data.
Summary
This dataset documents failures of Google's MedGemma-4B on hematologic oncology prompts — a clinical subspecialty absent… See the full description on the dataset page: https://huggingface.co/datasets/Mateenah/medgemma-4b-hematologic-oncology-blind-spots.Medical_oncology_01
📖 Dataset Summary
This dataset contains high-fidelity, deterministic synthetic patient records for Non-Small Cell Lung Cancer (NSCLC).
Unlike traditional generative AI that "guesses" data based on existing seeds, the Anode Zero-Seed Engine generates these records from first principles using medical logic, clinical guidelines, and genomic constraints. This ensures 100% biological and clinical consistency across all longitudinal fields.
🧬 Technical Specifications &… See the full description on the dataset page: https://huggingface.co/datasets/Sampade07/Medical_oncology_01.Anode_oncology_medicalmultimodal-oncology-atlas
LH2 Data — Multimodal Oncology Dataset
A large-scale, multimodal oncology dataset built around a principle rare in the field: placing non-Caucasian patient populations at the centre, not the periphery.
Dataset Summary
The vast majority of oncology datasets used to train diagnostic, prognostic, and treatment AI models are drawn overwhelmingly from Caucasian, Western populations — a well-documented limitation that undermines model generalisability and equity in real-world… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/multimodal-oncology-atlas.oncology-readability-collapse-risk-v0.3
What this dataset does
This dataset tests whether a model can detect pre-cancer instability risk from loss of signal readability rather than from stress burden alone.
The task is not cancer diagnosis.
The task is to classify whether a synthetic tissue ecology has entered readability collapse risk.
Core Stability Idea
The dataset represents a stability-transition hypothesis.
Cancer vulnerability may begin when tissue regulation loses the ability to correctly read… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-readability-collapse-risk-v0.3.oncologyMedical_oncologyoncology_ready_finetuneafrica-synth-cancer-geriatric-oncology-africa-eritrea
Geriatric Oncology Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-geriatric-oncology-africa-eritrea.Oncology-Cancer-Approved-Drugsoncology-signal-alignment-boundary-v0.4
What this dataset does
This dataset tests whether a model can detect signal-alignment failure in a synthetic tissue ecology.
The task is not cancer diagnosis.
The task is to classify whether readable biological signals can still coordinate repair.
Core Stability Idea
A tissue may still read damage, repair, immune, and metabolic signals but fail because those subsystems no longer align around coherent action.
This dataset moves beyond readability collapse.
It tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-signal-alignment-boundary-v0.4.oncology-missing-signal-detection-v0.6
What this dataset does
This dataset tests whether a model can detect missing-signal instability risk in a synthetic tissue ecology.
The task is not cancer diagnosis.
The task is to classify whether the observed signal set is sufficient to support stable sensing.
Core Stability Idea
A tissue may appear stable because a critical signal is absent.
Absence of signal is not the same as evidence of stability.
This dataset tests whether a model can distinguish true calm… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-missing-signal-detection-v0.6.oncology-signal-latency-boundary-v0.5
What this dataset does
This dataset tests whether a model can detect timing failure in a synthetic tissue ecology.
The task is not cancer diagnosis.
The task is to classify whether a tissue-state scenario can detect and act before the repair opportunity closes.
Core Stability Idea
A tissue may detect the correct signal, interpret it correctly, and coordinate a response, but still fail because action arrives too late.
This dataset tests the timing layer of… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-signal-latency-boundary-v0.5.oncology-precancer-constraint-geometry-v0.2
What this dataset does
This dataset tests whether a model can detect pre-cancer instability from constraint geometry rather than single-variable thresholds.
The task is not cancer diagnosis.
The task is to classify whether a synthetic tissue ecology has crossed into a persistent instability transition.
Core Stability Idea
The dataset represents a stability-transition hypothesis.
Cancer vulnerability may begin when tissue regulation loses self-correcting coherence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-precancer-constraint-geometry-v0.2.Anode_oncology_medical_databiomodels-sbml-odemodels-oncologymedical-text-generation-oncology-froncology-1k
Oncology Medical Dataset — 1,000 Record Free Sample
Enterprise-grade synthetic medical data. Zero PHI. 100% HIPAA-compliant.
Quality Metrics
Metric
Score
Industry Benchmark
Trinity Consensus Score (TAS)
98.0%
85-92% typical
Inter-Annotator Agreement
0.97
0.75-0.85 typical
Macro F1
0.970.80-0.90 typical
PHI Present
None
--
Generation Method
3-LLM Trinity Ensemble
Single model typical
What's Included (Free)
1,000… See the full description on the dataset page: https://huggingface.co/datasets/WitnessDataFactory/oncology-1k.Oncology-Companion-Diagnosticsfaers-oncology-drug-safety
Oncology — Drug Safety Intelligence (FAERS 2020–2025)
Version: 1.0.0 | Records: 928,991 | Source: FDA FAERS
Dataset Summary
Structured adverse event reports for oncology drugs from the FDA's FAERS
database, 2020–2025. Covers serious adverse events only (hospitalization,
life-threatening outcomes, death). Each record includes the suspect drug(s),
reported reactions (MedDRA coded), patient demographics, outcome codes,
reporter country, and seriousness flags.… See the full description on the dataset page: https://huggingface.co/datasets/RubyIntelligence/faers-oncology-drug-safety.
