oncology
OpenMed-NER-OncologyDetect-MultiMed-568MOpenMed-NER-OncologyDetect-BigMed-278MOpenMed-NER-OncologyDetect-SuperClinical-434MOpenMed-NER-OncologyDetect-SuperMedical-355MOpenMed-NER-OncologyDetect-ModernMed-395MOpenMed-NER-OncologyDetect-PubMed-335MOpenMed-NER-OncologyDetect-TinyMed-65MOpenMed-NER-OncologyDetect-BigMed-560M
Datasets
All datasets matching “oncology”novae
Description
Full novae dataset, including:
All the spatial transcriptomics samples used to train Novae
Protein samples used in the article
Some Visium and Visium HD samples
Synthetic data samples
You can download this dataset from the API, see novae.load_dataset
See here the list of available models trained on this dataset.
[!NOTE]
Note that Novae was trained on the image-based spatial transcriptomics samples. This means that it was not trained on the Visium/VisiumHD samples… See the full description on the dataset page: https://huggingface.co/datasets/prism-oncology/novae.oncology-trial-strategy
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper.
Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our
ICML 2026 workshop paper, accepted at two workshops:
GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline
Policy Training for Decision Agents".
Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.sandboxDatasets sandbox for tutorials or workshops.
E1.S3-Pediatric-Oncology-Genomic-Marker
Pediatric Oncology Synthetic Dataset
ML-Ready Genomic Marker–Driven Leukemia Subtype Classification Dataset
Overview
Property
Value
Patients
1,000
Primary target
Pediatric_Leukemia_Diagnosis (ALL-B / ALL-T / AML / Mixed-lineage / No leukemia)
Total Part A features
41
Total Part B visits
40,267
Avg visits per patient
40.3
Max sequence length
48
Leukemia-positive rate
~90% (tertiary referral center prevalence)… See the full description on the dataset page: https://huggingface.co/datasets/Auric-Grid/E1.S3-Pediatric-Oncology-Genomic-Marker.E1.S4-Synthetic-Thoracic-Oncology-Dataset
Synthetic Thoracic Oncology Dataset
Overview
1,000 synthetic patient records for ML-based treatment response prediction and
recurrence forecasting in early-stage non-small cell lung cancer (NSCLC).
Files
File
Description
thoracic_oncology_crosssectional.csv
Part A: 1,000 patients × 41 features
thoracic_oncology_longitudinal.csv
Part B: ~12,000 rows, one per clinical visit
static_features.csv
Demographics, biomarkers, baseline —… See the full description on the dataset page: https://huggingface.co/datasets/Auric-Grid/E1.S4-Synthetic-Thoracic-Oncology-Dataset.orena-segment-annotations
ORena FOCUS 2026 — SEGMENT supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 SEGMENT
track. This is the openly releasable half; the LapChole-FOCUS half is withheld under that
dataset's usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
1,300
vqa/hernia_mesh.jsonl — hernia videos
140
total
1,440
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-segment-annotations.
