CoolFace
Datasetpublic

geometric-intelligence/ogbench

OgBench: Benchmarking Graph Neural Networks on Omics Data OgBench is the first benchmark suite for graph-level prediction in the n ≪ p regime characteristic of omics data, where the number of patient samples n is much smaller than the number of nodes (genes or proteins) p per graph. Datasets This repository contains four preprocessed omics graph classification datasets: Dataset Modality n p Task HERITAGE Proteomics 654 4,977 Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.

sourceHugging Facecc-by-4.0updated 10d agoView on Hugging Face
0likes1.7kdownloads
Dataset Card

OgBench: Benchmarking Graph Neural Networks on Omics Data

OgBench is the first benchmark suite for graph-level prediction in the n ≪ p regime characteristic of omics data, where the number of patient samples n is much smaller than the number of nodes (genes or proteins) p per graph.

Datasets

This repository contains four preprocessed omics graph classification datasets:

DatasetModalitynpTask
HERITAGEProteomics6544,977Exercise responder (binary)
Parkinson'sTranscriptomics53521,755Cognitive status (binary)
AddNeuroMedTranscriptomics71117,198Clinical diagnosis (3-class)
BRCAEpigenomics64019,049Cancer subtype (4-class)

Source Data

  • HERITAGE: Robbins et al. (2021), Nature Metabolism. Available via MoTrPAC Data Hub (motrpac-data.org) under CC-BY 4.0.
  • Parkinson's: Shamir et al. (2017), Neurology. Available via NCBI GEO (GSE99039) under GEO public data access policy.
  • AddNeuroMed: Lovestone et al. (2009). Available via NCBI GEO (GSE63063) under GEO public data access policy.
  • BRCA: Yang et al. (2025), MLOmics, Scientific Data. Available on Figshare/Hugging Face under CC-BY 4.0.

Preprocessing

All datasets are preprocessed with a consistent pipeline including probe-to-gene aggregation, normalization, and covariate adjustment. Full preprocessing details are provided in Appendix B of the accompanying paper. Graphs are split 70/15/15 (train/val/test) with a fixed random seed.