CoolFace
Modelpublic

kozo2/edge-ML-node2vec

sourceHugging Facecc-by-4.0updated 5d agoView on Hugging Face
0likes
Model Card

Node2Vec embeddings for the edge_ML metabolomics graph

128-dimensional Node2Vec embeddings for the 18,494 nodes of an undirected metabolite co-response graph with 2,709,209 edges. Held-out link prediction reaches AUC 0.988, against 0.920 for a degree-only baseline.

The graph, node properties and full pipeline are in the companion dataset repository.

Files

FileContentsSize
edge_ML_expected_ge5_n2v.pt{embedding: [18494, 128] float32, node_id: [18494], args: {...}}9.7 MB
node2vec_model.pyModel definition, training loop, embedding export

Using it

python
import torch

ck = torch.load("edge_ML_expected_ge5_n2v.pt", weights_only=False)
z = ck["embedding"]                      # [18494, 128] float32
index = {nid: i for i, nid in enumerate(ck["node_id"])}
v = z[index["MTBLS1405_0002_00003332"]]  # one node's vector

node_id[i] is the original string ID for row i; the order is lexicographic over the union of the graph's two endpoint columns, matching the dataset's graph object. Scores were computed with cosine similarity, which is also the metric to use downstream.

Training

ParameterValue
embedding_dim128
walk_length20
context_size10
walks_per_node10
num_negative_samples1
p, q1.0, 1.0 (unbiased walks)
Batch size128 seed nodes, 145 batches per epoch
OptimiserSparseAdam, lr 0.01
Epochs20
Parameters2,367,232 (18,494 × 128)

Loss fell from 9.92 at initialisation to 0.880, flat from about epoch 14, at roughly 0.9 s/epoch on one H100. sparse=True on the model is what allows SparseAdam; changing either requires changing the other.

bash
uv run python node2vec_model.py --epochs 20

Node2Vec requires pyg-lib >= 0.6.0 for its random-walk kernel, which is not on PyPI; the dataset repository's pyproject.toml pins pyg-lib 0.9.0+pt214cu130 from data.pyg.org.

Evaluation

200,000 sampled positive edges against 200,000 non-edges verified absent from the full edge set, scored by cosine similarity, AUC by the Mann-Whitney rank identity.

ModelScored edgesScoreAUC
90/10 retrainHeld-out 10%, never seencosine0.9880
90/10 retrainIts own training edgescosine0.9892
90/10 retrainHeld-out 10%, never seendegree product d_u × d_v0.9201
Full graph (this release)Its own training edgescosine0.9892
Full graph (this release)Its own training edgesdot product0.9868

The held-out row is the one that matters: a second model was trained from scratch on 90% of the edges and scored on the 10% it never saw. Held-out 0.9880 against in-sample 0.9892 is a gap of 0.001, so the model learns graph structure rather than memorising pairs. The degree baseline matters because the graph is dense (median degree 90) — a high AUC that merely reproduced the degree distribution would carry little information.

Other checks on the released embeddings:

  • Neighbourhood recovery — of each node's 10 nearest embeddings, 49.9% are true graph neighbours against 1.6% expected by chance (31.7×); at top-50, 41.3% (26.2×).
  • Embedding health — all finite; L2 norms 0.94 / 1.92 / 9.49 (min / median / max); per-dimension standard deviation 0.15–0.28, so no dead dimensions; mean cosine over 200,000 random pairs is 0.0038, ruling out collapse.
  • Species purity — 90.7% of all nodes have ten nearest embeddings sharing their species, rising above 98% for the three largest species and falling to 68–79% for species with a few hundred nodes.

Limitations

  • Topology only. The walks are unweighted, so neither the graph's edge_attr (OddsRatio_log2, ChiTestsPValue) nor its node features x influence these embeddings. Letting association strength steer the walks needs a weighted sampler or a pre-thresholded edge set; using the node features needs a message-passing model.
  • Transductive. Node2Vec learns one vector per node in a fixed graph. There is no way to embed a node that was not present at training time.
  • Species and study are entangled. Edges form mostly within a study and a study is normally one species, so the clean species separation partly reflects how the graph was assembled, not an independent biological signal.
  • In the 90/10 evaluation split, 97 low-degree nodes were left isolated in the training graph and their vectors stay near initialisation. That affects only the held-out experiment; the released full-graph model has no isolated nodes.