kozo2 commited on
Commit
d538c61
·
verified ·
1 Parent(s): 9c7cdf3

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +115 -0
README.md ADDED
@@ -0,0 +1,115 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-4.0
3
+ library_name: pytorch
4
+ pipeline_tag: feature-extraction
5
+ tags:
6
+ - node2vec
7
+ - graph-embedding
8
+ - metabolomics
9
+ - pytorch-geometric
10
+ - metabolights
11
+ metrics:
12
+ - roc_auc
13
+ ---
14
+
15
+ # Node2Vec embeddings for the edge_ML metabolomics graph
16
+
17
+ 128-dimensional Node2Vec embeddings for the 18,494 nodes of an undirected metabolite
18
+ co-response graph with 2,709,209 edges. Held-out link prediction reaches **AUC 0.988**,
19
+ against 0.920 for a degree-only baseline.
20
+
21
+ The graph, node properties and full pipeline are in the companion dataset repository.
22
+
23
+ ## Files
24
+
25
+ | File | Contents | Size |
26
+ |---|---|---|
27
+ | `edge_ML_expected_ge5_n2v.pt` | `{embedding: [18494, 128] float32, node_id: [18494], args: {...}}` | 9.7 MB |
28
+ | `node2vec_model.py` | Model definition, training loop, embedding export | — |
29
+
30
+ ## Using it
31
+
32
+ ```python
33
+ import torch
34
+
35
+ ck = torch.load("edge_ML_expected_ge5_n2v.pt", weights_only=False)
36
+ z = ck["embedding"] # [18494, 128] float32
37
+ index = {nid: i for i, nid in enumerate(ck["node_id"])}
38
+ v = z[index["MTBLS1405_0002_00003332"]] # one node's vector
39
+ ```
40
+
41
+ `node_id[i]` is the original string ID for row `i`; the order is lexicographic over the
42
+ union of the graph's two endpoint columns, matching the dataset's graph object. Scores
43
+ were computed with **cosine** similarity, which is also the metric to use downstream.
44
+
45
+ ## Training
46
+
47
+ | Parameter | Value |
48
+ |---|---|
49
+ | `embedding_dim` | 128 |
50
+ | `walk_length` | 20 |
51
+ | `context_size` | 10 |
52
+ | `walks_per_node` | 10 |
53
+ | `num_negative_samples` | 1 |
54
+ | `p`, `q` | 1.0, 1.0 (unbiased walks) |
55
+ | Batch size | 128 seed nodes, 145 batches per epoch |
56
+ | Optimiser | `SparseAdam`, lr 0.01 |
57
+ | Epochs | 20 |
58
+ | Parameters | 2,367,232 (18,494 × 128) |
59
+
60
+ Loss fell from 9.92 at initialisation to 0.880, flat from about epoch 14, at roughly
61
+ 0.9 s/epoch on one H100. `sparse=True` on the model is what allows `SparseAdam`;
62
+ changing either requires changing the other.
63
+
64
+ ```bash
65
+ uv run python node2vec_model.py --epochs 20
66
+ ```
67
+
68
+ `Node2Vec` requires `pyg-lib >= 0.6.0` for its random-walk kernel, which is not on PyPI;
69
+ the dataset repository's `pyproject.toml` pins `pyg-lib` 0.9.0+pt214cu130 from
70
+ `data.pyg.org`.
71
+
72
+ ## Evaluation
73
+
74
+ 200,000 sampled positive edges against 200,000 non-edges verified absent from the full
75
+ edge set, scored by cosine similarity, AUC by the Mann-Whitney rank identity.
76
+
77
+ | Model | Scored edges | Score | AUC |
78
+ |---|---|---|---|
79
+ | 90/10 retrain | Held-out 10%, never seen | cosine | **0.9880** |
80
+ | 90/10 retrain | Its own training edges | cosine | 0.9892 |
81
+ | 90/10 retrain | Held-out 10%, never seen | degree product `d_u × d_v` | 0.9201 |
82
+ | Full graph (this release) | Its own training edges | cosine | 0.9892 |
83
+ | Full graph (this release) | Its own training edges | dot product | 0.9868 |
84
+
85
+ The held-out row is the one that matters: a second model was trained from scratch on 90%
86
+ of the edges and scored on the 10% it never saw. Held-out 0.9880 against in-sample 0.9892
87
+ is a gap of 0.001, so the model learns graph structure rather than memorising pairs. The
88
+ degree baseline matters because the graph is dense (median degree 90) — a high AUC that
89
+ merely reproduced the degree distribution would carry little information.
90
+
91
+ Other checks on the released embeddings:
92
+
93
+ - **Neighbourhood recovery** — of each node's 10 nearest embeddings, 49.9% are true graph
94
+ neighbours against 1.6% expected by chance (31.7×); at top-50, 41.3% (26.2×).
95
+ - **Embedding health** — all finite; L2 norms 0.94 / 1.92 / 9.49 (min / median / max);
96
+ per-dimension standard deviation 0.15–0.28, so no dead dimensions; mean cosine over
97
+ 200,000 random pairs is 0.0038, ruling out collapse.
98
+ - **Species purity** — 90.7% of all nodes have ten nearest embeddings sharing their
99
+ species, rising above 98% for the three largest species and falling to 68–79% for
100
+ species with a few hundred nodes.
101
+
102
+ ## Limitations
103
+
104
+ - **Topology only.** The walks are unweighted, so neither the graph's `edge_attr`
105
+ (`OddsRatio_log2`, `ChiTestsPValue`) nor its node features `x` influence these
106
+ embeddings. Letting association strength steer the walks needs a weighted sampler or a
107
+ pre-thresholded edge set; using the node features needs a message-passing model.
108
+ - **Transductive.** Node2Vec learns one vector per node in a fixed graph. There is no
109
+ way to embed a node that was not present at training time.
110
+ - **Species and study are entangled.** Edges form mostly within a study and a study is
111
+ normally one species, so the clean species separation partly reflects how the graph was
112
+ assembled, not an independent biological signal.
113
+ - In the 90/10 evaluation split, 97 low-degree nodes were left isolated in the training
114
+ graph and their vectors stay near initialisation. That affects only the held-out
115
+ experiment; the released full-graph model has no isolated nodes.