TaxEmbed — one hyperbolic embedding for all named cellular life
TaxEmbed places every taxon identifier of the NCBI Taxonomy of cellular life, 1,102,163
identifiers under "cellular organisms" (new_taxdump downloaded 2026-06-09), into a single
100-dimensional Poincaré ball. It is a lookup: 100 numbers per taxon, fixed-size and
differentiable, released so that a model that needs organismal context can take it as input instead
of building its own from the taxdump.
Paper: Koludarov I, Rost B. TaxEmbed: one hyperbolic embedding for all named cellular life (2026, preprint to follow). Code: https://github.com/jcoludar/taxembed (Apache-2.0).
What the geometry encodes
- Radius is taxonomic depth, by construction. Every norm is initialized at a depth-derived target and a radial nudge holds it there during training. On the released tensor the Pearson correlation of norm with depth is +0.957, equal to its value before the first gradient step (+0.957). The radius therefore records that the schedule was applied; it is not a learned property.
- Direction is learned lineage. Taxa at the same depth are ordered by lineage. The radius-free criterion S_angle (retrieve the k = 10 nearest same-depth taxa by cosine of the directions and score how well that ranking orders them by the depth of their most recent common ancestor with the query; random directions score 0, a perfect ordering 1) reads 0.965 on the released tensor (clade-clustered SE 0.007; 10,000 seeded queries) against an initialization null of −0.004, and 0.941 / 0.962 / 0.978 in the shallow / mid / deep depth bands.
- Embedded neighbours are tree neighbours. Half of a taxon's ten nearest embedded neighbours are among its ten nearest by tree path: precision@10 0.546 (95 % CI 0.535–0.555; 3,000 queries against 2,000 candidates; chance 0.005).
- Embedded distance is a taxonomic distance, not a clock. Against TimeTree divergence times (6,000 pre-registered pairs), the Spearman correlation of angular distance is +0.89 in Vertebrata (NCBI path length: +0.79) and +0.30 in Insecta (path length: +0.44). Most of the vertebrate advantage comes from ancestor depth, a property of the tree. A study that needs divergence times should use them directly.
Files
| File | Size | Contents |
|---|---|---|
cellular_embedding.safetensors |
441 MB | The embedding matrix. Tensor key embedding, shape (1102163, 100), float32. Dimension, taxon count, taxdump date and geometry are in the safetensors header (__metadata__). |
taxid_to_index.tsv |
16 MB | Two columns taxid → idx. Row i of the matrix is the taxon whose idx == i. Byte-identical to the training closure's mapping (MD5 a8ef06b048e03613230a0a14908f515e). |
edges_parent_child.tsv |
15 MB | 1,102,162 lines parent_idx child_idx on the same row indices; the root (cellular organisms, taxid 131567) is row 85051. Depths and tree-path distances follow from it without the taxdump. |
training_closure.npz |
8.6 MB | The training signal: all 21,399,053 ancestor–descendant pairs. Keys ancestor_idx, descendant_idx, depth_diff, ancestor_depth, descendant_depth, ancestor_taxid, descendant_taxid. |
closure_manifest.json |
— | Node, edge and pair counts of the closure (1,102,163 / 1,102,162 / 21,399,053; max depth 40). |
LICENSE |
— | Apache-2.0. |
Quick start
import numpy as np, pandas as pd
from safetensors.numpy import load_file
emb = load_file("cellular_embedding.safetensors")["embedding"] # (1102163, 100) float32
idx = pd.read_csv("taxid_to_index.tsv", sep="\t").set_index("taxid")["idx"]
def poincare_distance(u, v):
"""Geodesic distance in the Poincaré ball (carries the planted radius as well)."""
sq = np.sum((u - v) ** 2)
return np.arccosh(1 + 2 * sq / ((1 - u @ u) * (1 - v @ v)))
def angular_distance(u, v):
"""Angle between directions: where the learned lineage structure lives."""
return np.arccos(np.clip(u @ v / (np.linalg.norm(u) * np.linalg.norm(v)), -1.0, 1.0))
h, m = emb[idx[9606]], emb[idx[10090]] # Homo sapiens, Mus musculus
print(f"Poincaré {poincare_distance(h, m):.3f} angular {angular_distance(h, m):.3f} ‖human‖ {np.linalg.norm(h):.3f}")
# depths and tree distances from the edgelist alone
edges = np.loadtxt("edges_parent_child.tsv", dtype=np.int64) # parent_idx child_idx
parent = np.arange(len(emb)); parent[edges[:, 1]] = edges[:, 0] # root is its own parent
The lineage structure that training determines is in the direction of each vector; the norm is
set by depth through a fixed schedule. Compare taxa by the angle between their directions, or by
Poincaré distance among taxa at the same depth. To look up a taxon by name, join taxid_to_index.tsv
against names.dmp of the same new_taxdump release (2026-06-09).
What the rows are
- 1,102,163 identifiers rooted at "cellular organisms": Bacteria 216,507 · Archaea 7,131 · Eukaryota 878,524 · the root. This mirrors NCBI's own sampling (about 80 % eukaryotic, 0.6 % archaeal); it is not a balanced census of the domains.
- 1,025,397 rows are current taxids. The other 76,766 (7.0 %) are superseded identifiers that the
parser (taxopy) resolved through
merged.dmpand placed as leaves beside their replacement, so a lookup by an old taxid still returns a coordinate. They trained as ordinary siblings: the median cosine between an alias and its replacement equals that between the alias and a random sibling (0.998 for both). - 91,942 rows (8.3 %) carry placeholder binomials of the form " bacterium " (41 % of bacterial and 46 % of archaeal rows; none eukaryotic), placed in the hierarchy by the depositor's clade name. Masking them from the candidate pool at scoring time moves S_angle by +0.0002.
- Viruses are excluded. The build's
--cleanpass pruned unnamed "sp."/"cf."/"aff." placeholders and environmental, uncultured and unidentified samples bottom-up, without removing internal clades.
Training
Topology only: the depth-weighted transitive closure of the tree, no branch lengths and no molecular data. One recipe: a Euclidean tangent parametrization of the ball, a softmax objective over 300 sampled negatives drawn at the descendant's depth, a four-stage depth curriculum, a radial nudge (α = 0.05) toward a log depth schedule, and an effective batch of 2,048 (256 × 8); 200 epochs, Adam (learning rate 10^-3 on a cosine schedule warm-restarted at each curriculum boundary), mixed precision; 25.5 h on one V100. The final checkpoint is released, because angular structure keeps improving after the loss plateaus. A reduced-budget configuration of the same objective collapsed at 498,246 taxa (Metazoa): S_angle 0.619 against 0.973 for the full recipe, three seeds each. The trainer, the Slurm scripts and every evaluation are in the code repository.
Intended use
A fixed-size, differentiable, graded representation of taxonomic position: a soft prior or loss term in a model that needs organismal context, a factor to concatenate with a protein or genome embedding, or a retrieval key for taxonomic neighbours.
Limitations
- Transductive. Fixed to the 2026-06-09 taxdump; a taxon added later has no coordinate short of retraining.
- Taxonomic, not phylogenetic. The geometry reflects the internal consistency of one NCBI release. It agrees with an independent molecular clock to a clade-dependent degree (numbers above) and carries no branch-length information.
- Not seed-reproducible. The released tensor was trained before the trainer was seeded; runs from the current code reproduce exactly on CPU and nearly on GPU.
- Per-rank kNN purity and separation computed with Poincaré distance across depths (family purity@10 0.907, separation 7.6×) inherit the planted radial schedule and are descriptive only; use S_angle or same-depth comparisons as evidence of learned structure.
- Negative sampler. Negatives were drawn at the descendant's depth without excluding the anchor's own descendants, so 47.4 % of drawn negatives were descendants of their anchor. A retraining with the guard restored did not converge on three seeds; its effect on a converged model is open.
- Sampling skew inherited from NCBI (above) bounds any per-domain statement; the archaeal arm is thin.
License
Apache-2.0. Copyright 2025–2026 the TaxPointCare authors (@jcoludar and contributors). The implementation began as a fork of the Nickel & Kiela (2017) Poincaré-embeddings reference code and has since been fully reimplemented.
Citation
If you use this embedding, please cite:
Koludarov, I. & Rost, B. TaxEmbed: one hyperbolic embedding for all named cellular life (2026).