You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

TaxEmbed — one hyperbolic embedding for all named cellular life

TaxEmbed places every taxon identifier of the NCBI Taxonomy of cellular life, 1,102,163 identifiers under "cellular organisms" (new_taxdump downloaded 2026-06-09), into a single 100-dimensional Poincaré ball. It is a lookup: 100 numbers per taxon, fixed-size and differentiable, released so that a model that needs organismal context can take it as input instead of building its own from the taxdump.

Paper: Koludarov I, Rost B. TaxEmbed: one hyperbolic embedding for all named cellular life (2026, preprint to follow). Code: https://github.com/jcoludar/taxembed (Apache-2.0).

What the geometry encodes

  • Radius is taxonomic depth, by construction. Every norm is initialized at a depth-derived target and a radial nudge holds it there during training. On the released tensor the Pearson correlation of norm with depth is +0.957, equal to its value before the first gradient step (+0.957). The radius therefore records that the schedule was applied; it is not a learned property.
  • Direction is learned lineage. Taxa at the same depth are ordered by lineage. The radius-free criterion S_angle (retrieve the k = 10 nearest same-depth taxa by cosine of the directions and score how well that ranking orders them by the depth of their most recent common ancestor with the query; random directions score 0, a perfect ordering 1) reads 0.965 on the released tensor (clade-clustered SE 0.007; 10,000 seeded queries) against an initialization null of −0.004, and 0.941 / 0.962 / 0.978 in the shallow / mid / deep depth bands.
  • Embedded neighbours are tree neighbours. Half of a taxon's ten nearest embedded neighbours are among its ten nearest by tree path: precision@10 0.546 (95 % CI 0.535–0.555; 3,000 queries against 2,000 candidates; chance 0.005).
  • Embedded distance is a taxonomic distance, not a clock. Against TimeTree divergence times (6,000 pre-registered pairs), the Spearman correlation of angular distance is +0.89 in Vertebrata (NCBI path length: +0.79) and +0.30 in Insecta (path length: +0.44). Most of the vertebrate advantage comes from ancestor depth, a property of the tree. A study that needs divergence times should use them directly.

Files

File Size Contents
cellular_embedding.safetensors 441 MB The embedding matrix. Tensor key embedding, shape (1102163, 100), float32. Dimension, taxon count, taxdump date and geometry are in the safetensors header (__metadata__).
taxid_to_index.tsv 16 MB Two columns taxid → idx. Row i of the matrix is the taxon whose idx == i. Byte-identical to the training closure's mapping (MD5 a8ef06b048e03613230a0a14908f515e).
edges_parent_child.tsv 15 MB 1,102,162 lines parent_idx child_idx on the same row indices; the root (cellular organisms, taxid 131567) is row 85051. Depths and tree-path distances follow from it without the taxdump.
training_closure.npz 8.6 MB The training signal: all 21,399,053 ancestor–descendant pairs. Keys ancestor_idx, descendant_idx, depth_diff, ancestor_depth, descendant_depth, ancestor_taxid, descendant_taxid.
closure_manifest.json — Node, edge and pair counts of the closure (1,102,163 / 1,102,162 / 21,399,053; max depth 40).
LICENSE — Apache-2.0.

Quick start

import numpy as np, pandas as pd
from safetensors.numpy import load_file

emb = load_file("cellular_embedding.safetensors")["embedding"]           # (1102163, 100) float32
idx = pd.read_csv("taxid_to_index.tsv", sep="\t").set_index("taxid")["idx"]

def poincare_distance(u, v):
    """Geodesic distance in the Poincaré ball (carries the planted radius as well)."""
    sq = np.sum((u - v) ** 2)
    return np.arccosh(1 + 2 * sq / ((1 - u @ u) * (1 - v @ v)))

def angular_distance(u, v):
    """Angle between directions: where the learned lineage structure lives."""
    return np.arccos(np.clip(u @ v / (np.linalg.norm(u) * np.linalg.norm(v)), -1.0, 1.0))

h, m = emb[idx[9606]], emb[idx[10090]]                                     # Homo sapiens, Mus musculus
print(f"Poincaré {poincare_distance(h, m):.3f}  angular {angular_distance(h, m):.3f}  ‖human‖ {np.linalg.norm(h):.3f}")

# depths and tree distances from the edgelist alone
edges = np.loadtxt("edges_parent_child.tsv", dtype=np.int64)               # parent_idx child_idx
parent = np.arange(len(emb)); parent[edges[:, 1]] = edges[:, 0]           # root is its own parent

The lineage structure that training determines is in the direction of each vector; the norm is set by depth through a fixed schedule. Compare taxa by the angle between their directions, or by Poincaré distance among taxa at the same depth. To look up a taxon by name, join taxid_to_index.tsv against names.dmp of the same new_taxdump release (2026-06-09).

What the rows are

  • 1,102,163 identifiers rooted at "cellular organisms": Bacteria 216,507 · Archaea 7,131 · Eukaryota 878,524 · the root. This mirrors NCBI's own sampling (about 80 % eukaryotic, 0.6 % archaeal); it is not a balanced census of the domains.
  • 1,025,397 rows are current taxids. The other 76,766 (7.0 %) are superseded identifiers that the parser (taxopy) resolved through merged.dmp and placed as leaves beside their replacement, so a lookup by an old taxid still returns a coordinate. They trained as ordinary siblings: the median cosine between an alias and its replacement equals that between the alias and a random sibling (0.998 for both).
  • 91,942 rows (8.3 %) carry placeholder binomials of the form " bacterium " (41 % of bacterial and 46 % of archaeal rows; none eukaryotic), placed in the hierarchy by the depositor's clade name. Masking them from the candidate pool at scoring time moves S_angle by +0.0002.
  • Viruses are excluded. The build's --clean pass pruned unnamed "sp."/"cf."/"aff." placeholders and environmental, uncultured and unidentified samples bottom-up, without removing internal clades.

Training

Topology only: the depth-weighted transitive closure of the tree, no branch lengths and no molecular data. One recipe: a Euclidean tangent parametrization of the ball, a softmax objective over 300 sampled negatives drawn at the descendant's depth, a four-stage depth curriculum, a radial nudge (α = 0.05) toward a log depth schedule, and an effective batch of 2,048 (256 × 8); 200 epochs, Adam (learning rate 10^-3 on a cosine schedule warm-restarted at each curriculum boundary), mixed precision; 25.5 h on one V100. The final checkpoint is released, because angular structure keeps improving after the loss plateaus. A reduced-budget configuration of the same objective collapsed at 498,246 taxa (Metazoa): S_angle 0.619 against 0.973 for the full recipe, three seeds each. The trainer, the Slurm scripts and every evaluation are in the code repository.

Intended use

A fixed-size, differentiable, graded representation of taxonomic position: a soft prior or loss term in a model that needs organismal context, a factor to concatenate with a protein or genome embedding, or a retrieval key for taxonomic neighbours.

Limitations

  • Transductive. Fixed to the 2026-06-09 taxdump; a taxon added later has no coordinate short of retraining.
  • Taxonomic, not phylogenetic. The geometry reflects the internal consistency of one NCBI release. It agrees with an independent molecular clock to a clade-dependent degree (numbers above) and carries no branch-length information.
  • Not seed-reproducible. The released tensor was trained before the trainer was seeded; runs from the current code reproduce exactly on CPU and nearly on GPU.
  • Per-rank kNN purity and separation computed with Poincaré distance across depths (family purity@10 0.907, separation 7.6×) inherit the planted radial schedule and are descriptive only; use S_angle or same-depth comparisons as evidence of learned structure.
  • Negative sampler. Negatives were drawn at the descendant's depth without excluding the anchor's own descendants, so 47.4 % of drawn negatives were descendants of their anchor. A retraining with the guard restored did not converge on three seeds; its effect on a converged model is open.
  • Sampling skew inherited from NCBI (above) bounds any per-domain statement; the archaeal arm is thin.

License

Apache-2.0. Copyright 2025–2026 the TaxPointCare authors (@jcoludar and contributors). The implementation began as a fork of the Nickel & Kiela (2017) Poincaré-embeddings reference code and has since been fully reimplemented.

Citation

If you use this embedding, please cite:

Koludarov, I. & Rost, B. TaxEmbed: one hyperbolic embedding for all named cellular life (2026).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support