Cogito Estella (v0.17.0)
Non-autoregressive decoder heads that turn sentence embeddings into knowledge-graph
triples in one forward pass. They power cogito-mcp, a graph memory for agents that
answers questions from ingested documents in a few hundred tokens.
Code, tests and benchmark: https://github.com/DeliVali/cogito-estella
Encoders and the manifest
encoders.json maps each encoder to its assets and to the decode thresholds they were
validated under; cogito-mcp reads it on first run and downloads the default encoder's
assets from the m2m100-pool/ subfolder. --encoder sonar selects the SONAR-space heads
at the repository root.
| Encoder | Assets | Triple F1 | Licence |
|---|---|---|---|
m2m100-pool (default) |
m2m100-pool/cogito-prose-ontology-m2mpool{,-s2,-s3}.safetensors + .json sidecars, pool.safetensors + pool.json, vocab-onto-m2mpool.json |
0.852 | heads and pooling Apache-2.0; M2M-100 encoder MIT (facebook/m2m100_418M, revision 55c2e61b, downloaded from its own repository) |
sonar |
cogito-prose-ontology{,-s2,-s3}.pt, vocab-onto.json |
0.796 | heads Apache-2.0; SONAR encoder CC-BY-NC 4.0 (non-commercial), distributed separately by Meta |
Both ensembles decode 76 relations over a 20k-entity vocabulary. F1 is measured on held-out sentences whose entity/relation combinations never appeared in training, once per ensemble on a virgin slice after threshold selection; on 5,000 sentences held out under both protocols the M2M-100 ensemble scores 0.819 against 0.730. Each safetensors file carries its contract (encoder, revision, width, normalization, vocabulary and pooling sha256, F1, step) in the header and in the sidecar.
Earlier research heads from v0.7.0 stay in the repository, all in the SONAR space:
cogito-toolcalls-graphdecoder.pt (tool calls, F1 1.000), cogito-code-lora-adapters.pt
(Python code, 0.781; the LoRA modifies SONAR and inherits CC-BY-NC 4.0),
cogito-prose-candidates-{ft,cal,base,s2,s3}.pt (entity-conditioned prose, 0.827),
cogito-prose-openvocab{,-s4,-s5}.pt + cogito-prose-cascade-fallback.pt (open-vocab
prose, 0.651) and vocab-prose.json.
Additional m2m100-pool heads
Standalone research assets (not part of the default ask/ingest path), each measured
above its SONAR original on the m2m100-pool encoder:
| File | Task | Triple F1 | vs SONAR |
|---|---|---|---|
m2m100-pool/cogito-prose-candidates-m2mpool{,-s2,-s3,-s4,-s5}.safetensors + vocab-candidates-m2mpool.json |
entity-conditioned prose, 60 raw-verb relations (5-model ensemble) | 0.878 | 0.827 |
m2m100-pool/cogito-toolcalls-m2mpool.safetensors |
tool-call extraction | 1.000 | 1.000 |
m2m100-pool/cogito-code-lora-m2mpool.safetensors + vocab-code-lora-m2mpool.json |
Python code, calls/imports | 0.979 | 0.777 |
Candidates loads through CogitoGraphExtractor unchanged (pass the checkpoints and
vocab above with threshold=0.1, adj_threshold=0.6) and tool-calls through a plain
GraphDecoder (K=24, node_dim=448, V=8192, R=48) — both validated by construction or on
a held-out combination split like the other heads in this repository.
The code head needed more than a frozen encoder: a from-scratch attempt with the same
trunk recipe as the other heads plateaued at F1 0.38 while its training loss overfit to
zero. What closed the gap was adapting the encoder itself — the same lever SONAR's code
head used (0.652 frozen -> 0.777 LoRA) — applied here to M2M-100 instead: LoRA (r=32,
α=64, dropout 0.1) on every attention and feed-forward projection across its 12 encoder
layers, fine-tuned jointly with the pooling head and a fresh trunk+decoder. Because the
base encoder is MIT rather than SONAR's CC-BY-NC, this LoRA adapter carries no
non-commercial term either. A reference loader that reconstructs the adapted encoder and
runs extraction end to end ships at
load_code_lora.py:
from load_code_lora import load
extract = load()
extract("import numpy as np") # [('ROOT', 'imports', 'numpy')]
Usage
pip install "cogito-estella[mcp]"
cogito-mcp --dir .cogito # weights resolve from this repository on first run
from cogito_estella.integrations.llamaindex_connector import CogitoGraphExtractor
from cogito_estella.mcp.weights import resolve
w = resolve(None, None)
ex = CogitoGraphExtractor([str(c) for c in w.checkpoints], str(w.vocab), pool_path=w.pool,
threshold=w.operating_point[0], adj_threshold=w.operating_point[1])
ex.extract("The committee approved the new budget.") # [(subject, relation, object), ...]
Intended use and limitations
Structured-knowledge extraction for agent memory and GraphRAG ingestion. Entities are chosen among the nouns of the sentence (or caller-supplied candidates) within a 20k-lemma vocabulary and relations among 76 classes; content outside those spaces is not captured. The training oracle is dependency-parse SVO, so graphs over free prose inherit its noise; the pooling was trained on English sentences up to 128 tokens. The SONAR-space heads need the SONAR encoder at inference, which Meta distributes under CC-BY-NC 4.0. All heads and the pooling are Apache-2.0.
