Cogito Estella (v0.17.0)

Non-autoregressive decoder heads that turn sentence embeddings into knowledge-graph triples in one forward pass. They power cogito-mcp, a graph memory for agents that answers questions from ingested documents in a few hundred tokens.

Code, tests and benchmark: https://github.com/DeliVali/cogito-estella

demo

Encoders and the manifest

encoders.json maps each encoder to its assets and to the decode thresholds they were validated under; cogito-mcp reads it on first run and downloads the default encoder's assets from the m2m100-pool/ subfolder. --encoder sonar selects the SONAR-space heads at the repository root.

Encoder Assets Triple F1 Licence
m2m100-pool (default) m2m100-pool/cogito-prose-ontology-m2mpool{,-s2,-s3}.safetensors + .json sidecars, pool.safetensors + pool.json, vocab-onto-m2mpool.json 0.852 heads and pooling Apache-2.0; M2M-100 encoder MIT (facebook/m2m100_418M, revision 55c2e61b, downloaded from its own repository)
sonar cogito-prose-ontology{,-s2,-s3}.pt, vocab-onto.json 0.796 heads Apache-2.0; SONAR encoder CC-BY-NC 4.0 (non-commercial), distributed separately by Meta

Both ensembles decode 76 relations over a 20k-entity vocabulary. F1 is measured on held-out sentences whose entity/relation combinations never appeared in training, once per ensemble on a virgin slice after threshold selection; on 5,000 sentences held out under both protocols the M2M-100 ensemble scores 0.819 against 0.730. Each safetensors file carries its contract (encoder, revision, width, normalization, vocabulary and pooling sha256, F1, step) in the header and in the sidecar.

Earlier research heads from v0.7.0 stay in the repository, all in the SONAR space: cogito-toolcalls-graphdecoder.pt (tool calls, F1 1.000), cogito-code-lora-adapters.pt (Python code, 0.781; the LoRA modifies SONAR and inherits CC-BY-NC 4.0), cogito-prose-candidates-{ft,cal,base,s2,s3}.pt (entity-conditioned prose, 0.827), cogito-prose-openvocab{,-s4,-s5}.pt + cogito-prose-cascade-fallback.pt (open-vocab prose, 0.651) and vocab-prose.json.

Additional m2m100-pool heads

Standalone research assets (not part of the default ask/ingest path), each measured above its SONAR original on the m2m100-pool encoder:

File Task Triple F1 vs SONAR
m2m100-pool/cogito-prose-candidates-m2mpool{,-s2,-s3,-s4,-s5}.safetensors + vocab-candidates-m2mpool.json entity-conditioned prose, 60 raw-verb relations (5-model ensemble) 0.878 0.827
m2m100-pool/cogito-toolcalls-m2mpool.safetensors tool-call extraction 1.000 1.000
m2m100-pool/cogito-code-lora-m2mpool.safetensors + vocab-code-lora-m2mpool.json Python code, calls/imports 0.979 0.777

Candidates loads through CogitoGraphExtractor unchanged (pass the checkpoints and vocab above with threshold=0.1, adj_threshold=0.6) and tool-calls through a plain GraphDecoder (K=24, node_dim=448, V=8192, R=48) — both validated by construction or on a held-out combination split like the other heads in this repository.

The code head needed more than a frozen encoder: a from-scratch attempt with the same trunk recipe as the other heads plateaued at F1 0.38 while its training loss overfit to zero. What closed the gap was adapting the encoder itself — the same lever SONAR's code head used (0.652 frozen -> 0.777 LoRA) — applied here to M2M-100 instead: LoRA (r=32, α=64, dropout 0.1) on every attention and feed-forward projection across its 12 encoder layers, fine-tuned jointly with the pooling head and a fresh trunk+decoder. Because the base encoder is MIT rather than SONAR's CC-BY-NC, this LoRA adapter carries no non-commercial term either. A reference loader that reconstructs the adapted encoder and runs extraction end to end ships at load_code_lora.py:

from load_code_lora import load
extract = load()
extract("import numpy as np")   # [('ROOT', 'imports', 'numpy')]

Usage

pip install "cogito-estella[mcp]"
cogito-mcp --dir .cogito          # weights resolve from this repository on first run
from cogito_estella.integrations.llamaindex_connector import CogitoGraphExtractor
from cogito_estella.mcp.weights import resolve

w = resolve(None, None)
ex = CogitoGraphExtractor([str(c) for c in w.checkpoints], str(w.vocab), pool_path=w.pool,
                          threshold=w.operating_point[0], adj_threshold=w.operating_point[1])
ex.extract("The committee approved the new budget.")    # [(subject, relation, object), ...]

Intended use and limitations

Structured-knowledge extraction for agent memory and GraphRAG ingestion. Entities are chosen among the nouns of the sentence (or caller-supplied candidates) within a 20k-lemma vocabulary and relations among 76 classes; content outside those spaces is not captured. The training oracle is dependency-parse SVO, so graphs over free prose inherit its noise; the pooling was trained on English sentences up to 128 tokens. The SONAR-space heads need the SONAR encoder at inference, which Meta distributes under CC-BY-NC 4.0. All heads and the pooling are Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using DeliVali/cogito-estella 1