bankml / docs /embedding.md
Gregory-L's picture
bankML: the whole source (github.com/cryptoAGI/bankml @ 12ae409) and its page, with the bankML persona; the live engine (Dockerfile, hf/start.sh) ready for Docker hardware
28c70af verified
|
Raw History Blame Contribute Delete
10.1 kB

Embedding in bankml and Savante

Introduction

A language model answers; an embedding model measures meaning. It turns a piece of text into a list of numbers (a vector) so that texts which mean similar things get vectors that point in similar directions. Search by embedding finds what a question is about, not only the words it shares.

mindX, the system Savante belongs to, keeps its memories this way. Its memory store (agents/memory_pgvector.py) embeds every memory and document with bge-m3 into 1,024 numbers and stores them in PostgreSQL with pgvector. bankml now uses the same model, so Savante's local history and mindX's memory can be searched by meaning in the same way.

The embedding model is optional. Everything in bankml works without it. With it, the .history ragebar also understands paraphrase, and a published agent's history becomes searchable by meaning in PostgreSQL.

bankML does not compute embeddings itself yet. bge-m3 is an XLM-R encoder, and bankML has no encoder graph: bankml serve --native refuses /api/embed with HTTP 400 and the reason. The encoder graph, with /api/embed, /v1/embeddings and an oracle against llama.cpp's bge-m3 output, is phase O7 of OLLAMA.md, planned for 0.6.0 (TODO.md). Until then Savante asks the local Ollama, as described below.

Summary

model bge-m3 (BAAI), the model mindX uses by default (MINDX_EMBED_MODEL, default bge-m3)
licence MIT (read from the model's licence layer in the local Ollama store)
weights 1.16 GB GGUF, served by the local Ollama (ollama pull bge-m3)
pinned by the sha256 of that GGUF (Ollama's layer digest daec91ff…3062c on this laptop), recorded with every vector
output 1,024 dimensions, normalized to unit length; the width of mindX's VECTOR(1024) and bankml's bankml_exchanges.embedding
input each text cut at 4,000 characters, as mindX cuts it (≈ 1,000 tokens; bge-m3 reads up to 8,192)
used for 1. the .history ragebar: words (BM25) and meaning (bge-m3) ranked together; 2. embedding in PostgreSQL publishing
without it BM25 alone, and an empty embedding column; nothing fails
privacy vectors of private text stay beside .history (<history>.emb) and leave the machine only if you publish private lines

Explanation

Why an embedding model at all

The ragebar has always searched .history with BM25, which ranks exchanges by the words they share with the query, weighted by how rare each word is. It is fast, exact and explainable, and it is blind to paraphrase: "how fast is the ternary model" does not match an answer that says "Q2_0 decode took 2.4 s per token".

An embedding model closes that gap. bge-m3 maps the query and every exchange into the same 1,024-dimensional space, where the angle between two vectors reflects how close their meanings are. bankml does not replace BM25 with it. It fuses the two rankings, so an exact word match still counts and a match in meaning is added to it.

Why bge-m3

  • It is what mindX uses. Savante is one of mindX's offices. Using the same model and width means her vectors and mindX's live in the same space: a history published from bankml can sit in the same database and be compared with mindX's memories without re-embedding either.
  • It is open. MIT-licensed, like the rest of what bankml admits ("open source or go away").
  • It reads long text. Its 8,192-token window holds a whole question and answer; mindX cuts at 4,000 characters and bankml does the same, so the same text yields the same vector.
  • It is already on this computer. The local Ollama holds it; bankml adds no new runtime and downloads nothing.

How it stays out of the way

This laptop has about 1 GB of free memory while the chat model runs, and loading bge-m3 needs about 1.2 GB. So:

  • bankml asks for an embedding only when the ragebar or a publish needs one, and Ollama unloads the model 60 seconds after the last use (keep_alive), giving the memory back;
  • before loading it, bankml checks free memory (1.3 GB by default) and, if there is not enough, uses BM25 alone rather than push the machine into swap;
  • one embedding call runs at a time; a keystroke that finds it busy is answered by BM25, and the query's vector is cached so the next keystroke with the same text is instant;
  • exchanges are embedded once, in the background, and cached by the sha256 of the embedded text. The ragebar never waits for indexing; it uses what is ready and says how much is.

Technical

The model and its identity

bge-m3 is served by Ollama (0.13.3 here) from its store (/usr/share/ollama/.ollama/models). bankml reads the model's local manifest and takes two things from it: the model layer's digest, which is the sha256 of the GGUF Ollama loads, and the licence layer, which it classifies (MIT). It never asks the registry: what is identified is what is on disk. Every cached vector records that digest, and a cache written by different weights is ignored.

Calling it

POST http://127.0.0.1:11434/api/embed with {"model": "bge-m3", "input": [...], "keep_alive": "60s", "truncate": true}, the endpoint mindX uses (/api/embed, Ollama ≥ 0.3.4). The reply's vectors must be 1,024 long, or the call is refused. bankml divides each by its length so that the cosine of two vectors is their dot product.

The private cache

<history>.emb beside .history (for Savante, ~/.local/share/bankml/savante/savante.history.emb), one JSON line per exchange:

{"text_sha256": "…", "model": "bge-m3", "digest": "daec91ff…", "dims": 1024, "vec": "<1,024 float32, little-endian, base64>"}

The key is the sha256 of the exact text embedded: user + "\n" + assistant, cut at 4,000 characters. An edited or new exchange gets a new key; nothing is ever re-embedded twice. The file is derived data: delete it and it is rebuilt.

Ranking: reciprocal rank fusion

For a query, bankml ranks exchanges twice: by BM25 (mindX's rage.py when present, otherwise the built-in BM25), and by cosine similarity to the query's vector. The rankings are fused by reciprocal rank fusion (Cormack, Clarke and Büttcher 2009): each exchange scores Σ 1 / (60 + rank) over the rankings it appears in. RRF needs no tuning and no score calibration between the two systems; an exchange near the top of either list rises, one near the top of both rises most. The ragebar's status line names the engines and how many exchanges are embedded.

PostgreSQL

When an agent is published with its private lines (connectors.publish(…, include_private=True)), and the database has pgvector, every exchange's vector is written to bankml_exchanges.embedding (vector(1024)) in the same transaction as the lines themselves, and indexed by pgvectorscale's DiskANN when it is installed, else HNSW. A query can then find an agent's exchanges by meaning in SQL:

SELECT seq, line FROM bankml_exchanges WHERE agent = 'ada'
ORDER BY embedding <=> '[…1,024 numbers…]'::vector LIMIT 5;

Public-only publishing sends no lines and so no vectors: a vector of private text is itself private.

Code

file what
sAGI/embed.py status and provenance, embed(), the cache (index, index_async, cached), query_vector, fuse
sAGI/savante.py history_search() fuses BM25 and bge-m3; _semantic() indexes in the background
sAGI/connectors.py publish() writes embedding with the lines, in one transaction
testing/test_ui.py fusion, the cache, and the fallback to BM25 (offline: a fake Ollama)

Measured on this laptop (2026-09-29, Ryzen 3 3200U, Ollama 0.13.3)

first call (load + embed one text) 9.2 s
next call (three texts, model loaded) 1.40 s
memory while loaded 1.14 GB; free memory fell from 1.82 GB to 0.43 GB, which is why the 1.3 GB guard exists
indexing the three real exchanges in .history 11.5 s (once; cached afterwards)
meaning check (cosine to "How fast is the ternary model on this laptop?") a fact about decode speed 0.448 · the oversight office's oath 0.299
ragebar status RAGE (mindX rage.py, BM25) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank

The earlier attempt the same day was refused by the guard (1.0 GB free), as designed; this one ran when other applications had released memory.

Usage

Have it. bge-m3 is in the local Ollama:

ollama pull bge-m3          # once; 1.16 GB. On this laptop it is already there.
ollama list | grep bge-m3

Or adopt it through the Models tab's Ollama section like any other model (it is an embedding model, so it is not offered as a chat carrier).

Use it. Nothing to switch on. Open .history and type in the ragebar. The status line says, for example:

3 of 3 exchanges · RAGE-shaped BM25 (built in) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank

The first search starts indexing the history in the background; results improve as it completes. If bge-m3 cannot run (not pulled, Ollama down, not enough free memory), the status line names BM25 alone.

Publish with it. Agents tab → PostgreSQL → publish with private lines. The result reports embedded: N of M exchanges embedded with bge-m3 (daec91ff…).

Settings (environment):

variable default meaning
BANKML_EMBED_MODEL bge-m3 the Ollama model (must produce 1,024 dimensions)
BANKML_OLLAMA http://127.0.0.1:11434 the Ollama server (keep it local: history text is sent to it)
BANKML_EMBED_KEEP_ALIVE 60s how long Ollama keeps it loaded after a call
BANKML_EMBED_NEED_GB 1.3 free memory required before loading it

Check it. From the repository root, python3 -B -c "import sys; sys.path.insert(0, 'sAGI'); import embed; print(embed.status())" prints the model, its digest and licence, and whether it is ready.