File size: 10,065 Bytes
28c70af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 | # Embedding in bankml and Savante
## Introduction
A language model answers; an **embedding model** measures meaning. It turns a piece of text into a list of numbers
(a vector) so that texts which mean similar things get vectors that point in similar directions. Search by
embedding finds what a question is *about*, not only the words it shares.
mindX, the system Savante belongs to, keeps its memories this way. Its memory store (`agents/memory_pgvector.py`)
embeds every memory and document with **bge-m3** into 1,024 numbers and stores them in PostgreSQL with pgvector.
bankml now uses the same model, so Savante's local history and mindX's memory can be searched by meaning in the same
way.
The embedding model is **optional**. Everything in bankml works without it. With it, the `.history` ragebar also
understands paraphrase, and a published agent's history becomes searchable by meaning in PostgreSQL.
**bankML does not compute embeddings itself yet.** bge-m3 is an XLM-R encoder, and bankML has no encoder graph:
`bankml serve --native` refuses `/api/embed` with HTTP 400 and the reason. The encoder graph, with `/api/embed`,
`/v1/embeddings` and an oracle against llama.cpp's bge-m3 output, is phase O7 of [OLLAMA.md](OLLAMA.md), planned for
0.6.0 ([TODO.md](TODO.md#060--more-models)). Until then Savante asks the local Ollama, as described below.
## Summary
| | |
|---|---|
| model | **bge-m3** (BAAI), the model mindX uses by default (`MINDX_EMBED_MODEL`, default `bge-m3`) |
| licence | **MIT** (read from the model's licence layer in the local Ollama store) |
| weights | 1.16 GB GGUF, served by the local **Ollama** (`ollama pull bge-m3`) |
| pinned by | the sha256 of that GGUF (Ollama's layer digest `daec91ff…3062c` on this laptop), recorded with every vector |
| output | 1,024 dimensions, normalized to unit length; the width of mindX's `VECTOR(1024)` and bankml's `bankml_exchanges.embedding` |
| input | each text cut at 4,000 characters, as mindX cuts it (≈ 1,000 tokens; bge-m3 reads up to 8,192) |
| used for | 1. the `.history` ragebar: words (BM25) and meaning (bge-m3) ranked together; 2. `embedding` in PostgreSQL publishing |
| without it | BM25 alone, and an empty `embedding` column; nothing fails |
| privacy | vectors of private text stay beside `.history` (`<history>.emb`) and leave the machine only if you publish private lines |
## Explanation
### Why an embedding model at all
The ragebar has always searched `.history` with BM25, which ranks exchanges by the words they share with the query,
weighted by how rare each word is. It is fast, exact and explainable, and it is blind to paraphrase: "how fast is
the ternary model" does not match an answer that says "Q2_0 decode took 2.4 s per token".
An embedding model closes that gap. bge-m3 maps the query and every exchange into the same 1,024-dimensional space,
where the angle between two vectors reflects how close their meanings are. bankml does not replace BM25 with it. It
**fuses** the two rankings, so an exact word match still counts and a match in meaning is added to it.
### Why bge-m3
- **It is what mindX uses.** Savante is one of mindX's offices. Using the same model and width means her vectors and
mindX's live in the same space: a history published from bankml can sit in the same database and be compared with
mindX's memories without re-embedding either.
- **It is open.** MIT-licensed, like the rest of what bankml admits ("open source or go away").
- **It reads long text.** Its 8,192-token window holds a whole question and answer; mindX cuts at 4,000 characters
and bankml does the same, so the same text yields the same vector.
- **It is already on this computer.** The local Ollama holds it; bankml adds no new runtime and downloads nothing.
### How it stays out of the way
This laptop has about 1 GB of free memory while the chat model runs, and loading bge-m3 needs about 1.2 GB. So:
- bankml asks for an embedding only when the ragebar or a publish needs one, and Ollama unloads the model 60 seconds
after the last use (`keep_alive`), giving the memory back;
- before loading it, bankml checks free memory (1.3 GB by default) and, if there is not enough, uses BM25 alone
rather than push the machine into swap;
- one embedding call runs at a time; a keystroke that finds it busy is answered by BM25, and the query's vector is
cached so the next keystroke with the same text is instant;
- exchanges are embedded once, in the background, and cached by the sha256 of the embedded text. The ragebar never
waits for indexing; it uses what is ready and says how much is.
## Technical
### The model and its identity
bge-m3 is served by Ollama (0.13.3 here) from its store (`/usr/share/ollama/.ollama/models`). bankml reads the
model's **local manifest** and takes two things from it: the model layer's digest, which is the sha256 of the GGUF
Ollama loads, and the licence layer, which it classifies (MIT). It never asks the registry: what is identified is
what is on disk. Every cached vector records that digest, and a cache written by different weights is ignored.
### Calling it
`POST http://127.0.0.1:11434/api/embed` with `{"model": "bge-m3", "input": [...], "keep_alive": "60s", "truncate":
true}`, the endpoint mindX uses (`/api/embed`, Ollama ≥ 0.3.4). The reply's vectors must be 1,024 long, or the call
is refused. bankml divides each by its length so that the cosine of two vectors is their dot product.
### The private cache
`<history>.emb` beside `.history` (for Savante, `~/.local/share/bankml/savante/savante.history.emb`), one JSON line per
exchange:
```json
{"text_sha256": "…", "model": "bge-m3", "digest": "daec91ff…", "dims": 1024, "vec": "<1,024 float32, little-endian, base64>"}
```
The key is the sha256 of the exact text embedded: `user + "\n" + assistant`, cut at 4,000 characters. An edited or
new exchange gets a new key; nothing is ever re-embedded twice. The file is derived data: delete it and it is
rebuilt.
### Ranking: reciprocal rank fusion
For a query, bankml ranks exchanges twice: by BM25 (mindX's `rage.py` when present, otherwise the built-in BM25),
and by cosine similarity to the query's vector. The rankings are fused by **reciprocal rank fusion** (Cormack, Clarke
and Büttcher 2009): each exchange scores Σ 1 / (60 + rank) over the rankings it appears in. RRF needs no tuning and
no score calibration between the two systems; an exchange near the top of either list rises, one near the top of both
rises most. The ragebar's status line names the engines and how many exchanges are embedded.
### PostgreSQL
When an agent is published **with its private lines** (`connectors.publish(…, include_private=True)`), and the
database has pgvector, every exchange's vector is written to `bankml_exchanges.embedding` (`vector(1024)`) in the
same transaction as the lines themselves, and indexed by pgvectorscale's DiskANN when it is installed, else HNSW. A
query can then find an agent's exchanges by meaning in SQL:
```sql
SELECT seq, line FROM bankml_exchanges WHERE agent = 'ada'
ORDER BY embedding <=> '[…1,024 numbers…]'::vector LIMIT 5;
```
Public-only publishing sends no lines and so no vectors: a vector of private text is itself private.
### Code
| file | what |
|---|---|
| `sAGI/embed.py` | status and provenance, `embed()`, the cache (`index`, `index_async`, `cached`), `query_vector`, `fuse` |
| `sAGI/savante.py` | `history_search()` fuses BM25 and bge-m3; `_semantic()` indexes in the background |
| `sAGI/connectors.py` | `publish()` writes `embedding` with the lines, in one transaction |
| `testing/test_ui.py` | fusion, the cache, and the fallback to BM25 (offline: a fake Ollama) |
## Measured on this laptop (2026-09-29, Ryzen 3 3200U, Ollama 0.13.3)
| | |
|---|---|
| first call (load + embed one text) | 9.2 s |
| next call (three texts, model loaded) | 1.40 s |
| memory while loaded | 1.14 GB; free memory fell from 1.82 GB to 0.43 GB, which is why the 1.3 GB guard exists |
| indexing the three real exchanges in `.history` | 11.5 s (once; cached afterwards) |
| meaning check (cosine to "How fast is the ternary model on this laptop?") | a fact about decode speed 0.448 · the oversight office's oath 0.299 |
| ragebar status | `RAGE (mindX rage.py, BM25) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank` |
The earlier attempt the same day was refused by the guard (1.0 GB free), as designed; this one ran when other
applications had released memory.
## Usage
**Have it.** bge-m3 is in the local Ollama:
```sh
ollama pull bge-m3 # once; 1.16 GB. On this laptop it is already there.
ollama list | grep bge-m3
```
Or adopt it through the Models tab's **Ollama** section like any other model (it is an embedding model, so it is
not offered as a chat carrier).
**Use it.** Nothing to switch on. Open **.history** and type in the ragebar. The status line says, for example:
```
3 of 3 exchanges · RAGE-shaped BM25 (built in) + bge-m3 (meaning, 3 of 3 embedded), fused by reciprocal rank
```
The first search starts indexing the history in the background; results improve as it completes. If bge-m3 cannot
run (not pulled, Ollama down, not enough free memory), the status line names BM25 alone.
**Publish with it.** Agents tab → PostgreSQL → publish **with private lines**. The result reports
`embedded: N of M exchanges embedded with bge-m3 (daec91ff…)`.
**Settings** (environment):
| variable | default | meaning |
|---|---|---|
| `BANKML_EMBED_MODEL` | `bge-m3` | the Ollama model (must produce 1,024 dimensions) |
| `BANKML_OLLAMA` | `http://127.0.0.1:11434` | the Ollama server (keep it local: history text is sent to it) |
| `BANKML_EMBED_KEEP_ALIVE` | `60s` | how long Ollama keeps it loaded after a call |
| `BANKML_EMBED_NEED_GB` | `1.3` | free memory required before loading it |
**Check it.** From the repository root, `python3 -B -c "import sys; sys.path.insert(0, 'sAGI'); import embed;
print(embed.status())"` prints the model, its digest and licence, and whether it is ready.
|