Download docs/usage.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 60.5 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/usage.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/usage.md
-
curl -L -o usage.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/usage.md
Using bankml and Savante
How to install bankml, start it, and use it on your own computer. Everything runs on this machine: the model, the gate that verifies it, and Savante's page. New to Savante's page and her voice? Read playback.md. Every flag, environment variable and service setting is in install.md; each source module, its use and its limits, in modules/.
This guide covers the latest release, 0.3.6. Sections marked 0.3.7, 0.3.8 or 0.3.9 describe the next release, which is not yet tagged; the details are in the Unreleased sections of CHANGELOG.md.
- 1. Install
- 2. What you need
- 3. Start, stop, check
- 4. Talk to Savante
- 5. Models
- 6. By hand: build, verify, serve
- 6a. Ollama's API
- 6b. The C API
- 6c. The bankML console: bankML as itself (0.3.7)
- 7. Let others watch (view mode, on the LAN)
- 8. The files Savante keeps:
.history,.memory,.prompt - 8a. Proof of data without the data
- 8b. Custom agents from the Savante template
- 8c. THOT bundles: the dataset an iNFT points to
- 8d. PostgreSQL: publish an agent, load one back
- 8e. iNFT: mint an agent, load one from a token
- 9. Receipts, and how to check an answer
- 10. Savante's canon and the iNFT ledger
- 11. Testing and the release gate
- 12. Troubleshooting
- 13. Reference: commands, ports, environment
1. Install
One command, from a fresh machine to Savante answering on your own CPU:
git clone https://github.com/cryptoAGI/bankml && cd bankml
chmod +x install.sh
./install.sh
chmod +x install.sh makes the installer executable. A git clone keeps the bit already, so the line is harmless
there; a copy that arrived another way (a zip download, a USB stick, some file systems) may have lost it, and then
./install.sh answers permission denied. bash install.sh works either way.
or, without cloning first (the installer clones bankml to ~/bankml, or to $BANKML_DIR):
curl -fsSL https://raw.githubusercontent.com/cryptoAGI/bankml/main/install.sh | bash
It runs these steps in order, and stops at the first one that fails, saying why:
| step | what it does |
|---|---|
check |
Rust 1.99+ (pinned in rust-toolchain.toml), Python 3.10+, git, curl, tar, ss, sha256sum; reports AVX2, RAM and free disk |
build |
cargo build --release and cargo test --release |
engine |
downloads llama.cpp b11192 (17 MB) and refuses it unless its sha256 is the published one |
python |
uses your Python if it has Gradio and numpy; otherwise makes a venv with Gradio 3 |
canon |
clones Savante's canon to ~/cryptoAGI/savante (beside jaimla and luvai) if it is not there; an existing canon is left as it is, and a machine that still has the older ~/savante keeps using it |
model |
imports Bonsai-8B (1.16 GB) if absent, verifies it, and starts bankml serve with llama-server behind it |
start |
starts Savante on http://127.0.0.1:7873 |
Run any steps on their own: ./install.sh build, ./install.sh model start. Options: --skip-tests, --no-start,
--view (also start view mode, §7), --voice (also install Savante's voice; see playback.md).
Nothing needs sudo. What the installer chose is kept in ~/.local/share/bankml/install.env, and logs in
~/.local/share/bankml/logs/.
2. What you need
| what | where it comes from |
|---|---|
| Linux on x86_64 (AVX2 for the fast kernels; without it bankml is correct, just slower) | — |
Rust 1.99 (pinned in rust-toolchain.toml; rustup fetches it) |
rustup |
| Python 3.10+ | your system |
| about 2 GB of free RAM and 3 GB of free disk | the 1-bit 8B model maps 1.16 GB |
The installer fetches the rest: llama.cpp b11192 (llama-b11192-bin-ubuntu-x64.tar.gz, sha256
34cf6fa5de9da0db3932c78fe15fed2fbca17451e665dac0a4f6a3c8fc881ec7), Gradio 3 if needed, the model from
PYTHAI/Bonsai-8B-gguf-fork, and Savante's canon (§10). On
another platform, build llama.cpp b11192 yourself and set BANKML_LLAMA_SERVER to its llama-server.
The ternary model (2.31 GB) also works and gives better answers, but llama.cpp runs it about 5× slower on a CPU (see PERFORMANCE.md).
3. Start, stop, check
./install.sh start # bankml serve (if not up) is started by `model`; this starts Savante's page
./install.sh start --view # also the read-only page for your network
./install.sh status # what is running, and which model bankml verified
./install.sh stop # stops the page(s), bankml serve and llama-server
What runs where:
you (browser) ──► Savante's page (127.0.0.1:7873) ──► bankml serve (127.0.0.1:18093) ──► llama-server (127.0.0.1:18092)
verifies the model file runs the verified file
anyone on the LAN ──► view mode (0.0.0.0:7874, read-only)
- bankml serve is the gate. It refuses to start unless the model passes the guard, its sha256 equals the pin in
the model's
FORK.json, and llama-server serves that same file. Every answer carries a receipt (§9). - With the native engine (
--native; Savante'sautoengine setting chooses it for the ternary files), there is no llama-server:bankml servecomputes the answer itself and also listens on 18092 with llama-server's endpoints. - Savante's page is reachable only from this computer.
- View mode is a read-only page for others on your network (§7).
4. Talk to Savante
Open http://127.0.0.1:7873, type a question, press Send.
- The first answer takes about 2 minutes on a laptop CPU, while the model reads Savante's instructions. Later answers start sooner.
- Every answer is a draft · not a finding. The chat shows the answer alone; its clock, receipt (§9) and prompt
provenance are on the Admin tab under the last answer, and in
.historyand the Responses tab. - To ask for a review, say so: "review: is Bonsai-8B ready to serve mindX?". She then answers as FINDINGS / VERDICT / RATIONALE / CONDITIONS / RISKS WATCHED. The 1-bit carrier was graded REJECT for review duty in the sAGI carrier test, so treat its verdicts as drafts.
- Stop cancels an answer; New session starts a fresh conversation.
The Interaction tab is only the conversation: the chat, the question field, the clock and the portrait (press it for
the card; ← back to Interaction or Esc closes it, and every other tab has a ← Interaction button). The
Admin tab holds the settings: the system prompt (.prompt, §8), .memory, max tokens, temperature, the carrier,
and the engine's CPU threads and RAM budget (Apply restarts the engine with them, and restores your previous
settings if it will not start; the engine choice and n-gram speculation are under advanced). The other tabs hold
your history, notes, metrics, integrity checks and the verifier. playback.md walks through the
page, her card and her voice.
5. Models
The Models tab and sAGI/models.py import a model, verify it, and switch to it:
python3 sAGI/models.py list # what is here, pinned or not, and which one is running
python3 sAGI/models.py catalog # the curated catalogue, and whether each fits this machine
python3 sAGI/models.py import qwen3-1.7b # a catalogue id, a Hugging Face URL, or ollama:NAME:TAG
python3 sAGI/models.py use Qwen3-1.7B-Q8_0.gguf # switch (rolls back if the new model fails)
An import is kept only if its sha256 equals the one its publisher lists, bankml guard passes, and a FORK.json
pin is written; bankml serve checks that pin before every start. Only open-source models are accepted (Gemma and
Llama are refused), and an import that would not fit this machine's RAM or disk is refused with the numbers.
Converted models (0.3.4). SmolLM2-135M-Instruct and mindX's mindx-gen39 have no published F16 GGUF. They are
made from their pinned safetensors by llama.cpp b11192's own converter, and sAGI/models.py pins the result against
the recorded conversion (CONVERTED in the file says how each was made):
python3 llama.cpp/convert_hf_to_gguf.py SmolLM2-135M-Instruct/ --outtype f16 --outfile .models/SmolLM2-135M-Instruct-F16.gguf
# (b11192, commit 171e8846; the safetensors at HuggingFaceTB/SmolLM2-135M-Instruct@12fd25f7; torch is needed here only)
python3 -c "import sys; sys.path.insert(0, 'sAGI'); import models; print(models.adopt('SmolLM2-135M-Instruct-F16.gguf'))"
# → hashes the file against the recorded e9aba089…, re-reads the source's LFS sha256 at the revision, writes FORK.json
What bankML's own forward pass plays (bankml serve --native, bankml generate, the C API): Qwen3 in Q1_0 and
Q2_0_g64 (Bonsai-8B, Ternary-Bonsai-8B, Bonsai-1.7B) and the Llama architecture in F16 (SmolLM2-135M-Instruct,
mindx-gen39). Anything else is refused with the reason; /api/tags lists it with that reason.
6. By hand: build, verify, serve
What the installer does, step by step:
cargo build --release && cargo test --release
target/release/bankml guard .models/Bonsai-8B-Q1_0.gguf # play | refuse (reason) | need_more
target/release/bankml verify .models/Bonsai-8B-Q1_0.gguf --fork FORK.json # guard, then the sha256 pin
target/release/bankml serve .models/Bonsai-8B-Q1_0.gguf --fork FORK.json \
--spawn /path/to/llama-b11192/llama-server --threads 3 --ctx 2048
python3 sAGI/savante.py --mode interact
verify exits 0 on play, 2 on refuse, 1 on an I/O error, 3 on a truncated file. serve can also sit in front of a
llama-server you started (--upstream 127.0.0.1:18092); it then checks through /props that the server runs the
verified file. It listens on 127.0.0.1:18093 and speaks the OpenAI API:
curl -s 127.0.0.1:18093/bankml # what was verified, when, and what is upstream
curl -s 127.0.0.1:18093/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"Say hello. /no_think"}],"max_tokens":32}'
JSON mode and grammars (0.3.3)
/v1/chat/completions (and the C API's bankml_chat) takes what llama-server takes, resolved the way it resolves it:
B=127.0.0.1:18093; J='Content-Type: application/json'
curl -s $B/v1/chat/completions -H "$J" -d '{"messages": [{"role": "user", "content": "Describe a cat."}],
"response_format": {"type": "json_object"}, "temperature": 0, "max_tokens": 64}'
curl -s $B/v1/chat/completions -H "$J" -d '{"messages": [{"role": "user", "content": "Is the sun a star?"}],
"grammar": "root ::= (\"yes\" | \"no\") \".\""}'
echo 'Describe a cat.' | target/release/bankml generate .models/Bonsai-8B-Q1_0.gguf --json --max 64
response_format: {"type": "json_object"}, and the schemas{}and{"type": "object"}(inresponse_formator the top-leveljson_schema), give llama-server b11192's grammar for this template, byte for byte. Its root begins with the generation prompt, which the sampler takes in before the first token, as the server does. Under that grammar the model may open with a fenced```jsonblock, andcontentis the JSON value alone. That is the object once it closes, or what there is of it whenmax_tokenscuts the answer. A stream sends the content's growth.grammar: any GBNF that llama.cpp parses, with rootroot. It is taken as a user grammar, with no prefill, andcontentis the text as generated. A grammar llama.cpp would not parse is refused with its parser's reason.- Every token is drawn as llama-server draws it under a grammar. The usual sampler chain runs first. If its token breaks the grammar, the logits are masked and the chain runs again, which takes a second draw from the seeded generator. Greedy and seeded answers are token-identical to the server (the oracle is in oracles.md).
- Any other JSON schema (0.3.5): in
response_format(json_schema, orjson_objectwith aschema), the top-leveljson_schema, or Ollama'sformat: {…}, it is converted into the grammar llama-server b11192 builds for it on the model's own template (bankML/schema.rs; byte-identical on 173 of 173 schemas on each template against llama.cpp's own code), and answered as JSON mode is: prefilled, fenced or bare,contentthe value alone (a top-level string, number or literal included). The answers are token-identical to llama-server's on every native model, greedy and seeded. A schema llama.cpp refuses is refused with its message. - An answer cut by
max_tokensgets llama-server's content: the value as far as it got (an unfinished escape dropped), or, when nothing of the value came yet ("```json\n"), the raw text, as the server answers a whole request. A stream sends only the value's growth. - Refused, with the reason: a schema llama.cpp b11192 refuses; a
patternllama.cpp would turn into a grammar that is not UTF-8; a non-object ornulltop-leveljson_schema(llama-server fails those requests);response_formattogether withgrammar;json_schematogether withgrammar; aresponse_formattype other thantext,json_objectorjson_schema.
The context limit, slots and logprobs (0.3.8)
In the next release (not yet tagged); each is checked against llama-server b11192's own answers (CHANGELOG.md, modules/serve.md, modules/prompt_cache.md).
At the context limit the native engine answers as llama-server does with context shift off: a generation that reaches
--ctxstops withfinish_reason: "length"; a prompt that does not fit is refused at once with HTTP 400 and llama-server'sexceed_context_size_errorbody, which names both counts — so a client can trim and retry.Slots. Start
bankml serve --nativewith--slot-dir DIR(created at start). Save the conversation's KV cache by name, restore it later — in another process too — and skip recomputing the prompt. Without--slot-dirthe slot actions answer 501, as llama-server's do:curl -s 127.0.0.1:18092/slots/0?action=save -H 'Content-Type: application/json' -d '{"filename":"warm.bin"}' curl -s 127.0.0.1:18092/slots/0?action=restore -H 'Content-Type: application/json' -d '{"filename":"warm.bin"}' curl -s 127.0.0.1:18092/slots/0?action=erase -H 'Content-Type: application/json' -d '{}'The file is checked when it is read (the model's sha256 and a sha256 trailer): a file from another model or a damaged one is refused and the slot emptied (never half-filled), so the next answer simply recomputes. Filenames are plain names inside the directory; a path is refused.
Conversations that take turns. The native engine has one slot and serves requests one at a time, as llama-server
-np 1. Like llama-server, it keeps the states of other conversations in RAM (BANKML_CACHE_RAM, MiB; 0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after another does not recompute its whole history.Your bankML from a web page (0.3.9).
./install.sh start --spacelets the bankML Space's page (PYTHAI/bankml, Your own bankML) talk to your ownbankml servefrom your browser: your CPU, your verified model, a receipt on every answer, nothing sent anywhere else. It starts serve with--allow-origin https://pythai-bankml.static.hf.space, so only that one page gets CORS headers;./install.sh start --no-spacetakes it back.A smaller conversation memory (0.3.9, in progress).
BANKML_CACHE_TYPE=q8_0keeps the KV cache in q8_0 instead of f16, about half the memory, so a long context fits on a small machine. It is llama.cpp's--cache-type-k q8_0 --cache-type-v q8_0exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the same tokens as llama-server so configured (testing/kv_oracle.py). A slot saved with one cache type is refused by an engine with the other.Logprobs.
"logprobs": true, "top_logprobs": 5on/v1/chat/completionsreturns each token's log-probability and the five most likely alternatives (top_logprobsdefaults to 20), the same floats llama-server reports. Streamed answers carry each token's entry in the chunk its text makes.top_logprobswithoutlogprobsis refused with llama-server's message. llama-server's own/completionendpoint (and itsn_probs) is not served; use/v1.
6a. Ollama's API
Since 0.3.1 bankml serve --native also speaks Ollama's API, over the same engine, gate and receipts. Native
only: bankML answers what its own verified forward pass can do and refuses the rest with HTTP 400
{"error": "…"} that says why; it never hands a request to llama-server or to Ollama. What mindX asks of Ollama,
and what is still missing, is in OLLAMA.md.
target/release/bankml serve .models/Bonsai-8B-Q1_0.gguf --fork ~/.local/share/bankml/forks/Bonsai-8B-Q1_0.gguf.FORK.json \
--native --registry --ctx 2048 # --registry [DIR]: every model pinned there, by name; --keep-alive 5m
B=127.0.0.1:18093; J='Content-Type: application/json'
curl -s $B/api/version
curl -s $B/api/tags # every pin: name, size, digest = its sha256, details, "bankml": {"native", "reason"}
curl -s $B/api/ps # the resident model and its expires_at
curl -s $B/api/show -H "$J" -d '{"model": "bonsai-8b-q1_0"}'
curl -s $B/api/chat -H "$J" -d '{"model": "bonsai-8b-q1_0", "messages": [{"role": "user", "content": "Name a prime."}],
"stream": false, "options": {"temperature": 0.3, "seed": 7, "num_predict": 32}}'
curl -s $B/api/generate -H "$J" -d '{"model": "bonsai-8b-q1_0", "system": "Answer in one word.", "prompt": "Capital of France?", "stream": false}'
curl -s $B/api/generate -H "$J" -d '{"model": "bonsai-8b-q1_0", "keep_alive": 0}' # unload now (done_reason "unload")
curl -s $B/api/generate -H "$J" -d '{"model": "ternary-bonsai-8b-q2_0_g64", "keep_alive": "30m"}' # verify and load (done_reason "load")
Names. A model's name is its file's stem, lower-cased:
bonsai-8b-q1_0,ternary-bonsai-8b-q2_0_g64,smollm2-135m-instruct-f16,mindx-gen39-f16.:latestand the file name are accepted as aliases.- Since 0.3.4 so is the name without its weight-type suffix, when exactly one pin has that base:
mindx-gen39(Ollama's tag for mindX's generation 39),smollm2-135m-instruct,bonsai-1.7b. - A request without a model uses the one
servestarted with. - An unknown name answers 404.
Where models come from. Without
--registry, only the startup model is served. With it, every GGUF pinned inDIR(default$BANKML_FORKS, else~/.local/share/bankml/forks) can be asked for. The file is looked for beside the startup model and inDIR./api/tagslists every pin, honestly. A pin bankML cannot play is listed with"native": falseand the reason. Asking for one answers 400 with that reason. Two examples:qwen3-0.6b-q8_0: Q8_0 weights, O3;qwen3-0.6b(Ollama's blob): Q4_K weights, O3.
One resident model. A request for another model drops the resident one, then verifies the new one before it loads it: the guard, then the sha256 pin, as at the start.
keep_aliveis read as Ollama reads it:"5m","1h","1h30m", or a number of seconds;0unloads after the answer;- a negative value keeps the model for good.
/api/*requests that name none get--keep-alive(default5m). An idle model is dropped when its time runs out, and its memory map with it. The OpenAI and llama-server endpoints name no keep-alive. They keep the model resident, as 0.3.0 did, and load the startup model again if nothing is resident.An empty prompt only loads or unloads. That is a
/api/generatewithout aprompt, or a/api/chatwithoutmessages. The answer'sdone_reasonisloadorunload.Streaming is on by default, as in Ollama: NDJSON (
application/x-ndjson), one object per line, then a finaldone: trueobject. That object carries:done_reason(stoporlength);total_duration,load_duration,prompt_eval_durationandeval_duration, in nanoseconds;prompt_eval_count: the whole prompt;eval_count: generated tokens, the end-of-turn token included, as llama-server counts it;bankml_cache_n: prompt tokens reused from the cache;bankml_receipt, as on/v1/chat/completions.
Options.
- Honoured:
temperature,top_k,top_p,min_p,seed,num_predictandstop(also on/v1).top_kis reproduced from 1 to 128; 0 or a larger value is refused, because llama.cpp sorts larger sets another way. num_ctx: accepted up to the served--ctx, refused above it. Under it (0.3.5) a conversation that does not fit loses its oldest messages first, as Ollama's do (the last message and the system messages stay); a prompt that still does not fit is refused with its size.- Penalties (O2):
repeat_penalty,repeat_last_n,presence_penaltyandfrequency_penaltyare honoured as llama-server b11192 applies them, token for token on its oracle: the window holds the lastrepeat_last_ntokens (default 64) including the prompt's, a repeated token's logit is divided by the repeat penalty (multiplied when not positive), thencount × frequency + presenceis taken off.repeat_last_nbelow 0 and a repeat penalty of 0 or less are refused with llama-server's message. Also on/v1and as a ModelfilePARAMETER. - The rest of the chain (0.3.7, next release; modules/sampler.md):
typical_pon/api/*and/v1; on/v1(not Ollama options) alsotop_n_sigma,xtc_probability,xtc_threshold,dynatemp_range,dynatemp_exponent,dry_multiplier,dry_base,dry_allowed_length,dry_penalty_last_nanddry_sequence_breakers— each as llama-server b11192 applies it, token for token on its oracle, in its default order. Clamped as it clamps (top_p,min_p, the XTC fields to [0, 1]); refused as it refuses (a negative DRY window or allowed length, empty breakers).mirostatand a customsamplersorder are refused: they are not reproduced. - Ignored, because they do not change the answer bankML gives:
num_thread,num_batch,num_gpu,use_mmapand the other resource options. - Refused: any other option.
- Honoured:
format: "json"is JSON mode (0.3.3). It is answered by the same grammar llama-server b11192 uses forresponse_format: {"type": "json_object"}(below), so/api/chatwithformat: "json"gives llama-server's tokens.format: {}andformat: {"type": "object"}are the same request. Since 0.3.5 any other schema object is answered with the grammar llama-server builds for it (below).formatwithraw: trueis refused (the grammar begins with the template's generation prompt, which a raw prompt does not have).Refused, with the reason:
- a
formatschema llama.cpp b11192 refuses (with its message); toolsandtool_calls(O6);images;suffix,template,context;think: true(the template is rendered with thinking off;think: falseis accepted);/api/embed(O7),/api/pull,/api/push, andDELETE /api/deleteof a pinned file (derived models: below).
- a
raw: truetokenizespromptas given, with special tokens and no template.systemis ignored then, as in Ollama.The loopback rules still apply. POST bodies need
Content-Type: application/json(curl -dalone sends a form type and gets 415), andHostmust be a loopback name.
Derived models and conversion: bankml create, bankml convert (O5, 0.3.5)
ollama create, natively. A derived model is a layer (system prompt, parameters, stop strings, example messages,
licence) over a pinned base, written as <registry>/<name>.MODEL.json. Weights are never copied. Loading it
verifies the base exactly as a pinned model is verified. The details are in
OLLAMA.md. mindX's persona layer, made this way, answers token-identically
to llama-server given the same persona (testing/persona_oracle.py).
# mindX's flow, without Ollama: the merged safetensors directory → a pinned GGUF → the persona layer
cat > Modelfile <<'MF'
FROM /home/mindx/mindX/data/godel/ascend/gen39/out/runs/<recipe>/ollama_push/merged
MF
bankml create mindx-gen39 -f Modelfile # converts (byte-identical to llama.cpp) → mindx-gen39-F16.gguf + its FORK.json
cat > Modelfile.persona <<'MF'
FROM mindx-gen39
SYSTEM """You are mindX generation 39, …"""
PARAMETER stop <|im_end|>
PARAMETER num_ctx 2048
MF
bankml create mindx-gen39 -f Modelfile.persona # re-created in place: the layer over the same pinned base
# the same over HTTP (bankml serve … --native --registry)
curl -s $B/api/create -H "$J" -d '{"model": "mindx-persona", "from": "mindx-gen39", "system": "You are mindX.", "parameters": {"temperature": 0.7}}'
curl -s $B/api/show -H "$J" -d '{"model": "mindx-persona"}' # the Modelfile back, details.parent_model
curl -s $B/api/copy -H "$J" -d '{"source": "mindx-persona", "destination": "mindx-persona-b"}'
curl -s -X DELETE $B/api/delete -d '{"model": "mindx-persona-b"}'
# conversion alone, with the FORK.json that pins it
bankml convert DIR -o mindx-gen39-F16.gguf --fork mindx-gen39-F16.gguf.FORK.json --source "PYTHAI/mindXascension@4bd31b9d weights/gen39"
- Modelfile. Ollama's grammar:
"""…"""blocks,#comments, case-insensitive instructions.- Taken:
FROM,SYSTEM,PARAMETER(temperature,top_k,top_p,min_p,seed,num_ctx,num_predict,stop),MESSAGE,LICENSE,REQUIRES. TEMPLATE: only when it is the base's own.- Refused, with the reason:
ADAPTER(merge the LoRA first), other parameters,quantize, uploadedfiles. - A relative
FROM ./xis relative to the Modelfile.
- Taken:
- Where.
--registry DIR(default$BANKML_FORKS, else~/.local/share/bankml/forks).FROMa name looks there, and in--models DIR.FROMa GGUF needs its FORK.json there.serve --native --registrylists the derived models and answers/api/create. Without--registryit refuses, because it has nowhere to write.
- A derived model's name wins over the suffix-less alias of a pin (
mindx-gen39is the layer,mindx-gen39-f16the pin), and/api/psnames it, with its digest and parent, as Ollama names the model that loaded its runner. A layer'sPARAMETER num_ctxfits the conversation on/apiand/v1as the request's option does. bankml convert. The Llama architecture (SmolLM2,mindx-genN) to GGUF F16, byte-identical to llama.cpp b11192'sconvert_hf_to_gguf.py --outtype f16:e9aba089…for SmolLM2-135M-Instruct,6b64c748…for gen39.- llama.cpp names the model from the directory's name. A directory called
mergedgivesgeneral.nameMerged, as llama.cpp's would.--model-namesets it, as llama.cpp's flag does. - A
README.mdis refused, because llama.cpp would copy the card's metadata in. - A tokenizer that llama.cpp has not been seen to name is refused (
testing/convert_oracle.py --chkhsh DIRmeasures one).
- llama.cpp names the model from the directory's name. A directory called
6b. The C API
Since 0.3.2 bankML is also a C library, for a program that embeds it rather than calling a server. The full reference, with the ownership and threading rules and the log formatter, is CAPI.md.
cargo build --release -p bankml-capi # target/release/libbankml.so, libbankml.a
cc app.c -I capi/include -L target/release -lbankml -Wl,-rpath,$PWD/target/release -o app
#include <bankml.h>
static int piece(const char *p, size_t n, void *u) { fwrite(p, 1, n, stdout); return 1; }
int main(void) {
char *err = NULL, *result = NULL;
bankml_t *h = bankml_open(".models/Bonsai-8B-Q1_0.gguf", "Bonsai-8B-Q1_0.gguf.FORK.json", 2048, &err);
if (!h) { fprintf(stderr, "%s\n", err); bankml_free(err); return 2; } /* refused: the reason, as serve gives it */
bankml_chat(h, "{\"messages\": [{\"role\": \"user\", \"content\": \"Say hello.\"}], \"max_tokens\": 32}", piece, NULL, &result);
puts(result); /* the /v1/chat/completions object, with usage and bankml_receipt */
bankml_free(result);
bankml_close(h);
}
bankml_open runs the same verification as bankml serve, and bankml_chat gives the answer and receipt that
serve --native gives (the gate's C API oracle compares them, turn by turn). bankml_set_log and the printf-style
bankml_log send the library's messages to your own callback.
6c. The bankML console: bankML as itself (0.3.7)
In the next release (not yet tagged). The design is in modules/console.md and modules/metrics.md.
python3 sAGI/console.py # http://127.0.0.1:7875 — talks to bankml serve at BANKML_SERVE_LISTEN (127.0.0.1:18093)
--port moves it; any --host other than loopback is refused, because the console can restart the engine.
./install.sh stop does not stop it (it stops 7874, 7873, 18093 and 18092); stop it with Ctrl-C.
A simple page with four tabs:
- Interaction: ask; the answer streams, and under it the receipt —
✓ answer = receiptwhen the sha256 of the text you received is the receipt's, the model, the tokens, the time to first token, the generation speed. bankML answers as itself (sAGI/personas/bankml.persona): each question carries a SELF block measured at that moment from/bankml/usageand/bankml/metrics, so "how many tokens have you used, how fast, how much power" is answered from measurement, and "not measured" where it was not. - Admin: three sliders — CPU threads, the RAM budget (weights plus KV cache, which sets the context) and the GPU limit (0 = off) — applied as one verified restart of the engine, rolled back if it fails; and D3 charts of what bankML measures: tokens per second, TTFT, CPU, memory, GPU busy against its limit, power.
- Logging: every exchange with its timings and receipt, refusals included, and the engine's log.
- Infotags: the metadata an iNFT publication of the session carries (ERC-721 shape): the engine, the model and its sha256, the persona and its doctrine root, the token totals, and an RFC 6962 Merkle root over the exchanges with their CID. The exchanges stay on this machine; the root lets a holder check any one of them. Download it as JSON.
Power appears only after ./install.sh power (opt-in, the one step that uses sudo; see
install.md §2); until then it reads "not measured".
7. Let others watch (view mode, on the LAN)
./install.sh start --view # or: python3 sAGI/view.py --host 0.0.0.0 --port 7874
Others open http://<this computer's LAN address>:7874 (ip -4 addr shows it). The page is read-only: the live
testing log, this machine's load, the verified model, the release records, CI, Savante's office and ledger, and her
recordings (see playback.md). It never shows the content of .history or .memory, only their
commitments (§8a). It is the Python standard library, not Gradio, because the Gradio installed for the chat has
path-traversal bugs and must stay on this computer.
8. The files Savante keeps: .history, .memory, .prompt
Nothing is ever written into the canon (~/cryptoAGI/savante). What the UI writes lives in BANKML_UI_STATE (default
~/.local/share/bankml/savante/):
savante.history: one JSON object per exchange (JSONL):{"ts": 1790633549.123, "sent_at": "2026-09-28T15:12:29.123-0700", "first_token_s": 113.02, "response_s": 125.11, "answered_at": "2026-09-28T15:14:34.233-0700", "session": "e2e-test", "user": "In one sentence: …", "assistant": "Savante knows …", "assistant_raw": "Savante knows …", "shown": "… (with the footer)", "prompt": "persona · system_prompt (canon, ledgered)", "prompt_provenance": "persona.system_prompt · persona sha256 d96556b11989… (ledgered)", "receipt": {"bankml": "0.0.7", "model_sha256": "284a335a…", "prompt_tokens": 316, "completion_tokens": 28, "ttft_ms": 113000, "wall_ms": 125100, "response_sha256": "80010408…"}}sent_atis the press of Send;first_token_sandresponse_sare measured from it. The receipt'sttft_ms/wall_msare measured by bankml serve from when it forwarded the request.assistant_raw(0.1.7+) is the text exactly as the model wrote it, which is what the receipt hashes;assistantis the same text with any<think>block and surrounding whitespace removed. Each record is appended as one locked, fsynced write; a line damaged by a crash is skipped (and counted), never allowed to hide the rest. The history is plain text: back it up, grep it, or delete it.savante.memory: one JSON object per note:{"ts", "at", "text", "sha256", "source": {"kind": "typed" | "response", "session", "sent_at", "response_sha256"}}.Savante.prompt: the Space template's prompt, cached the first time it is chosen.
The history the model sees.
- Without a context limit: the last 12–17 exchanges, each cut to 4,000 characters. The window starts at 12 exchanges and moves in steps of six, so between moves each prompt is the previous prompt plus one exchange, and the engine's prompt cache reuses all of it.
- Within the engine's context. The window must also fit the context the engine actually runs with (
n_ctxfrom/props), counted by its own tokenizer.- The start stays where it was last turn while that still fits, so the cache is reused.
- When it no longer fits, the start jumps so that the newest half of what fits is kept, leaving room for the next turns. When only two or three exchanges fit, all of them are kept.
- The answer's trail (Admin → the last answer) then says so: "history: 4 of 13 exchanges fit the engine's 2048-token context — raise the RAM budget on the Admin tab for more".
- Exchanges older than the 12–17 window are dropped without a note, as they always were.
- If the system prompt and the question alone do not fit, the answer is refused with the numbers.
- On a 2048-token context the persona prompt plus
.memoryleaves room for only a few exchanges; 4096 tokens (about 0.3 GB more RAM for the 8B model) gives the window room to stay put for several turns. After an engine restart, the first question restores the saved KV of the system prompt (a 51 MB file per model, context and system prompt, in~/.local/share/bankml/savante/slots/), so it skips the system prompt's prefill: on this laptop 132 s became 15 s, with the same answer. (A window that slid by one exchange per turn, as before 0.1.8, changed the text right after the system prompt and forced the whole history to be re-read every turn; over 60 turns the stable window moves 8 times instead of 48.) The chat's footer (clock, receipt) is never sent to the model.
8a. Proof of data without the data
.history and .memory stay on this computer. What can leave it is a commitment:
| part | what |
|---|---|
| leaf | sha256(0x00 ‖ one JSONL line, exactly as written, without its newline) |
| Merkle root | RFC 6962 / RFC 9162: node = sha256(0x01 ‖ left ‖ right), the tree split at the largest power of two below n (an odd node is promoted, never paired with itself). Scheme rfc6962-sha256; 0.1.6 and earlier used an unprefixed tree, so their roots differ |
| file sha256 / CIDv1 | of the whole file; CIDv1 raw (0x55), sha2-256, base32 lower: the construction Savante's iNFT ledger uses |
🔏 proof (Responses tab) produces, for one exchange, {"commitment": …, "proof": {"scheme", "record", "records", "leaf", "path": [hash, …], "merkle_root"}, "verifies": true}. Give someone the exchange's line and the proof: they
verify it as RFC 9162 §2.1.3.2 describes (the index and the tree size decide at each step which side the sibling is
on) and compare the result with the root and record count you published; the pair is the commitment. That proves the exchange is in your history, unchanged,
and reveals nothing else. The same works for any data kept in a browser's local storage: keep the data, publish the
hash or CID. Metrics and the view page show the commitments; the view page never serves a line.
To check a proof by hand, in Python:
import hashlib, json
line = b'...the exact JSONL line...'
p = json.load(open("proof.json"))["proof"]
h = hashlib.sha256(line).hexdigest()
for s in p["path"]:
a, b = (h, s["hash"]) if s["side"] == "right" else (s["hash"], h)
h = hashlib.sha256(bytes.fromhex(a) + bytes.fromhex(b)).hexdigest()
print(h == p["merkle_root"])
8b. Custom agents from the Savante template
Agents tab. Savante's canon is the template and is never written. Derive a new agent creates, in
~/.local/share/bankml/agents/<name>/ (BANKML_AGENTS):
| file | what |
|---|---|
<name>.persona |
the persona (mindX .persona v1 shape): a new name and identity, Savante's BDI/skills/safety as a starting point; the token bindings start empty (Savante's are never copied) |
<name>.prompt |
the system prompt, as text: the .prompt the chat uses |
<name>.agentcard.json |
EIP-721 metadata ∪ ERC-8004 registration-v1; status not_yet_minted; derived_from names Savante's persona sha256 and doctrine root |
<name>.commitments.json |
the ledger: sha256 + CIDv1 of each file, and the doctrine root (keccak256 over the same 15 JSON pointers as Savante's ledger) |
<name>.history, <name>.memory |
the agent's own conversations and notes |
Use this agent switches the chat, .history, .memory, Responses and Metrics to it. Edit (custom agents
only) opens the .persona as JSON and the .prompt as text. Save and re-ledger refuses a persona that breaks
the binder's preflight (non-ASCII keys, floats, integers beyond ±2⁵³), lacks a doctrine clause, or changes one:
the 15 doctrine clauses (name, source, format, system prompt, mantra, oath, beliefs, skills, safety, embodiment, tool
allowlist …) are fixed when the agent is derived, and their root is recorded then and carried forward. To change them,
derive a new agent. Everything else (the .prompt file, voice examples, description, the aivatar) stays editable.
Otherwise it regenerates the card and ledger. An agent whose files do not match its ledger refuses to speak.
The keccak256 is written in pure Python. It reproduces Savante's published doctrine root
(0x92fe83eb…ae137d0) from her persona, and it equals pycryptodome's keccak for every input length from 0 to 400
bytes (tested).
8c. THOT bundles: the dataset an iNFT points to
Agents → THOT bundle builds <name>.thot.json to sagi.thot_manifest/1 (~/sagi/engine/THOT_MANIFEST.md):
- Facets. Core facets
persona(the ternary head, leaf 0) andprompt, plus declared custom facetsx-bankml.agentcard,x-bankml.historyandx-bankml.memory. The history and memory are committed by digest only: the manifest holds no conversation text and can be published. - Bundle root.
bundle_root= keccak256 overlabel ‖ 0x1f ‖ sha256_hex ‖ 0x1e, core facets in registry order, then custom facets by name. - Merkle root. A 64-leaf keccak256 tree, padded with
keccak256(b""), with unsorted pairs and no prefixes. - Identity.
thot:<sha256>, a CIDv1,thot-<cid>, andcontentRoot(keccak256), all over the canonical manifest without its identity block. - Lineage. Genesis is generation 1. Every facet change (a new chat, a new note, an edited prompt) makes the next
build generation n+1, with
parent= the previous manifest's CID. A re-bind with no change keeps the generation.relationsrecordsderived_fromSavante's bundle. - Rung. Each agent directory is its own git repository, so the rung's locator
(
localhost/<user>/<name>@<commit>) names real bytes. History and memory are git-ignored and listed aslocator_lacks.
The builder is checked against the spec's own test vectors: Savante @1fcca89, Jaimla @8b57ccf and LuvAI @0c1eef7 give the same bundle roots, Merkle roots and identity CIDs. Savante's current generation-9 manifest verifies with no findings.
8d. PostgreSQL: publish an agent, load one back
Agents → PostgreSQL. The connection is BANKML_PG_DSN (default dbname=bankml, the local socket, your role).
One-time setup, which needs sudo:
sudo -u postgres psql -c "CREATE ROLE $USER LOGIN CREATEDB" -c "CREATE DATABASE bankml OWNER $USER"
sudo -u postgres psql -d bankml -c "CREATE EXTENSION IF NOT EXISTS vector"
- Publish the agent in use. This rebuilds and verifies its THOT bundle, then upserts
bankml_agents: persona (the exact bytes, plus a jsonb copy for queries), prompt, card, ledger, manifest, THOT CID, contentRoot and generation.- Off by default:
.historyand.memorylines go in (bankml_exchanges,bankml_memory, exact line text, sequence and sha256) only when include .history and .memory lines is ticked. Otherwise the database holds only their commitments, inside the manifest.
- Off by default:
- Load. Rebuilds a published agent as a local one. It is refused unless the stored manifest's identity recomputes and every restored file (persona, prompt, and history and memory when included) re-hashes to the manifest's digest. A tampered row cannot be loaded.
- Vectors.
bankml_exchanges.embeddingisvector(1024)(bge-m3's width, as mindX uses), indexed with DiskANN when pgvectorscale is installed and HNSW (pgvector) otherwise. It is filled with bge-m3 vectors when private lines are published and bge-m3 can run (see embedding.md); otherwise it stays empty. - Safety. Values travel to psql as COPY data into a temporary table; no value is ever part of SQL text.
testing/test_connectors.pyruns every step against a throwaway cluster (initdb in a temp dir, pgvector, its own socket), including a tampered prompt, a tampered history line and an SQL-injection string, all handled.
8e. iNFT: mint an agent, load one from a token
Agents → iNFT, for the agent in use. Savante herself is not minted from here: her verdict on minting is DEFER.
The contract is the house ERC-7857 iNFT_7857 (DeltaVerse iNFT4), and an open (unsealed) agent is minted with:
mintOpenAgent(address to, bytes32 contentRoot, string storageURI, bytes32 metadataRoot,
uint256 dimensions, uint8 parallelUnits, string tokenURI) — MINTER_ROLE only
| argument | from the agent |
|---|---|
contentRoot |
its THOT manifest's identity.contentRoot (keccak256 of the canonical manifest). The contract accepts a content root once, ever |
metadataRoot |
keccak256 of the agent card's canonical bytes |
storageURI, tokenURI |
local://thot/<cid> and local://card/<cid> until the bundle and card are stored somewhere (the rung stays referenced) |
dimensions, parallelUnits |
one of 8 … 1048576 (768 by default), and 1 |
- Plan and simulate builds the calldata and runs it as an
eth_callfrom the minter. You get the token id it would receive, or the decoded revert (AccessControlUnauthorizedAccount,ContentRootAlreadyMinted,InvalidDimension, …). - Unsigned transaction gives the transaction (chain id, to, data, gas estimate) and the same call as a
cast send … --account <your keystore>line. You sign it, in your own wallet: bankml never signs on a public chain. - Mint on the local devnet sends only when the chain id is 31337 (anvil). On any other chain it refuses.
- Load the agent from this token reads
getPayload,tokenURI,ownerOfandopenMint, finds the agent whose THOT generation has thatcontentRoot(every manifest is archived by CID in<agent>/thot/), verifies its bundle, and walks the THOT parent links from the current generation back to the minted one. The genesis stays attached to the token forever; later generations are proven to descend from it. - Start a local devnet runs anvil (127.0.0.1:8545, chain 31337), deploys
iNFT_7857from its compiled artifact (DeltaVerse/deploy/iNFT4/out) with account 0 as admin and minter, and fills in the fields. The whole path then runs on this computer.
Where it stands: iNFT_7857 is not deployed on any public chain, its audit (iNFT4/audit.json) is not cleared,
and a real mint needs a wallet holding MINTER_ROLE. testing/test_chain.py runs the whole path on a throwaway
anvil: deploy, simulate (and the refusal without the role), unsigned transaction, mint, read-back, the double-mint
refusal, load, two more generations and load again, and the refusal to send off a devnet.
9. Receipts, and how to check an answer
Every answer from bankml serve carries:
| field | meaning |
|---|---|
bankml |
the version that served it |
engine |
what computed it: llama.cpp b11192 llama-server (loopback), behind bankml P0, or with --native bankML <version> native: its own forward pass, … (and the GPU when one took part) |
model_sha256, guard |
the file that was verified and pinned, and the guard's verdict |
prompt_tokens, completion_tokens |
the engine's own counts |
ttft_ms, wall_ms |
time to first token and total, at the gateway (ttft_ms is null on a non-streamed answer in 0.3.6; the next release sets it there too) |
response_sha256 |
sha256 of the answer text exactly as the model produced it (choices[0], UTF-8; an unpaired \u surrogate counts as U+FFFD) |
request_sha256 |
(0.1.7+) sha256 of the request body bankml forwarded: which prompt this answer is to |
signed |
always false: a receipt is not a signature (below) |
The UI recomputes the answer's sha256 and shows ✓ when it matches. A mismatch shows (≠ received!). To check an
answer yourself, hash assistant_raw from .history (older records: assistant) and compare it with
receipt.response_sha256.
What a receipt proves, and what it does not. It binds an answer to the request it answered and to the sha256 of the model file bankml verified before serving, and it catches any change to the text afterwards. It is not signed: anyone can write a receipt, so it shows integrity between you and your own bankml serve, not to a third party who does not trust your machine. bankml serve also re-checks, before each answer, that the model file is the one it verified (device, inode, size and modification time) and that the engine still serves that file; if either changed, it refuses to answer rather than issue a receipt for weights it did not verify.
10. Savante's canon and the iNFT ledger
Savante is defined by her canon, ~/cryptoAGI/savante (github.com/cryptoAGI/savante), bound for an iNFT by the ledger
savante.commitments.json. The ledger commits by sha256 (and CIDv1) to:
- the persona;
- the charter;
- the sagi skill;
- six
sAGI.*facets; - the agent card;
- the image;
- the thot bundle;
- a keccak256 doctrine root over 15 clauses an owner may not edit.
At start the UI re-hashes every one of those files and shows the result in Integrity. If the persona does not
verify, the UI refuses to speak as Savante. The Verifier tab runs the full offline check. Nothing here mints: the
card's status is not_yet_minted, and the decision to mint belongs to the owner's signature. Point the UI at another
checkout with SAVANTE_CANON=/path.
11. Testing and the release gate
cargo test --release # offline
BANKML_GGML_LIB=/path/to/llama-b11192 testing/release_gate.sh # everything → testing/results/<version>.txt
testing/live.sh "title" <command…> # one step, shown live in view mode
The gate runs, in order:
- the build, the unit tests and the end-to-end CLI tests;
- clippy (
-D warnings); - the licence headers (
testing/spdx_check.py); - the Python suites: the guard, the UI data layer, the PostgreSQL connector (a throwaway cluster), the iNFT path (a throwaway anvil devnet), the model importer (with a real carrier on spare ports);
- the Rust and Python guards agreeing on every case; the C API's printf oracle (
bankml_logagainst libcsnprintf, from C); - every oracle: bankml's kernels must equal llama.cpp's compiled kernels, bit for bit, on every weight of the real models;
- the A/B speed tests, the prefill tile, the memory floor and the whole-token budgets;
- the conversation oracles:
serve --nativethrough/api/chatand/v1, and the C API'sbankml_chatagainstserve --native(ternary) and llama-server's record (1-bit); - the live oracles that start
serve --nativeand compare it with llama-server's answers: JSON mode and schemas, the persona layer, the penalties, and in the next release the samplers, the context limit, slots, sessions, logprobs and the q8_0 KV cache.
A speed counts only if every oracle passed on the same code. See testing/README.md, and oracles.md for what each oracle compares.
12. Troubleshooting
| symptom | cause | fix |
|---|---|---|
| chat says Cannot reach bankml serve | serve not running | ./install.sh status, then ./install.sh model (§3) |
| the installer stops at a step | that step's prerequisite is missing | it says which; fix it and run ./install.sh again (steps already done are quick) |
serve prints refuse: … != pinned |
the file is not the one the fork pinned | re-download; bankml sha256 FILE |
serve prints upstream … serves X, not the verified Y |
the llama-server on that port runs another file | restart it with the verified file, or use --spawn |
serve prints upstream … not healthy |
llama-server not up (a cold load can take a minute) | wait; check its log |
| the first answer takes minutes | CPU prefill of the ~300-token system prompt (≈3 tok/s on a laptop) | expected; later turns reuse the cache |
| refused: savante.persona does not verify | the canon was edited or is incomplete | git -C ~/cryptoAGI/savante status; run the Verifier |
| view mode says swap full | the machine is under memory pressure | close other work before measuring |
| view from another computer does not load | firewall, or bound to 127.0.0.1 | --host 0.0.0.0; allow TCP 7874 |
13. Reference: commands, ports, environment
| command | what |
|---|---|
./install.sh [step …] [--skip-tests] [--no-start] [--view] [--voice] |
install and start (§1); steps: check build engine python canon model voice start stop status, and (0.3.7, opt-in, sudo) power [--remove] (install.md §2) |
bankml guard FILE [--engine mainline|prism] [--json] |
header check: play / refuse / need_more |
bankml sha256 FILE |
the file's sha256 |
bankml pin FILE --fork FORK.json |
sha256 against the fork's record |
bankml verify FILE --fork FORK.json [--engine mainline|prism] [--json] |
guard, then pin |
bankml serve FILE --fork FORK.json [--upstream H:P | --spawn BIN] [--listen H:P] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR] |
the gate in front of llama-server (n-gram speculation opt-in; --slot-dir becomes llama-server's --slot-save-path) |
bankml chat-template MODEL.gguf < messages.json |
the prompt a conversation becomes, as llama.cpp's /apply-template (P3; byte-identical on its oracle; tools and assistant prefills refused) |
BANKML_GPU=off · BANKML_GPU_SHARE=0.3 |
the GPU worker (0.2.14): a verified card takes a calibrated share of every 1-bit matrix's rows; off disables it, a number overrides the share |
BANKML_GPU_LIMIT=0.8 |
the GPU limiter (0.3.7, next release): bankML's GPU buffers stay within this share of the card's heap (of the RAM they could use, on an integrated card) and the card rests so it is busy at most this share of the time; what does not fit or arrives while it rests runs on the CPU, the same bits. Each matrix shape also decides, measured in the pipeline, whether the card's share pays (on the Vega 3 most shapes choose the CPU) |
bankml serve FILE --fork FORK.json --native [--listen H:P] [--upstream H:P] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR] |
answers from bankML's own forward pass (0.3.0); since 0.3.1 also Ollama's API (§6a) and, with --registry, any pinned model by name, one resident at a time; token-identical to llama-server on its oracle; also serves llama-server's endpoints on the engine address, so Savante reaches it unchanged. --slot-dir (0.3.8, next release): llama-server's slot save/restore (§6). Savante's Admin tab engine setting (under advanced) chooses it: auto (bankML for the ternary files), native, llama.cpp |
bankml create NAME -f Modelfile [--registry DIR] [--models DIR] |
O5 (0.3.5): a derived model, a layer (SYSTEM, PARAMETER, stop, MESSAGE, LICENSE) over a pinned base, FROM a name, a pinned GGUF or a safetensors directory (converted); written as DIR/NAME.MODEL.json, no weights copied (§6a) |
bankml convert DIR -o OUT.gguf [--model-name N] [--fork F --source S] [--ignore-model-card] |
O5 (0.3.5): Llama safetensors → GGUF F16, byte-identical to llama.cpp b11192's convert_hf_to_gguf.py --outtype f16; --fork writes the pin with every input's sha256 |
bankml gpu --verify |
runs the bit-exact kernel oracle on every usable card (0.2.13); a card that fails is named and never used |
bankml gpu [--remote] |
every video card found (Vulkan, merged with /sys/class/drm) and which bankml will use; --remote adds the GPUs Hugging Face rents (22 NVIDIA flavors with card counts and prices; listed, never started); BANKML_GPU=off turns the component off, BANKML_GPU=0,2 picks cards (0.2.12; the GPU kernels are the next steps) |
bankml generate MODEL.gguf [--max N] [--sample [--temp T] [--top-k K] [--top-p P] [--min-p P] [--seed S]] < messages.json (or plain text) |
bankml's own forward pass, greedy, or with --sample llama-server's sampler chain (the model's defaults unless given; same seed, same tokens as llama-server, 0.2.11), streamed (P3, 0.2.7): token-identical to llama-server b11192 on its oracle for the 1-bit and ternary models, prompts of any length and contexts of any length (all three of ggml's CPU attention kernels, since 0.2.10); BANKML_LLAMA_THREADS (default 3) must equal the -t of the llama.cpp being matched, because its long-context decode kernel chunks by thread; BANKML_THREADS sets the threads. Ternary: 2.3–2.4 tokens/s against llama-server's 0.30; 1-bit: 1.8 against 2.8 |
bankml tokenize MODEL.gguf [--no-special] < text |
token ids, as llama.cpp's /tokenize (P3's tokenizer; token-identical on its oracle) |
bankml usage [PID …] |
memory, cores, and each process's resident memory and CPU % (bankml's psutil, from /proc); bankml serve answers the same at GET /bankml/usage, with (0.3.7, next release) package watts when RAPL is readable (./install.sh power), each GPU's busy %, VRAM and GTT, and the GPU limiter's state |
GET /bankml/metrics |
(0.3.7, next release) bankML's own measurements of its last 256 answers: TTFT, prompt and generation tokens per second, grammar time, energy and joules per token where measured; totals |
python3 sAGI/savante.py --mode interact [--port 7873] |
talk to Savante (loopback) |
python3 sAGI/view.py [--host 0.0.0.0] [--port 7874] |
the read-only page for the LAN |
python3 sAGI/console.py [--port 7875] |
(0.3.7, next release) the bankML console, loopback only (§6c) |
python3 sAGI/models.py list | catalog | search Q | import ID|URL|ollama:NAME:TAG | use FILE | first-run |
the model importer and carrier switch |
python3 sAGI/speak.py [--prune] |
render every voice clip, write both exports (--prune: drop unused clips) |
| port | service |
|---|---|
| 18092 | llama-server b11192 (upstream, loopback) |
| 18093 | bankml serve (loopback) |
| 7873 | interact mode (loopback) |
| 7874 | view mode (LAN) |
| 7875 | the bankML console (loopback; 0.3.7, next release) |
The variables below are the ones a user meets most. The complete list, with where each is read, is in install.md §6.
| variable | default | meaning |
|---|---|---|
SAVANTE_CANON |
~/cryptoAGI/savante |
Savante's canon (read-only) |
BANKML_SERVE |
http://127.0.0.1:18093 |
where the UI finds bankml serve |
BANKML_UI_STATE |
~/.local/share/bankml/savante |
.history and caches |
BANKML_REPO |
the checkout | where view mode reads testing/ |
BANKML_AGENTS |
~/.local/share/bankml/agents |
custom agents (one git repository each) |
BANKML_PG_DSN |
dbname=bankml |
PostgreSQL for publishing and loading agents |
RAGE_PATH |
~/mindX/mindx/godel/mindxtrain/hf/space_ui |
where the ragebar finds mindX's rage.py (built-in BM25 otherwise) |
BANKML_GGML_LIB |
— | llama.cpp b11192 release dir, for the oracles |
BANKML_THREADS |
pool: all cores; budgets: 1,2,3,4 (Q1_0), 1,3 (Q2_0) |
the thread pool's size, or a comma list of thread counts for the decode budgets |
BANKML_NO_SHANI |
unset | set to hash with the portable SHA-256 instead of the CPU's SHA extensions |
BANKML_OLLAMA_MODELS |
/usr/share/ollama/.ollama/models |
the local Ollama store the importer adopts from |
BANKML_PIPER, BANKML_PRONUNCIATION |
~/.local/share/bankml/piper, the house table |
Savante's voice engine and pronunciation table |
BANKML_ESPEAK, BANKML_ESPEAK_VOICES |
the DeltaVerse vendor dirs | the eSpeak fallback voice |
BANKML_OLLAMA, BANKML_EMBED_MODEL, BANKML_EMBED_KEEP_ALIVE, BANKML_EMBED_NEED_GB |
see embedding.md | the embedding model |
BANKML_MODELS |
.models in the checkout |
where imported models go |
BANKML_FORKS |
~/.local/share/bankml/forks |
FORK.json pins, one per imported model |
BANKML_LLAMA_SERVER |
~/sAGI/bonsai/llama-b11192/llama-server |
the engine the carrier spawns |
BANKML_BIN |
target/release/bankml |
the bankml binary the importer and switch call |
BANKML_FIRST_RUN |
1 |
0 stops interact from starting the Bonsai-8B carrier by itself |
BANKML_CTX, BANKML_THREADS_SERVE |
2048, 3 |
the engine's context and threads when the importer starts the carrier |
BANKML_SERVE_LISTEN, BANKML_UPSTREAM |
127.0.0.1:18093, 127.0.0.1:18092 |
the ports the switch manages |
BANKML_VOICE_DIR, BANKML_EXPORT_DIR |
sAGI/voice/cache, sAGI/voice/export |
voice clips and the two exports |
BANKML_VOICE_ASYNC |
1 |
0 stops the UI rendering missing clips in the background |
BANKML_CACHE_RAM |
8192 MiB, at most a quarter of the memory available at load | (0.3.8, next release) the native engine's host prompt cache in MiB, as llama-server's --cache-ram (0 off, -1 no limit) |
BANKML_CACHE_TYPE |
f16 |
(0.3.9, in progress) q8_0 keeps the native KV cache as llama.cpp's --cache-type-k/v q8_0 (§6) |