Download docs/install.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 70.4 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/install.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/install.md
-
curl -L -o install.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/install.md
Installing and configuring bankML: the complete reference
This is the full reference for installing bankML 0.3.6 and the Savante UI, and for every setting they accept. Settings from the next release (0.3.7, 0.3.8, and 0.3.9 in progress; not yet tagged) are marked with their version; CHANGELOG.md has the details. What each source module does is in modules/. For a first install, read usage.md Β§1βΒ§3; the README's Install and use gives the short version. This page assumes you have read one of them. It covers every step, flag, port, path and environment variable, what each one is for, and what it costs.
Everything here comes from the source: install.sh, bankML/main.rs, bankML/serve.rs, bankML/native.rs, the
kernels, sAGI/*.py and testing/. If a setting is not on this page, bankML does not read it.
- 1. Requirements in detail
- 2. install.sh in full
- 3. Building by hand
- 4. Models: where they live, how they are pinned
- 5. Running bankml serve, and every other command
- 6. Environment variables: the complete list
- 7. Performance and resource tuning
- 8. Running as a service
- 9. Security and exposure
- 10. Development and the release gate
- 11. Troubleshooting
1. Requirements in detail
Platform and toolchain
| what | detail |
|---|---|
| OS and architecture | Linux on x86_64. bankML itself builds wherever Rust 1.99 does. The llama.cpp engine the installer downloads (llama-b11192-bin-ubuntu-x64.tar.gz) is Linux x86_64 only. On any other platform, build llama.cpp b11192 yourself and set BANKML_LLAMA_SERVER to its llama-server |
| Rust | 1.99.0, pinned in rust-toolchain.toml (channel 1.99.0, component clippy), and rust-version = "1.99" in both crates. The C API (capi/) defines a C-variadic function (bankml_log), and Rust stabilized those in 1.99. With rustup installed, the first cargo run in the checkout fetches 1.99.0. The pin applies only inside the checkout, not to the rest of the machine |
| crates | none. Both packages in the workspace (bankml and bankml-capi) build offline, with no download |
| tools the installer checks | git, curl, tar, ss (iproute2), and sha256sum (or shasum) |
| Python | 3.10 or newer, for the Savante UI and the model importer |
| Gradio | gradio>=3.37,<4, plus numpy. The UI is written for Gradio 3, and its layout and scripts work around Gradio 3. Gradio 4 is not supported |
CPU features: chosen at run time
The release build has no target-cpu flag. Each kernel checks the CPU when it runs, so one binary works on any
x86_64 machine:
| feature | what uses it | without it |
|---|---|---|
AVX2 + FMA + F16C, all three (q1_0::has_avx2) |
the fast 1-bit, ternary and F16 kernels, attention and prefill | the portable kernels: the same bits, slower (the AVX2 1-bit path is 15.2Γ the portable C port, PERFORMANCE.md) |
| SHA-NI + SSE4.1 + SSSE3 | SHA-256 of the model file at every serve start, model load and switch (5.5Γ the portable rounds; 2.9 s for the 1.16 GB 8B model) |
the portable SHA-256, same digest; also chosen when BANKML_NO_SHANI is set |
./install.sh check reports AVX2. A machine without it is still correct, only slower.
Memory and disk, per model
The models bankML's own forward pass plays (serve --native, bankml generate, the C API):
| model | file | weights (mapped) | notes |
|---|---|---|---|
| Bonsai-8B, 1-bit | Bonsai-8B-Q1_0.gguf |
1.16 GB | the default carrier |
| Ternary-Bonsai-8B | Ternary-Bonsai-8B-Q2_0_g64.gguf |
2.31 GB | better answers; llama.cpp runs it about 5Γ slower on a CPU, bankML's kernel does not |
| Bonsai-1.7B, 1-bit | Bonsai-1.7B-Q1_0.gguf |
0.25 GB | the oracles' test model |
| SmolLM2-135M-Instruct | SmolLM2-135M-Instruct-F16.gguf |
0.27 GB | converted, Llama architecture |
| mindx-gen39 | mindx-gen39-F16.gguf |
0.27 GB | mindX's generation 39, converted |
Other catalogue models (Q8_0, Q4_K_M) are served only through llama-server, behind bankml serve without --native.
What else takes memory:
- The KV cache is f16 and grows with the context. Qwen3-8B uses 147,456 bytes per token
(
sAGI/models.py), so a 2048-token context holds about 0.30 GB and 4096 tokens about 0.6 GB. The other models need less; the guard reports each one'skv_f16_bytes_per_token. In the native engine,BANKML_CACHE_TYPE=q8_0(0.3.9, in progress) keeps it in 53 % of those bytes, as llama.cpp's q8_0 cache does (Β§6). - llama-server adds about 250 MB of compute buffers and runtime beside the weights and KV (
OVERHEADinsAGI/models.py, an order-of-magnitude measurement). - The whole stack (8B 1-bit model, engine, UI) runs in about 2 GB of free RAM. Plan for 3 GB of free disk.
The importer refuses an import when either limit is crossed, and gives the numbers:
- Disk: the file's size plus a 1.5 GB margin must be free.
- Memory: weights larger than 70 % of the machine's RAM are refused, because they would page to swap.
Optional components
| component | needed for | detail |
|---|---|---|
| a Vulkan GPU | the GPU worker (1-bit matrices only) | the loader libvulkan.so.1 is opened at run time, so no link-time dependency. Integrated and discrete cards with a compute queue are used; software renderers (Mesa llvmpipe) are refused. A card is used only after it passes the on-card oracle (bankml gpu --verify) |
ffmpeg |
Savante's neural voice | Piper's output is pitched and equalised with ffmpeg; without it the eSpeak stand-in is used |
Ollama with bge-m3 |
embeddings for the ragebar and PostgreSQL vectors | optional; without it, search is BM25 alone (embedding.md) |
PostgreSQL with pgvector, psql |
Agents β PostgreSQL | usage.md Β§8d |
Foundry's anvil and the iNFT4 build |
the local iNFT devnet | usage.md Β§8e |
| a C compiler | the C API's examples and its printf oracle | CAPI.md |
Python: the venv rule
The python step first tries the interpreter it was given (BANKML_PYTHON, else the one a previous run recorded,
else python3). If that interpreter imports both gradio and numpy, it is used as it is. Otherwise the step
creates ~/.local/share/bankml/venv with --system-site-packages and installs gradio>=3.37,<4 and numpy into it.
The system Python is never modified. The chosen interpreter is recorded as INSTALL_PYTHON in install.env, and
later runs use it. If python3 -m venv fails (on Debian and Ubuntu, python3-venv is missing), install that package,
or point BANKML_PYTHON at an interpreter that already has Gradio 3 and numpy.
2. install.sh in full
./install.sh [step β¦] [--skip-tests] [--no-start] [--view] [--voice] [-h|--help]
Nothing in it needs sudo. Unless steps are named, it runs check build engine python canon model, then voice if
--voice was given, then start unless --no-start was given. It stops at the first step that fails, and says why.
The steps
| step | what it does | what it writes |
|---|---|---|
check |
Rust β₯ 1.99 (cargo --version), python3 β₯ 3.10, git, curl, tar, ss, sha256sum/shasum; warns if the machine is not Linux/x86_64 and if AVX2 is missing; reports RAM (total, available) and free disk. Any missing item stops the run |
nothing |
build |
cargo build --release, then cargo test --release --quiet (unit and end-to-end, offline) unless --skip-tests; prints bankml version |
target/release/bankml |
engine |
uses an existing llama-server if one is set or found (see below); otherwise downloads llama.cpp b11192 (17 MB) and refuses it unless its sha256 is 34cf6fa5β¦881ec7. A tarball that is already there and matches is reused; one that does not match is deleted and downloaded again |
~/.local/share/bankml/llama-b11192-bin-ubuntu-x64.tar.gz, ~/.local/share/bankml/llama-b11192/, INSTALL_LLAMA_SERVER in install.env |
python |
the venv rule above | ~/.local/share/bankml/venv/ if needed, INSTALL_PYTHON in install.env |
canon |
clones github.com/cryptoAGI/savante to the canon path if it is absent; an existing canon is left untouched (it is read-only) |
the canon checkout |
model |
needs build and engine done. Runs sAGI/models.py first-run: imports Bonsai-8B if absent (checked against the publisher's sha256, guarded, pinned), then starts bankml serve on it with llama-server behind it. Then checks http://127.0.0.1:18093/bankml |
.models/Bonsai-8B-Q1_0.gguf, its FORK.json in the forks directory, logs/first-run.json, savante/carrier.log |
voice |
downloads Piper 2023.11.14-2 and the en_GB-cori-high voice (.onnx, .onnx.json, MODEL_CARD); each file is skipped if already there |
$BANKML_PIPER (default ~/.local/share/bankml/piper/) |
start |
starts interact mode (sAGI/savante.py --mode interact --port 7873), and with --view also view mode (sAGI/view.py --host 0.0.0.0 --port 7874), each detached (setsid nohup), waiting up to 60 s for its port. A port that already listens is left as it is |
logs/interact.log, logs/view.log |
stop |
for ports 7874, 7873, 18093 and 18092, finds the process listening there (by ss, never by matching command lines) and sends it SIGTERM |
nothing |
status |
which of the four ports listen; GET /bankml (verified, model, sha256 prefix); bankml version; the engine and Python in use |
~/.local/share/bankml/.status.json |
power (0.3.7, opt-in, sudo) |
lets a rapl group your user joins read the CPU package energy counter, so bankML can report watts and joules per token; explains the side channel (PLATYPUS, CVE-2020-8694) and asks first (BANKML_POWER_YES=1 to agree non-interactively); ./install.sh power --remove restores root-only |
/etc/udev/rules.d/60-bankml-rapl.rules |
start does not start bankml serve; model does. Interact mode also starts the first-run carrier in the
background when bankml serve does not answer, unless BANKML_FIRST_RUN=0. So after a reboot ./install.sh start
alone brings the whole stack back. Use ./install.sh model start to wait for the carrier in the foreground.
The options
| option | effect |
|---|---|
--skip-tests |
build skips cargo test --release |
--no-start |
the default run ends after model (or voice): everything installed, Savante not started. bankml serve is still started by model |
--view |
start also starts view mode on 0.0.0.0:7874. It works with named steps too: ./install.sh start --view |
--voice |
the default run includes voice. With named steps, name it instead: ./install.sh voice |
--space / --no-space |
0.3.9: let bankML's Hugging Face page (PYTHAI/bankml) talk to your bankml serve from your browser β its Your own bankML mode: your CPU, your verified model, a receipt the page checks. Serve starts with --allow-origin https://pythai-bankml.static.hf.space (that one page, nothing else); the choice is remembered in install.env (INSTALL_ALLOW_ORIGIN; BANKML_ALLOW_ORIGIN overrides it). ./install.sh start --space applies it at once (the running services stop first); --no-space forgets it. ./install.sh status says which page is allowed |
-h, --help |
prints the header of install.sh and exits |
Any other argument stops the installer with exit code 2.
What it reads
| variable | default | effect |
|---|---|---|
BANKML_DIR |
~/bankml |
only on the curl β¦ | bash route: where bankml is cloned (or the existing clone used) before install.sh re-runs from there |
BANKML_DATA |
~/.local/share/bankml |
the installer's data directory: install.env, logs/, the engine, the venv, Piper |
BANKML_LLAMA_SERVER |
β | the llama-server to use; skips the download |
INSTALL_LLAMA_SERVER |
from install.env |
what a previous engine step chose; used when BANKML_LLAMA_SERVER is unset |
BANKML_PYTHON |
β | the Python for the UI and importer |
INSTALL_PYTHON |
from install.env |
what a previous python step chose; used when BANKML_PYTHON is unset |
SAVANTE_CANON |
~/cryptoAGI/savante; ~/savante if that is the only one present |
the canon to clone into or use |
BANKML_PIPER |
$BANKML_DATA/piper |
where voice installs Piper |
How the engine is found: BANKML_LLAMA_SERVER, else INSTALL_LLAMA_SERVER, else the first executable of
$BANKML_DATA/llama-b11192/llama-server and ~/sAGI/bonsai/llama-b11192/llama-server; if none, engine downloads
it. The Python is BANKML_PYTHON, else INSTALL_PYTHON, else python3. (check tests the version of python3 on
the PATH, whatever the UI will use.)
LLAMA_TAG (b11192) and LLAMA_TGZ (llama-b11192-bin-ubuntu-x64.tar.gz) are constants in the script, together
with the archive's sha256. They are not settings: bankML is verified against that one release.
The installer passes its choices on. The importer and both UIs run with BANKML_LLAMA_SERVER (the chosen engine)
and BANKML_BIN (target/release/bankml); the UIs also get SAVANTE_CANON. When you start sAGI/savante.py by
hand, set the same variables yourself (Β§11).
What it leaves on disk
~/.local/share/bankml/ ($BANKML_DATA)
βββ install.env INSTALL_LLAMA_SERVER=β¦, INSTALL_PYTHON=β¦ (shell-quoted; sourced by later runs)
βββ logs/ first-run.json, interact.log, view.log
βββ llama-b11192-bin-ubuntu-x64.tar.gz, llama-b11192/ the engine
βββ venv/ only if your Python lacked Gradio or numpy
βββ piper/ only after `voice`
βββ .status.json the last `status` answer
βββ forks/ FORK.json pins ($BANKML_FORKS)
βββ savante/ the UI's state ($BANKML_UI_STATE): savante.history, savante.memory,
β Savante.prompt, savante.aivatar, carrier.log, resources.json, slots/
βββ agents/ custom agents, one git repository each ($BANKML_AGENTS)
<checkout>/target/release/bankml the binary
<checkout>/.models/ imported models ($BANKML_MODELS)
~/cryptoAGI/savante/ the canon ($SAVANTE_CANON), read-only
forks/, savante/ and agents/ are written by the importer and the UI, not by the installer. They sit under
~/.local/share/bankml/ by their own defaults, so setting BANKML_DATA does not move them; set BANKML_FORKS,
BANKML_UI_STATE and BANKML_AGENTS as well.
Re-running, and single steps
Every step checks before it acts, so a second run is quick and changes nothing that is already right. The tarball is
re-verified, not re-downloaded; the model is imported only if absent; a running UI is not restarted. Run any steps
alone, in any order: ./install.sh build, ./install.sh engine python, ./install.sh model start --view.
To rebuild after a git pull, run ./install.sh build, then ./install.sh stop and ./install.sh model start.
The piped route
curl -fsSL https://raw.githubusercontent.com/cryptoAGI/bankml/main/install.sh | bash
curl -fsSL https://raw.githubusercontent.com/cryptoAGI/bankml/main/install.sh | BANKML_DIR=/opt/src/bankml bash
Outside a checkout, the script clones github.com/cryptoAGI/bankml into $BANKML_DIR (default ~/bankml), or uses
the clone already there, and runs that copy's install.sh with the same arguments. Pass arguments through bash:
curl β¦ | bash -s -- --no-start.
3. Building by hand
cargo build --release # target/release/bankml (the root package is the default member)
cargo test --release # unit tests and testing/cli.rs, offline
cargo build --release -p bankml-capi # target/release/libbankml.so and libbankml.a
cargo test --release -p bankml-capi # the C API's unit tests
cargo clippy --release --workspace --all-targets -- -D warnings
- The workspace. It has two members. The root package
bankmlbuilds the binary and the library (bankML/bankml.rs).capi/(bankml-capi) builds the C librarylibbankml, ascdylibandstaticlib, over the root crate by path. A plaincargo buildbuilds only the root package. Headers:capi/include/bankml.h. Linking, ownership and threading rules are in CAPI.md. - The release profile uses
lto = true,codegen-units = 1andpanic = "abort", for one small binary. Do not add-C target-cpu=native: the kernels already choose AVX2 at run time (Β§1), and a binary built for one CPU may not start on another. --lockedis what the release gate uses (cargo build --release --locked). With no external crates,Cargo.locklists only the two workspace packages.- Installing the binary elsewhere. The binary is self-contained. Copy
target/release/bankmlwhere you need it (Β§8 keeps the previous version beside it for rollback). The UI and importer find it throughBANKML_BIN.
4. Models: where they live, how they are pinned
Directories
| what | default | set with |
|---|---|---|
model files (*.gguf) |
<checkout>/.models |
BANKML_MODELS (BANKML_REPO moves the checkout the default is relative to) |
pins (*.FORK.json) |
~/.local/share/bankml/forks |
BANKML_FORKS |
derived models (NAME.MODEL.json) |
the forks directory | bankml create --registry DIR |
A pin is <file>.FORK.json: a JSON object with provenance (source_repo, source_url, source_revision,
licence_tag, imported_at_utc, pinned_from) and files: [{"path": "<file>.gguf", "bytes": β¦, "sha256": "β¦"}].
bankml verify, bankml serve and bankml_open pass a file only if its sha256 equals the files entry for its
name, and only after the guard says play.
The importer: sAGI/models.py
python3 sAGI/models.py list # every GGUF in .models: pinned or UNPINNED, size, source, licence; the carrier
python3 sAGI/models.py catalog # the curated catalogue, and whether each entry fits this machine
python3 sAGI/models.py search QUERY # search ollama.com
python3 sAGI/models.py import ID # a catalogue id
python3 sAGI/models.py import https://huggingface.co/OWNER/REPO[/blob|resolve/REV/FILE]
python3 sAGI/models.py import ollama:NAME[:TAG]
python3 sAGI/models.py use FILE # make FILE the carrier (verified switch, rollback on failure)
python3 sAGI/models.py first-run # what the installer's `model` step runs
With no subcommand, list runs. The Models tab in Savante does the same things.
The rules an import must pass:
- Licence. The licence must be OSI-approved or a public-domain dedication. The accepted SPDX tags are apache-2.0, mit, bsd-2-clause, bsd-3-clause, bsd, isc, mpl-2.0, gpl-2.0, gpl-3.0, lgpl-2.1, lgpl-3.0, agpl-3.0, unlicense, cc0-1.0, zlib, artistic-2.0, epl-2.0, ecl-2.0, upl-1.0 and 0bsd. Anything else is refused outright, which excludes Gemma and Llama.
- A published sha256. The download is hashed as it streams and kept only if it equals the sha256 the publisher lists: a catalogue entry is re-checked against its repository at the pinned revision; a Hugging Face file against the repository's LFS sha256 at the resolved revision; an Ollama model against its registry layer digest. With no published sha256 there is nothing to pin to, and the import is refused.
- The guard.
bankml guard --jsonmust sayplay. A refused file is deleted. - The fit. The disk and memory limits of Β§1.
Then the importer writes the FORK.json.
The sources:
- Catalogue ids:
bonsai-8b(the default),ternary-bonsai-8b,bonsai-4b,bonsai-1.7b,qwen3-0.6b,qwen3-1.7b,qwen3-4b,qwen3-8b,smollm2-1.7b,smollm3-3b,granite-3.3-2b,qwen2.5-coder-1.5b,qwen2.5-coder-7bandqwen3.8-27b(16.5 GB, for a machine with 24 GB or more).catalogprints them with sizes and whether each fits. - Hugging Face: a repository URL with several GGUF files asks you to pick one.
- Ollama: a model the local Ollama already holds is adopted with no download. The importer links to the blob in
BANKML_OLLAMA_MODELS(default/usr/share/ollama/.ollama/models) after hashing it. Otherwise the model comes fromregistry.ollama.ai.
The only network calls are GETs to huggingface.co, ollama.com and registry.ollama.ai. Nothing is uploaded.
Pinning a file that is already here. use FILE pins an unpinned catalogue file first: it hashes the file and
compares the hash with the repository's published one. The two converted models, SmolLM2-135M-Instruct-F16 and
mindx-gen39-F16, are pinned against their recorded conversion. pin_converted hashes the GGUF and re-reads the
source safetensors' LFS sha256 at the recorded revision. There is no subcommand for this; call it as
usage.md Β§5 shows:
python3 -c "import sys; sys.path.insert(0, 'sAGI'); import models; print(models.adopt('SmolLM2-135M-Instruct-F16.gguf'))"
Switching the carrier (use). use refuses:
- while an answer is being written;
- a file that is not in
.models; - weights over 70 % of RAM;
- an embedding architecture (bert, nomic-bert, jina-bert-v2, xlm-roberta, modern-bert, neo-bert, t5encoder).
Otherwise it stops the carrier and starts bankml serve on the new file. The switch succeeds only when
/bankml answers verified with that file's sha256. If the new model fails, the previous one is restarted.
The registry, and derived models
bankml serve --native --registry [DIR] serves every GGUF pinned in DIR by name: the file's stem in lower case,
with :latest, the file name, and (when unique) the name without its weight-type suffix accepted as aliases. A
pinned file is looked for in two places: beside the startup model, and in DIR. A pin whose file is missing is
listed in /api/tags with the reason. So a registry can be a directory of FORK.json files with the GGUFs symlinked
in (Β§8). Derived models (bankml create, /api/create) are layers over a pinned base, written as
DIR/NAME.MODEL.json, with no weights copied. Their use is described in
usage.md Β§6a and OLLAMA.md.
5. Running bankml serve, and every other command
Two engines, one gate
bankml serve FILE --fork FORK.json [--upstream HOST:PORT | --spawn LLAMA_SERVER] [--listen HOST:PORT] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR] [--engine mainline|prism]
bankml serve FILE --fork FORK.json --native [--listen HOST:PORT] [--upstream HOST:PORT] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR] [--engine mainline|prism]
The two modes start the same way. serve checks the guard and the sha256 pin, refuses if the file changed while it
was being hashed, then either:
- llama.cpp mode (the default) puts llama-server b11192 behind the gate, and checks through
/propsthat llama-server serves the verified file; - native mode (
--native) answers from bankML's own forward pass, token-identical to llama-server on its oracle. It speaks the OpenAI API and Ollama's API, and it also answers llama-server's endpoints on the engine address, so Savante reaches it unchanged.
serve binds --listen only after verification. In llama.cpp mode it also waits until the upstream is healthy.
A client should wait for GET /bankml to answer, not for the port to open. install.sh, sAGI/models.py and the
oracles all do this.
Every serve flag
| flag | default | mode | effect |
|---|---|---|---|
FILE |
required | both | the GGUF to verify and serve |
--fork FORK.json |
required | both | the pin. A missing flag prints the usage and exits 1; an unreadable file exits 1 |
--engine mainline|prism |
mainline |
both | which llama.cpp the guard judges for. mainline refuses PrismML-fork-only types (PQ2_0, PTQ1_0); prism accepts them |
--listen HOST:PORT |
127.0.0.1:18093 |
both | the gateway's address |
--upstream HOST:PORT |
127.0.0.1:18092 |
both | llama.cpp mode: the llama-server to check, or where --spawn starts one. Native mode: a second listening address that answers llama-server's endpoints (/health, /props, /tokenize, /apply-template, /v1/chat/completions). Native mode cannot start while something else holds this port |
--spawn LLAMA_SERVER |
β | llama.cpp | serve launches this llama-server on the verified path and on --upstream, waits up to 900 s for /health, and stops it when serve stops (on SIGTERM, SIGINT, SIGHUP). Refused if something already answers on --upstream. Without --spawn, an existing upstream must be healthy within 10 s |
--threads N |
3 |
llama.cpp | passed as -t N to the spawned llama-server. Native mode ignores it; use BANKML_THREADS (Β§7) |
--ctx N |
4096 |
both | the context: -c N for the spawned llama-server, n_ctx for the native engine. A prompt of N tokens or more is refused (0.3.8: with llama-server's exceed_context_size_error 400 body on /v1); an Ollama num_ctx above it is refused |
--spec-ngram |
off | llama.cpp | --spec-type ngram-simple in the spawned engine: n-gram speculative decoding, exact at temperature 0, opt-in because it measured within noise |
--slot-dir DIR |
β | both | llama.cpp mode: --slot-save-path DIR for the spawned engine. Native mode (0.3.8): POST /slots/0?action=save|restore|erase with {"filename"}, as llama-server's slot API; the directory is created at start. Either way, a restart restores the system prompt instead of computing it again. Without it, the slot actions answer 501 |
--allow-origin ORIGIN |
β | both | 0.3.9: one web page (https://host[:port]) whose scripts may call this gateway from a browser β the bankML Space's page, for example (https://pythai-bankml.static.hf.space). bankML answers that origin's CORS preflight (with Chrome's Access-Control-Allow-Private-Network) and adds Access-Control-Allow-Origin to its answers, streamed ones included; any other origin gets no CORS headers and a 403 preflight. The loopback Host and JSON-POST rules are unchanged |
--native |
off | β | native mode |
--registry [DIR] |
off; bare: $BANKML_FORKS, else ~/.local/share/bankml/forks |
native | serve every model pinned in DIR by name, one resident at a time, each verified again when it loads; answer /api/create, /api/copy, /api/delete for derived models. A following argument that starts with -- is not taken as DIR |
--keep-alive DUR |
5m |
native | how long an /api/* request that names no keep_alive keeps the model resident: "5m", "1h30m", seconds, 0 (unload after the answer), negative (for good). OpenAI and llama-server endpoints keep the model resident |
A spawned llama-server is always started as:
llama-server -m <verified path> --host <upstream host> --port <upstream port> -t <threads> -c <ctx> -np 1 --jinja --reasoning off --no-webui [--spec-type ngram-simple] [--slot-save-path DIR]
When verification or startup fails, serve prints bankml serve: refuse: <reason> and exits 2.
The gateway's fixed limits
These are constants in serve.rs, not settings:
| limit | value | over it |
|---|---|---|
| connections at once | 32 | 503 bankml serve: too many connections |
| request line or header line | 16 KiB | 400 |
| header lines | 100 | 400 |
| request body | 16 MiB | 413 request too large |
| chunked request bodies | not accepted | 400: send Content-Length |
| read timeout per connection | 30 s | connection closed |
| write timeout per connection | 120 s | connection closed |
| upstream response body or chunk | 64 MiB | the request fails |
Host |
127.0.0.1, localhost or [::1], any port |
403 |
POST Content-Type |
must begin application/json |
415 |
Endpoints
| mode | endpoints |
|---|---|
| both | GET /bankml (what was verified, when, and what is upstream; in native mode also the resident model and every name), GET /bankml/usage (memory and CPU of serve and its engine, sampled at most once a second; 0.3.7 adds package watts, each GPU's busy %, VRAM and GTT, and the GPU limiter's state), GET /bankml/metrics (0.3.7: the native engine's own measurements of its last 256 answers) |
| llama.cpp | GET /health, /v1/models, /props (proxied); POST /v1/chat/completions (receipted) |
| native | GET /health, /props, /v1/models; POST /tokenize, /apply-template, /v1/chat/completions, /slots/0?action=β¦ (0.3.8); GET|HEAD /; Ollama's /api/version, /api/tags, /api/ps, /api/show, /api/chat, /api/generate, /api/create, /api/copy, DELETE /api/delete. /api/embed, /api/pull and /api/push answer 400 with the reason |
What each request may contain, and what is refused, is in usage.md Β§6, Β§6a and OLLAMA.md.
Ports
| port | service | bound to |
|---|---|---|
| 18092 | llama-server b11192, or native serve's engine address | loopback |
| 18093 | bankml serve |
loopback |
| 7873 | Savante, interact mode (Gradio) | loopback only; anything else is refused |
| 7874 | view mode (standard library HTTP) | 0.0.0.0 by default |
| 7875 | the bankML console, sAGI/console.py (0.3.7) |
loopback only; anything else is refused. ./install.sh stop does not stop it |
| 11434 | Ollama, for embeddings (BANKML_OLLAMA) |
loopback |
| 8545 | anvil, the local iNFT devnet, when started from the Agents tab | loopback |
The importer moves the carrier's two ports with BANKML_SERVE_LISTEN and BANKML_UPSTREAM, and the UI's view of it
with BANKML_SERVE. Keep the three consistent.
Choosing the engine for the carrier
The carrier is what sAGI/models.py starts. On Savante's Admin tab the engine setting (under advanced, saved
in savante/resources.json) chooses how:
auto(the default): native for the ternary (Q2_0) files, where bankML is about 8Γ llama-server with the same tokens; llama-server for everything else, since llama-server is still faster on the 1-bit files.native: alwaysbankml serve --native.llama.cpp: always llama-server behind the gate.
The two carriers are started this way:
- Native carrier:
bankml serve FILE --fork F --native --upstream $BANKML_UPSTREAM --listen $BANKML_SERVE_LISTEN --ctx N --slot-dir ~/.local/share/bankml/savante/slots, withBANKML_THREADSset to the chosen thread count andBANKML_GPU_LIMITto the saved GPU limit (BANKML_GPU=offwhen it is 0). - llama.cpp carrier:
bankml serve FILE --fork F --spawn $BANKML_LLAMA_SERVER --upstream β¦ --listen β¦ --threads N --ctx N [--spec-ngram] --slot-dir ~/.local/share/bankml/savante/slots.
Either is given 1800 s to answer verified.
The same tab sets threads (1 to the number of cores) and a RAM budget. The budget becomes a context: whatever
remains after the weights and the 250 MB overhead, divided by the KV bytes per token, rounded down to 256, between
512 and 32768. Apply restarts the carrier with the new settings and restores the previous ones if it will not
start. Without a saved budget, the context is BANKML_CTX (2048) and the threads BANKML_THREADS_SERVE (3).
Every other command
| command | options and defaults | exit codes |
|---|---|---|
bankml guard FILE |
--engine mainline|prism (default mainline), --json |
0 play Β· 2 refuse Β· 3 need_more (truncated) Β· 1 cannot read |
bankml sha256 FILE |
β | 0 Β· 1 |
bankml pin FILE --fork F |
β | 0 pinned Β· 2 refuse Β· 1 cannot read the fork |
bankml verify FILE --fork F |
--engine, --json |
0 play Β· 2 refuse Β· 3 truncated Β· 1 I/O |
bankml tokenize MODEL.gguf < text |
--no-special (do not parse special tokens) |
0 Β· 1 |
bankml chat-template MODEL.gguf < messages.json |
β | 0 Β· 2 |
bankml generate MODEL.gguf < messages.json|text |
--max N (256); --json (JSON mode); --sample with --temp, --top-k, --top-p, --min-p, --seed (each defaults to the model's GGUF value). Without --sample, greedy |
0 Β· 2 Β· 1 |
bankml create NAME -f Modelfile |
-f/--file required; --registry DIR (default $BANKML_FORKS, else ~/.local/share/bankml/forks); --models DIR (where FROM a name also looks) |
0 Β· 2 refuse Β· 1 usage |
bankml convert DIR -o OUT.gguf |
-o/--outfile required; --outtype f16 (the only one accepted); --model-name NAME; --fork F writes the pin, --source SRC names the source in it (default the directory's canonical path); --ignore-model-card |
0 Β· 2 refuse Β· 1 usage or write error |
bankml gpu |
--remote (also list the GPUs Hugging Face rents; listed, never started); --verify (run the on-card oracle on every selected card) |
0 Β· 2 a card failed --verify |
bankml usage [PID β¦] |
memory, cores, and each process's RSS and CPU % over 0.5 s (default: bankml itself) | 0 Β· 1 |
bankml version |
also --version, -V |
0 |
The Python entry points
| command | options |
|---|---|
python3 sAGI/savante.py |
--mode interact|view (default interact); --host (default 127.0.0.1; interact accepts only 127.0.0.1 or localhost); --port (default 7873). --mode view hands over to view.py, on 0.0.0.0:7874 unless a host and port are given |
python3 sAGI/view.py |
--host (default 0.0.0.0), --port (default 7874) |
python3 sAGI/console.py |
(0.3.7) --host (default 127.0.0.1; loopback only), --port (default 7875); reaches bankml serve at BANKML_SERVE_LISTEN |
python3 sAGI/models.py |
list, catalog, search Q, import ID|URL|ollama:NAME[:TAG], use FILE, first-run |
python3 sAGI/agents.py |
adopt PERSONA [--slug SLUG] (install an existing .persona as its own agent), list, verify SLUG |
python3 sAGI/speak.py |
renders every voice clip, then writes both exports; --shard K/N renders every N-th statement starting at K, for parallel renders, and skips the exports; --prune drops clips no current text uses |
6. Environment variables: the complete list
Every variable bankML or its tools read, where it is read, its default, and what it does. Where the installer sets a variable for the processes it starts, that is noted.
Engine and kernels (Rust: bankml, libbankml)
| variable | read by | default | effect |
|---|---|---|---|
BANKML_THREADS |
bankML/par.rs (Pool::from_env) |
all available cores | the thread pool of the native engine: serve --native, generate, the C API. Every thread count gives the same bits. sAGI/models.py sets it for a native carrier. In the decode-budget benchmarks (q1_0.rs, q2_0.rs) it is instead a comma list of counts to measure, default 1,2,3,4 (Q1_0) and 1,3 (Q2_0) |
BANKML_LLAMA_THREADS |
bankML/forward.rs |
3 |
the -t of the llama.cpp run being reproduced. In decode at 512 or more KV cells, llama.cpp's split attention kernel chunks by its thread count, so the bits depend on it. Keep it equal to the reference's -t for token identity. It does not change bankML's own thread count |
BANKML_NO_SHANI |
bankML/sha256.rs |
unset | set to any value to hash with the portable SHA-256 instead of the CPU's SHA extensions |
BANKML_FORKS |
bankML/main.rs |
~/.local/share/bankml/forks |
the registry directory for serve --registry without a DIR, and for bankml create without --registry |
HOME |
bankML/main.rs |
β | the base of the default forks directory |
BANKML_CACHE_TYPE |
bankML/native.rs (0.3.9) |
f16 |
q8_0 keeps the native engine's KV cache as llama.cpp's --cache-type-k q8_0 --cache-type-v q8_0 does, Hadamard rotation included: 53 % of the f16 cache's memory, the same tokens as llama-server so configured |
BANKML_CACHE_RAM |
bankML/native.rs (0.3.8) |
8192 MiB, at most a quarter of the memory available at load | the host prompt cache's limit in MiB, as llama-server's --cache-ram (0 off, -1 no limit) |
GPU (Rust)
| variable | read by | default | effect |
|---|---|---|---|
BANKML_GPU |
bankML/gpu/mod.rs |
unset: every usable card, discrete first, largest memory first | off, none or cpu turn the GPU component off. A comma list (0,2) picks cards by backend index (as bankml gpu lists them). The worker uses the first selected card |
BANKML_GPU_SHARE |
bankML/gpu/worker.rs |
calibrated at open: card rate Γ· combined rate (26β35 % on a Vega 3) | the fraction of each 1-bit matrix's rows the card computes, clamped to 0β0.95. A share under 0.02, set or calibrated, leaves the card unused |
BANKML_GPU_LIMIT |
bankML/gpu/worker.rs (0.3.7) |
0.8 |
the share of the card's memory (its heap; on an integrated card, the RAM it could use) and of its time bankML may take; clamped 0.05β1. The console's GPU slider and models.resources()["gpu_limit"] set it (0 there means BANKML_GPU=off) |
Installer (install.sh)
| variable | default | effect |
|---|---|---|
BANKML_DIR |
~/bankml |
piped route only: where to clone and run |
BANKML_DATA |
~/.local/share/bankml |
the installer's data directory |
BANKML_LLAMA_SERVER |
found or downloaded (Β§2) | the engine; passed on to the importer and UIs |
INSTALL_LLAMA_SERVER |
recorded in install.env |
the previous choice of engine |
BANKML_PYTHON |
INSTALL_PYTHON, else python3 |
the UI's Python |
INSTALL_PYTHON |
recorded in install.env |
the previous choice of Python |
SAVANTE_CANON |
~/cryptoAGI/savante (or ~/savante) |
the canon; passed on to the UIs |
BANKML_PIPER |
$BANKML_DATA/piper |
where voice installs Piper (also read by speak.py) |
BANKML_POWER_YES |
unset | (0.3.7) 1 answers yes to the power step's question, for a non-interactive run |
Model importer and carrier (sAGI/models.py)
| variable | default | effect |
|---|---|---|
BANKML_REPO |
the checkout (sAGI/..) |
the base of .models and target/release/bankml, and the carrier's working directory |
BANKML_MODELS |
$BANKML_REPO/.models |
where models are imported and looked for |
BANKML_FORKS |
~/.local/share/bankml/forks |
where pins are written and read |
BANKML_LLAMA_SERVER |
~/sAGI/bonsai/llama-b11192/llama-server |
the llama-server a llama.cpp carrier spawns. The installer sets it to its own choice for the processes it starts |
BANKML_BIN |
$BANKML_REPO/target/release/bankml |
the bankml binary the importer, guard and switch call |
BANKML_OLLAMA_MODELS |
/usr/share/ollama/.ollama/models |
the local Ollama store to adopt models from |
BANKML_SERVE_LISTEN |
127.0.0.1:18093 |
the carrier's --listen, and where the importer checks /bankml |
BANKML_UPSTREAM |
127.0.0.1:18092 |
the carrier's --upstream |
BANKML_CTX |
2048 |
the carrier's context when no RAM budget is saved |
BANKML_THREADS_SERVE |
3 |
the carrier's threads when none are saved (--threads, or BANKML_THREADS for a native carrier) |
BANKML_UI_STATE |
~/.local/share/bankml/savante |
where carrier.log, resources.json and slots/ are kept |
Savante UI and its state (sAGI/savante.py, view.py, agents.py, connectors.py)
| variable | read by | default | effect |
|---|---|---|---|
SAVANTE_CANON |
savante.py, speak.py |
~/cryptoAGI/savante; ~/savante if only that exists |
Savante's canon, read-only |
BANKML_SERVE |
savante.py |
http://127.0.0.1:18093 |
where the UI reaches bankml serve |
BANKML_UI_STATE |
savante.py, models.py |
~/.local/share/bankml/savante |
savante.history, savante.memory, Savante.prompt, savante.aivatar, the carrier's log, resources and slots |
BANKML_AGENTS |
agents.py |
~/.local/share/bankml/agents |
custom agents, one directory and git repository each |
BANKML_AGENT |
savante.py |
mindx |
the agent the UI starts with; Savante when that agent is not installed. Empty: Savante |
BANKML_FIRST_RUN |
savante.py |
1 |
0 stops interact mode from starting the Bonsai-8B carrier by itself when bankml serve does not answer |
BANKML_REPO |
savante.py, view.py |
the checkout | where the UI and view mode read testing/ (the live log, the release records) |
BANKML_PG_DSN |
connectors.py |
dbname=bankml |
the libpq connection string for publishing and loading agents (local socket, your role) |
RAGE_PATH |
savante.py |
~/mindX/mindx/godel/mindxtrain/hf/space_ui |
where the ragebar finds mindX's rage.py; built-in BM25 otherwise |
Voice (sAGI/speak.py)
| variable | default | effect |
|---|---|---|
BANKML_PIPER |
~/.local/share/bankml/piper |
Piper and the en_GB-cori-high voice. The neural voice needs piper/piper, the .onnx and ffmpeg |
BANKML_VOICE_DIR |
sAGI/voice/cache |
rendered clips |
BANKML_EXPORT_DIR |
sAGI/voice/export |
the two complete exports (Savante.opus, Savante-reading.opus) |
BANKML_VOICE_ASYNC |
1 |
0 stops the UI from rendering missing clips in the background |
BANKML_PRONUNCIATION |
~/mindX/data/config/pronunciation.json |
the pronunciation table; a built-in table is the fallback |
BANKML_ESPEAK |
~/DeltaVerse/vendor/espeak-ng/0.3.5-en |
the eSpeak NG stand-in voice's engine |
BANKML_ESPEAK_VOICES |
~/DeltaVerse/vendor/espeak-ng/mindx-voices |
its voice files |
Embeddings (sAGI/embed.py)
| variable | default | effect |
|---|---|---|
BANKML_OLLAMA |
http://127.0.0.1:11434 |
the Ollama server. Keep it local: history text is sent to it |
BANKML_EMBED_MODEL |
bge-m3 |
the embedding model; must give 1,024 dimensions |
BANKML_EMBED_KEEP_ALIVE |
60s |
how long Ollama keeps it loaded after a call |
BANKML_EMBED_NEED_GB |
1.3 |
free memory required before loading it; below that, search is BM25 alone |
Testing and oracles
| variable | read by | default | effect |
|---|---|---|---|
BANKML_GGML_LIB |
q1_0.rs, q2_0.rs tests, release_gate.sh, the testing/*_oracle.py recorders |
β | the llama.cpp b11192 release directory (libggml-base.so, libggml-cpu-haswell.so, libllama*.so). Without it, the gate skips every oracle, A/B and budget |
LLAMA_SRC |
grammar_oracle.py, content_oracle.py, schema_oracle.py, release_gate.sh |
β | a llama.cpp source checkout at tag b11192, for the oracles that compile against its headers |
BANKML_LLAMA_SRC |
convert_oracle.py, release_gate.sh |
upstream/llama.cpp |
the b11192 source tree whose convert_hf_to_gguf.py and gguf-py the conversion oracles run. In the gate it also enables oracle_name_heuristics |
BANKML_CONVERT_DIR |
convert.rs (oracle_convert_b11192) |
β | the safetensors directory to convert (its name matters: llama.cpp names the model from it) |
BANKML_CONVERT_ORACLE |
same | β | llama.cpp's GGUF of that directory, to compare byte for byte |
BANKML_CONVERT_OUT |
same | $TMPDIR/bankml-convert-oracle.gguf |
where bankML's conversion is written |
BANKML_CONVERT_KEEP |
same | unset | set to keep that file after the test |
BANKML_NAMES_ORACLE |
convert.rs (oracle_name_heuristics) |
β | the JSONL that convert_oracle.py --names writes |
BANKML_PENALTY_STEMS |
native.rs (oracle_penalties) |
mindx-gen39-F16,Bonsai-1.7B-Q1_0 |
which models' penalty records to replay |
BANKML_SAMPLER_STEMS |
native.rs (oracle_samplers, 0.3.7) |
mindx-gen39-F16,Bonsai-1.7B-Q1_0 |
which models' sampler records to replay |
BANKML_ORACLE_MODEL |
schema_oracle.py |
the Bonsai and O4 models in .models |
one GGUF to read the chat template from, instead of every reproduced template |
BANKML_FORK |
serve_oracle.py, json_oracle.py, json_schema_oracle.py |
$BANKML_FORKS/Bonsai-8B-Q1_0.gguf.FORK.json |
the Bonsai-8B pin the live oracles serve with |
BANKML_FORKS |
the live oracles (serve_, json_, json_schema_, persona_, penalty_oracle.py) |
~/.local/share/bankml/forks |
where they find each model's pin |
BANKML |
convert_oracle.py |
target/release/bankml |
the binary --bankml runs |
BANKML_TEST_CARRIER |
test_models.py |
unset | 1 adds a real carrier switch and rollback on spare ports (needs Bonsai-1.7B and llama-server); the gate sets it |
BANKML_INFT_ARTIFACT |
test_chain.py |
~/DeltaVerse/deploy/iNFT4/out/iNFT_7857.sol/iNFT_7857.json |
the compiled iNFT_7857 the devnet test deploys |
BANKML_PIN_CPUS |
testing/pinned.sh |
the last two CPUs this shell may use | the cores a benchmark is pinned to (taskset) |
BANKML_PIN_MEM |
testing/pinned.sh |
2500M |
the hard memory cap, with no swap (a user cgroup through systemd-run) |
The Python test suites set BANKML_UI_STATE, BANKML_AGENTS, BANKML_MODELS, BANKML_FORKS, RAGE_PATH,
SAVANTE_CANON, BANKML_VOICE_DIR and BANKML_EXPORT_DIR to temporary directories themselves, so they never touch
your state.
7. Performance and resource tuning
The figures below are from PERFORMANCE.md unless stated. They were measured on an AMD Ryzen 3 3200U laptop (2 cores, 4 threads, 5.8 GB) and a 2-vCPU EPYC VPS.
Threads: three knobs that do different things
BANKML_THREADSsizes bankML's own pool: native serve,generate, the C API. The default is every core. Answers are bit-identical at any thread count, so this is purely a speed and sharing choice. On the laptop (2 cores, 4 threads), one token's ternary matmuls took 0.45β0.47 s on 1 thread, 0.31 s on 2 and 0.23β0.25 s on 3; a fourth thread added nothing. The 1-bit kernel is at parity with ggml at every count. On a shared machine, give bankML fewer threads than cores, as the service in Β§8 does with 1.--threadsis passed to a spawned llama-server (-t). It has no effect in native mode. The importer's carrier usesBANKML_THREADS_SERVE(3) or the Admin tab's choice for whichever engine it starts.BANKML_LLAMA_THREADS(default 3) is about exactness, not speed. It must equal the-tof the llama.cpp run whose tokens you compare against, because llama.cpp's split attention kernel (decode at 512 or more KV cells) chunks by thread count. Leave it at 3 unless your reference ran with another-t.
Context size and KV memory
The KV cache is f16 and grows as the context fills. For the 8B Qwen3 models it costs about 0.15 MB per token: a
2048-token context is about 0.3 GB, 4096 about 0.6 GB. The importer's carrier defaults to 2048 (BANKML_CTX); serve
on its own defaults to 4096. The trade-off (usage.md Β§8):
on a 2048-token context, Savante's persona prompt and .memory leave room for only a few exchanges. With 4096, the
history window can stay put for several turns and keep the prompt cache warm. Set the context through the Admin
tab's RAM budget, so it is planned from the model's actual KV size, or with --ctx.
One resident model, and keep-alive
Native serve holds one model at a time, as Ollama does with MAX_LOADED_MODELS=1. A request for another model
drops the resident one, then verifies and loads the new one: the guard and the full sha256, about 2.9 s for the
1.16 GB model with SHA-NI. Memory is therefore bounded by the largest model you serve plus its KV cache, never two
models at once. --keep-alive (default 5m) decides how long an idle model's memory map stays. A short keep-alive
gives memory back sooner but pays the verify and load again; -1 keeps the model for good. OpenAI-style and
llama-server-style requests keep the model resident.
Prompt cache and slots
Both engines reuse the longest common prefix of the previous prompt. The UI's history window moves in steps of six
exchanges for this reason, so each prompt extends the last one. Both serve one request at a time on one slot (the
spawned llama-server runs with -np 1). The native engine (0.3.8) also keeps llama-server's host prompt cache: when a
request shares little with the slot, the slot's state goes to RAM and a cached state that keeps more of the prompt
comes back, so conversations that take turns do not recompute their history. BANKML_CACHE_RAM bounds it (Β§6). With --slot-dir (the importer always sets it; for the native
carrier since 0.3.8), a restarted engine restores the system prompt's saved KV instead of computing it again. On the laptop that took
the first answer after a restart from 132 s to 15 s. Each saved slot is about 51 MB per model, context and system
prompt.
The GPU share
The GPU worker helps only the 1-bit kernels. On the integrated Vega 3 it measured unchanged within noise (1.93β2.00
tokens/s with the card, 1.96β1.97 without): the card is worth about one CPU core there, and the per-matrix submit
cancels the gain. A discrete card gains directly. The card's buffers come out of system RAM on an integrated GPU.
On a memory-tight machine, BANKML_GPU=off frees that memory and changes no token, because the GPU path is bit-exact
to the CPU path by its own oracle. BANKML_GPU_SHARE overrides the calibration if you have measured a better split.
In the next release (0.3.7), BANKML_GPU_LIMIT (default 0.8) caps the card's memory and busy time, and each matrix
shape decides by measurement whether the card pays. On the Vega 3 (Bonsai-1.7B) that measured neutral, 10.36 tokens/s
off against 10.30β10.38 on, with card memory down from 66 MB to about 13 MB (CHANGELOG.md).
SHA-NI
SHA-NI is on whenever the CPU has it. It makes every start, load and switch faster: restarting the 8B carrier went
from 25 s to 7.0 s. Set BANKML_NO_SHANI=1 only to compare against the portable path or to rule it out while
debugging.
Fixed-resource benchmarking
BANKML_PIN_CPUS=2,3 BANKML_PIN_MEM=2500M testing/pinned.sh <command β¦>
pinned.sh pins the command to fixed cores with taskset and caps its memory, with no swap, in a user cgroup
(systemd-run --user --scope). It records load and free memory before and after. It says plainly when the cgroup
cap could not be enforced, in which case only the CPU is pinned. Compare runs only when they were made on the same
pins.
Small machines: the memory pitfalls the docs record
- The ternary file must stay in the page cache. With the chat engine and a browser holding memory, the 2.31 GB ternary file was only partly resident, so every token re-read it from disk (4.5β5.2 s per token instead of about 0.23 s of matmuls). Measuring the whole-token figure needs at least 3 GB free.
- One Piper render at a time.
speak.pytakes a machine-wide lock ($BANKML_PIPER/.render.lock), because memory, not speed, is the limit. On a small machine runpython3 sAGI/speak.pyalone, or setBANKML_VOICE_ASYNC=0so the UI does not render in the background. - bge-m3 needs about 1.2 GB to load and holds 1.14 GB while loaded. Hence the 1.3 GB free-memory guard
(
BANKML_EMBED_NEED_GB) and the 60 s keep-alive. On a laptop with about 1 GB free while the chat model runs, search falls back to BM25 rather than push the machine into swap. - Weights over 70 % of RAM are refused by the importer for the same reason.
8. Running as a service
This is the shape of the deployed native server on the mindX VPS (2 vCPU EPYC, 7.8 GB). It runs a registry of the pinned models, on one core, with a hard memory cap. Replace the paths with yours.
The registry directory
The server reads every *.FORK.json in the registry directory and looks for each pinned file beside the startup
model and in that directory. So a registry can hold the pins and symlinks to the GGUFs wherever they really
live:
/srv/bankml/registry/
βββ Bonsai-8B-Q1_0.gguf -> /srv/bankml/models/Bonsai-8B-Q1_0.gguf
βββ Bonsai-8B-Q1_0.gguf.FORK.json
βββ Ternary-Bonsai-8B-Q2_0_g64.gguf -> /srv/bankml/models/Ternary-Bonsai-8B-Q2_0_g64.gguf
βββ Ternary-Bonsai-8B-Q2_0_g64.gguf.FORK.json
βββ Bonsai-1.7B-Q1_0.gguf -> β¦
βββ Bonsai-1.7B-Q1_0.gguf.FORK.json
βββ mindx-gen39-F16.gguf -> β¦
βββ mindx-gen39-F16.gguf.FORK.json
βββ mindx-gen39.MODEL.json (a derived model, if you create one: `bankml create β¦ --registry` this dir)
Copy the FORK.json files from the machine that imported the models (~/.local/share/bankml/forks/), or import on
the server itself. The link is never trusted: each file is verified against its pin when it loads. Verification
hashes the canonical (resolved) path, and before every answer the gateway re-checks that file's device, inode, size
and modification time.
The unit
# /etc/systemd/system/bankml.service
[Unit]
Description=bankml serve (native, registry)
After=network.target
[Service]
User=bankml
Group=bankml
Environment=BANKML_THREADS=1
Environment=BANKML_FORKS=/srv/bankml/registry
ExecStart=/opt/bankml/bankml serve /srv/bankml/registry/Bonsai-8B-Q1_0.gguf \
--fork /srv/bankml/registry/Bonsai-8B-Q1_0.gguf.FORK.json \
--native --registry /srv/bankml/registry --ctx 2048 --listen 127.0.0.1:18093
CPUQuota=100%
MemoryMax=3G
Nice=10
Restart=on-failure
[Install]
WantedBy=multi-user.target
Line by line:
User=,Group=: an unprivileged account that can read the registry and the models.servewrites nothing in native mode, unless/api/create,/api/copyorDELETE /api/deletewrite derived models into the registry. Give that account write access to the registry only if you want those endpoints to work.Environment=BANKML_THREADS=1: one pool thread. Native mode ignores--threads; this variable is its thread count. Unset, the pool takes every core.Environment=BANKML_FORKS=β¦: the forks directory for a bare--registry, and forbankml createrun as this user without--registry. Here--registrynames the directory explicitly, so the variable keeps the two in agreement rather than changing what the server serves.ExecStart=: the startup model and its pin (the default model, served when a request names none), native mode, the registry, a 2048-token context and the loopback gateway address.--upstreamis left at127.0.0.1:18092: native mode also listens there with llama-server's endpoints. If something else holds 18092 on this host, add--upstream 127.0.0.1:<free port>.CPUQuota=100%: at most one core's worth of CPU time, matchingBANKML_THREADS=1. The VPS keeps its second core for everything else.MemoryMax=3G: a hard cap sized for the worst case. One model is resident at a time, so the largest one decides: the ternary model at 2.31 GB mapped, plus the f16 KV cache at--ctx 2048(about 0.3 GB for the 8B models), plus the process. If you raise--ctx, raise the cap with it.Nice=10: lower scheduling priority than interactive work on the same host.Restart=on-failure:serveexits 2 when it refuses to start (a file that fails its pin, a busy port). systemd restarts it after any failure, and the journal keeps the reason each time.
sudo systemctl daemon-reload && sudo systemctl enable --now bankml
journalctl -u bankml -f # "bankml serve 0.3.6 --native: β¦ verified (sha256 β¦); models β¦; listening on β¦"
Verifying a deployed server
curl -s 127.0.0.1:18093/bankml # verified, model, the resident model, every name in the registry
curl -s 127.0.0.1:18093/api/tags # every pin: name, size, digest = its sha256, "bankml": {"native", "reason"}
curl -s 127.0.0.1:18093/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"Say hello."}],"max_tokens":16}' # look for bankml_receipt.model_sha256
curl -s 127.0.0.1:18093/bankml/usage # what serve uses now (RSS, CPU %)
Run these on the server itself, or through an SSH tunnel (ssh -L 18093:127.0.0.1:18093 host). The gateway answers
loopback Host names only (Β§9). Compare model_sha256 in the receipt with the pin in the FORK.json.
Upgrades and rollback
Keep each release's binary beside the last one and point the unit at a symlink:
cp target/release/bankml /opt/bankml/bankml-0.3.6
ln -sfn bankml-0.3.6 /opt/bankml/bankml && sudo systemctl restart bankml
# rollback: ln -sfn bankml-0.3.5 /opt/bankml/bankml && sudo systemctl restart bankml
The binary has no dependencies to roll back with it. The pins and models do not change between versions.
9. Security and exposure
| service | bound to | why |
|---|---|---|
bankml serve |
loopback | it has no authentication. It requires a loopback Host (127.0.0.1, localhost, [::1], else 403) and a JSON Content-Type on POST (else 415), so a web page can drive it neither by DNS rebinding nor by a simple cross-origin form POST. It caps connections, header sizes and body sizes, and times out slow clients |
| interact mode (7873) | loopback only | the installed Gradio 3.37 has published path-traversal CVEs (CVE-2023-51449, fixed in 4.11), and interact mode runs the Verifier, devnet mints and PostgreSQL publishing. savante.py refuses any --host other than 127.0.0.1 or localhost and exits with the reason. It also adds a trusted-host check, so a request whose Host is not loopback is refused |
| the bankML console (7875, 0.3.7) | loopback only | it can restart the engine with new CPU, RAM and GPU limits. console.py refuses any --host other than loopback |
| view mode (7874) | 0.0.0.0 by default |
the page for the LAN, built on the Python standard library, not Gradio. It has a fixed set of GET routes (/, /api/state, /api/result?name= for a name testing/results/ lists, /savante.png, /favicon.ico, /knobs.js, /audio/*.ogg, /export/*.opus), and every POST answers 405. Every response carries X-Content-Type-Options: nosniff, Referrer-Policy: no-referrer and a Content-Security-Policy: default-src 'none', with images, scripts, media and connections from 'self' only and inline styles and scripts allowed; file downloads get default-src 'none' alone. It shows commitments, never the content of .history or .memory |
| Ollama (embeddings) | keep it local | history text is sent to it |
| anvil devnet | loopback | started on 127.0.0.1:8545 |
Never expose to a network:
bankml serveand its engine port 18092. A reverse proxy that forwards the publicHostgets 403, and one that rewritesHostwould publish an unauthenticated model server.- Interact mode.
- The PostgreSQL DSN.
- The
install.envand UI state directories.
Use view mode for others on the LAN, and an SSH tunnel for yourself.
Receipts are not signatures. A receipt binds an answer to its request and to the verified model's sha256, between
you and your own bankml serve. It proves nothing to a third party who does not trust your machine
(usage.md Β§9).
The installer needs no sudo. Every download is checked before use: the llama.cpp archive against its published sha256, models against their publisher's sha256 and the guard. On a public chain bankML never signs; the iNFT path sends transactions only to a chain with id 31337 (anvil).
10. Development and the release gate
cargo test --release # offline, no models needed
BANKML_GGML_LIB=/path/to/llama-b11192 LLAMA_SRC=/path/to/llama.cpp-b11192 \
BANKML_LLAMA_SRC=/path/to/llama.cpp-b11192 testing/release_gate.sh # β testing/results/<version>.txt
BANKML_GGML_LIBis the unpacked b11192 release, the same directory the installer'senginestep makes (~/.local/share/bankml/llama-b11192). The oracles load ggml's own compiled kernels from it in-process, and compare bankML's results bit for bit.LLAMA_SRCis a llama.cpp checkout at tag b11192 (git clone --depth 1 --branch b11192 https://github.com/ggml-org/llama.cpp). It enables the schema and content oracles, which compile against its headers.BANKML_LLAMA_SRC(defaultupstream/llama.cpp) is the source tree whose converter the conversion oracles run. It enablesoracle_name_heuristics.- Without
BANKML_GGML_LIB, the gate runs only its offline part: build, tests, C API tests, clippy, licence headers, the Python suites, guard agreement and the printf oracle. It printsBANKML_GGML_LIB unset: oracles, A/B and budgets skipped. - Models. Every oracle that needs a model reads it from
.models/, and every live stage needs that model's pin inBANKML_FORKS. A missing model or record is reported as skipped, not passed.
Running stages individually. Each oracle is an ignored Rust test. The gate finds its full path and runs it alone:
cargo test --release -q -- --list --ignored | grep oracle_json_mode # its full name
cargo test --release -- --ignored --exact <full::name> --nocapture --test-threads=1
testing/live.sh "json mode" cargo test --release -- --ignored --exact <full::name> --nocapture --test-threads=1
python3 -B testing/serve_oracle.py --bankml Bonsai-1.7B-Q1_0 # a live stage on its own
testing/live.sh appends a step's output to testing/live.log, which view mode shows as it runs. The live stages
start their own serve --native on spare loopback ports (serve_oracle.py uses 18195 and 18196), so they do not
collide with a running carrier on 18092 and 18093.
Time and memory. A full gate runs every oracle one after another (--test-threads=1), and takes a long time on a
laptop: plan to leave the machine to it. Memory is the usual limit:
- The live stages load real models. The 0.3.5 run was stopped by the development session's memory guard, and the
0.3.6 run stalled at the same stage, because the Vega 3 GPU worker's buffers come out of system RAM. In both, the
remaining stages ran with
BANKML_GPU=offand the record marks where. The next release addsBANKML_GPU_LIMIT(0.3.7) for this. - The whole-token budgets (
decode_budget_*) measure the disk, not the kernel, unless the model stays in the page cache. That needs at least 3 GB free for the ternary model. - Close other work and stop the carrier (
./install.sh stop) before a gate run, and setBANKML_GPU=offon an integrated-GPU machine short of memory.
A speed counts only if every oracle passed on the same code. The full list of tests and their records is in testing/README.md and oracles.md.
11. Troubleshooting
Installing
| symptom | cause | fix |
|---|---|---|
./install.sh: Permission denied |
the copy lost its executable bit (zip download, USB stick) | bash install.sh, or chmod +x install.sh |
unknown argument: β¦ (exit 2) |
a typo, or an option install.sh does not have |
./install.sh --help |
Rust X is older than 1.99 |
an older toolchain without rustup | install rustup (it reads rust-toolchain.toml), or rustup toolchain install 1.99.0 |
python3 -m venv failed |
python3-venv not installed |
sudo apt install python3-venv, or BANKML_PYTHON=/path/to/python with Gradio 3 and numpy |
β¦ sha256 β¦ is not the published β¦: refused |
a corrupted or substituted llama.cpp download | run ./install.sh engine again; the bad file was deleted |
no llama.cpp b11192 build for <OS>/<arch> here |
the release archive is Linux x86_64 only | build llama.cpp b11192 and set BANKML_LLAMA_SERVER |
bankml serve did not come up: see β¦/carrier.log |
the carrier failed to verify or start | read ~/.local/share/bankml/savante/carrier.log; the last lines name the refusal |
a UI step says did not start: see β¦/logs/<name>.log |
Python, Gradio or the canon | read that log |
Starting bankml serve
| symptom | cause | fix |
|---|---|---|
| the port does not open for a while after start | serve binds only after hashing the whole file and, with --spawn, after llama-server is healthy (up to 900 s for a cold load) |
wait for GET /bankml, not for the port. Hashing the 1.16 GB model takes about 3 s with SHA-NI |
refuse: β¦ != pinned |
the file is not the one the FORK.json pins |
re-download it; bankml sha256 FILE |
β¦ changed while it was being hashed: refused |
the file was being written or replaced during the start | let the copy finish, then start again |
<upstream> already answers: an engine is running there β¦ |
--spawn with something already on --upstream |
stop it, or drop --spawn and use --upstream to check it by its /props |
upstream β¦ serves X, not the verified Y: refused |
the llama-server on that port runs another file | restart it on the verified file, or use --spawn |
upstream β¦ not healthy after N s |
llama-server not up | check its log; a cold load can take a minute |
the engine exited (β¦) before it was healthy |
llama-server could not load the model (memory, a wrong build) | run the same llama-server by hand to see its error |
cannot listen on the engine address 127.0.0.1:18092 β¦ (is llama-server running there?) |
native mode also binds --upstream, and something holds it |
stop that process, or pass --upstream 127.0.0.1:<free port> |
cannot listen on 127.0.0.1:18093 |
another serve is running |
./install.sh status; ./install.sh stop |
the carrier's ports are still held by [pids]; stop them by hand |
the importer can stop only this user's bankml and llama-server |
stop those processes yourself |
Requests
| symptom | cause | fix |
|---|---|---|
403 β¦ loopback clients only |
the Host header is not 127.0.0.1, localhost or [::1] (a LAN address, a proxy) |
call it on loopback, through an SSH tunnel if remote |
415 POST bodies must be Content-Type: application/json |
curl -d alone sends a form type |
add -H 'Content-Type: application/json' |
400 chunked requests are not accepted |
a client that streams the request body | send Content-Length |
413 request too large |
a body over 16 MiB | send less |
503 too many connections |
32 connections already open | fewer concurrent clients; one model answers one request at a time anyway |
400 request (N tokens) exceeds the available context size (M tokens) (0.3.8) |
the prompt does not fit --ctx |
trim the conversation, or restart with a larger --ctx if it fits in memory |
501 This server does not support slots action |
a /slots request to a server started without --slot-dir |
restart with --slot-dir DIR |
404 model 'β¦' not found: bankml serves pinned models only (β¦) |
a name that no pin has | use one of the names listed; add --registry to serve more than the startup model |
400 num_ctx N β¦ is larger than the context this server holds |
a request or Modelfile asks for more than --ctx |
restart with a larger --ctx, if it fits in memory |
a 400 naming a refused option, sampler, tools, images, think: true or a schema |
not something bankML reproduces exactly | the message says what; usage.md Β§6a lists every refusal |
| the chat says Cannot reach bankml serve | the carrier is not running | ./install.sh status, then ./install.sh model |
Savante, models and resources
| symptom | cause | fix |
|---|---|---|
refused: interact mode serves loopback only β¦ |
--host 0.0.0.0 (or any non-loopback host) on interact mode |
use view mode for the LAN |
savante.py started by hand cannot start the carrier |
without the installer's environment, BANKML_LLAMA_SERVER defaults to ~/sAGI/bonsai/llama-b11192/llama-server |
export BANKML_LLAMA_SERVER to the engine in install.env (INSTALL_LLAMA_SERVER), or start with ./install.sh start |
an import says licence β¦ is not open source |
the model's licence is not on the OSI list | not importable, by design |
needs X GB + 1.5 GB margin; Y GB free on disk |
not enough disk | free space or set BANKML_MODELS to a larger disk |
X GB of weights on a Y GB machine would page to swap |
weights over 70 % of RAM | a smaller model |
β¦ exists but is not the published file (sha256 differs); move it aside first |
a stale or partial file with that name | move it, import again |
β¦ is an embedding model β¦; it cannot be the chat carrier |
bge-m3 and other encoders | use them through Ollama for embeddings (embedding.md) |
| Apply on the Admin tab says the budget cannot hold the model | the RAM budget is under weights + 250 MB + 512 tokens of KV | raise the budget; the message gives the minimum |
| search says BM25 alone | Ollama down, bge-m3 not pulled, or less than 1.3 GB free | ollama pull bge-m3; free memory, or lower BANKML_EMBED_NEED_GB if you know the risk |
| refused: savante.persona does not verify | the canon was edited or is incomplete | git -C ~/cryptoAGI/savante status; run the Verifier |
the log says bankml: GPU not used: β¦ not verified |
the card failed the on-card oracle | expected behaviour: bankML stays on the CPU. bankml gpu --verify shows the row that differs |
| the machine swaps during a gate or a long session | the integrated GPU's buffers, a second model, Piper, bge-m3 | BANKML_GPU=off, one Piper render at a time, stop the carrier before a gate |
| view from another computer does not load | firewall, or view bound to 127.0.0.1 | --host 0.0.0.0; allow TCP 7874 |
The basic symptoms are also in usage.md Β§12.