mindXtrain / docs /bankml.md
Gregory-L's picture
bankml: the verified CPU engine as backend, serve target and imprint probe (#1)
730c5bb
|
Raw History Blame Contribute Delete
10.9 kB
# bankml β€” the verified CPU engine
[bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime for 1-bit (`Q1_0`),
ternary (`Q2_0_g64`) and F16 GGUF models, **token-identical to llama.cpp b11192** on every oracle in
its release gate. `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*`
from its own forward pass on `127.0.0.1:18093`, and puts a **receipt** on every answer. mindXtrain
uses it as a serve target, an operator backend and a second imprint instrument.
mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml code is vendored
(the clean-room policy in [`CLAUDE.md`](../CLAUDE.md)). bankml's own docs:
[README](https://github.com/cryptoAGI/bankml#readme) Β·
[usage](https://github.com/cryptoAGI/bankml/blob/main/docs/usage.md) Β·
[bankML as mindX's Ollama](https://github.com/cryptoAGI/bankml/blob/main/docs/OLLAMA.md).
## What bankml runs, and what it refuses
| | bankml |
|---|---|
| **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
| **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop`; from **0.3.6** also `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`, token for token as llama-server b11192 applies them |
| **Refused with HTTP 400 and a reason** | penalties on bankml before 0.3.6, `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` |
| **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |
A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
mindXtrain mirrors that β€” it **never retries a refused request with altered parameters**, and never
drops a Modelfile instruction to make bankml accept it.
## Operator backend β€” `MINDXTRAIN_BACKEND=bankml`
`mindxtrain/operator/backends/bankml.py` registers `bankml` (a subclass of `openai_compat`).
```bash
bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry # 127.0.0.1:18093
MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
```
- `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
- `MINDXTRAIN_BANKML_OPTIONS`: a JSON object of sampling fields to send on every request (or
`BankmlBackend(options=...)`), e.g. `{"repeat_penalty": 1.3}` for a small generation on 0.3.6+.
Nothing extra is sent unless it is set.
- `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event.
- HTTP 400 β†’ `BankmlRefusal(reason)` β†’ the operator answers 400 with bankml's reason. Other non-2xx
and a mid-stream `{"error": …}` β†’ `BankmlError`.
- Health: `/health`, `/readyz` and `/coach/api/health` probe `GET /bankml` (the identity endpoint) and
list the model from `/v1/models`.
- **Auto-detect** (no `MINDXTRAIN_BACKEND`): ollama β†’ vllm β†’ bankml β†’ vllm. bankml is chosen only
when ollama is down, `GET /bankml` answers and vLLM does not, so a host that resolved to ollama or
vllm before keeps doing so.
- Governance panels (`resolve_chat_base_url`): `MINDXTRAIN_BACKEND=bankml` uses bankml's URL; else
`MINDXTRAIN_BANKML_BASE_URL` is appended *last* to the OPENAI β†’ VLLM β†’ OLLAMA chain, and only
when it is set.
## Serving a trained run β€” `mindxtrain serve --to bankml`
```bash
uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
[--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \
[--bankml-system "You are mindX, generation 80."] \
[--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>']
```
A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured
on gen39, the bare pin answers runs of `,` with or without a per-request penalty, and the layered
tag answers in words.
1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
2. Checks the binary: `bankml version` and the verbs in `bankml --help`. `bankml create` and
`bankml convert` arrive in **bankml 0.3.5**; an older binary is reported as
`bankml_too_old` (exit 2), never as a crash. The verbs are the truth β€” an unreleased build may
still say 0.3.4 and already carry them.
3. Merges the LoRA (`merge_lora_adapter`, needs `uv sync --extra ml`), or takes the checkpoint as an
already-merged directory when it holds `config.json` and no `adapter_config.json`.
4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
The directory is passed through a link named after the tag (`<work>/<tag>/<tag>`): llama.cpp
names a model after the directory it reads, so `merged/` would give `general.name` "Merged"
and a different sha256. Through the link gen39 converts to `6b64c748…`, the pin mindX serves.
Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before.
`--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
(best-effort, as for ollama).
### The Modelfile subset (`bankml_sanitize`)
| instruction | bankml |
|---|---|
| `FROM` merged dir / pinned GGUF / registry name | taken |
| `SYSTEM`, `MESSAGE`, `LICENSE`, `REQUIRES` | taken (recorded) |
| `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
| `ADAPTER` | **refused** β€” merge first (`push_to_bankml` does) |
| `TEMPLATE` | **refused** unless equal to the base's own chat template |
| penalties, `repeat_last_n` | taken on bankml **0.3.6+**; **refused** before (checked against `bankml version`) |
| mirostat*, `typical_p` | **refused** β€” not reproduced |
| `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** β€” a resource option is not part of a model |
Each refusal is returned with its reason; nothing is dropped silently. Python API:
`mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult` (never raises; `status` is
one of `created`, `refused`, `bankml_missing`, `bankml_too_old`, `merge_failed`, `failed`, `error`).
## A second imprint instrument β€” `mindxtrain imprint-bankml`
```bash
uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \
[--seed 0] [--num-predict 48] [--system "…"] [--base-url http://127.0.0.1:18093/v1]
```
Poses the script's user-turns to two tags on bankml's `/api/chat` with `temperature 0`, a fixed
`seed`, `num_predict 48` and **no penalties**, then scores with the existing `score_imprint`. The
report (`BankmlImprintReport`) carries the decoding, every utterance's receipt and the distinct
`model_sha256` values, and `report.method` is tagged `<scorer>/bankml-greedy`.
**It is not comparable with the canonical gate.** `mindxtrain imprint` decodes with transformers
greedy, `repetition_penalty 1.3` and `no_repeat_ngram_size 3`; every number in an ascent log comes
from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so
(`canonical_gate: false`, `comparable_with: "bankml-greedy only"`). What it buys:
- **reproducible** β€” the same seed and weights give the same tokens, identical to llama.cpp;
- **auditable** β€” a score is tied to the exact weights by sha256, not to a tag name;
- **cheap** β€” a 135M F16 actor answers on one CPU core, with no torch in the probing process.
Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes
without a system turn (runs of `,` and `?||`), which is the very behaviour the 1.3 penalty in the
canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops
repeating; pass the persona's `--system` as the coach does.
## Console and published Modelfiles
- `mindxtrain/ui/console.py` no longer sends a penalty or mirostat option left at the engine's own
default (`repeat_penalty 1.1`, presence / frequency `0`, `typical_p 1`, `mirostat 0` and its
tau / eta while it is off), so bankml accepts a console request with default settings. A
deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means
its own default, except that a Modelfile's `PARAMETER repeat_penalty` now applies where the
console used to override it with 1.1.
- `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
penalty line, which bankml can load. The default stays 1.3.
## The chat and judge models on a bankml node
The governance panel and the LLM judges used to name `llama3.2` when no model was given; bankml
serves no such model. Both now read the environment at call time:
| variable | for | example on the VPS |
|---|---|---|
| `MINDXTRAIN_CHAT_MODEL` | boardroom members and dojo judges with no `model` | `bonsai-8b-q1_0` |
| `MINDXTRAIN_JUDGE_MODEL` | `CorrectnessEvaluator`, `PairwiseEvaluator`, `GuidelineEvaluator`, classroom (falls back to the chat model) | `bonsai-8b-q1_0` |
| `MINDXTRAIN_CHAT_OPTIONS` | extra fields on every `chat_once` body | `{"repeat_penalty": 1.3}` |
`classroom(use_judge=True)` with no model now uses that judge instead of silently skipping it.
Speeds (bankml's `docs/PERFORMANCE.md`, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M
F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5
(the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge).
The VPS runs bankml on one thread: expect about a third of that.
Deploying all of this to the mindX VPS, step by step: [install.md](install.md).
## Tests
`tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`,
`tests/test_bankml_extras.py` β€” no
network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
`tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
cannot change what the other auto-detect tests resolve to.