# bankml — the verified CPU engine [bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime for 1-bit (`Q1_0`), ternary (`Q2_0_g64`) and F16 GGUF models, **token-identical to llama.cpp b11192** on every oracle in its release gate. `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*` from its own forward pass on `127.0.0.1:18093`, and puts a **receipt** on every answer. mindXtrain uses it as a serve target, an operator backend and a second imprint instrument. mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml code is vendored (the clean-room policy in [`CLAUDE.md`](../CLAUDE.md)). bankml's own docs: [README](https://github.com/cryptoAGI/bankml#readme) · [usage](https://github.com/cryptoAGI/bankml/blob/main/docs/usage.md) · [bankML as mindX's Ollama](https://github.com/cryptoAGI/bankml/blob/main/docs/OLLAMA.md). ## What bankml runs, and what it refuses | | bankml | |---|---| | **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) | | **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop`; from **0.3.6** also `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`, token for token as llama-server b11192 applies them | | **Refused with HTTP 400 and a reason** | penalties on bankml before 0.3.6, `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` | | **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` | A refusal is the product, not a defect: bankml answers only what its verified forward pass does. mindXtrain mirrors that — it **never retries a refused request with altered parameters**, and never drops a Modelfile instruction to make bankml accept it. ## Operator backend — `MINDXTRAIN_BACKEND=bankml` `mindxtrain/operator/backends/bankml.py` registers `bankml` (a subclass of `openai_compat`). ```bash bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry # 127.0.0.1:18093 MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080 ``` - `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`). - `MINDXTRAIN_BANKML_OPTIONS`: a JSON object of sampling fields to send on every request (or `BankmlBackend(options=...)`), e.g. `{"repeat_penalty": 1.3}` for a small generation on 0.3.6+. Nothing extra is sent unless it is set. - `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`. Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event. - HTTP 400 → `BankmlRefusal(reason)` → the operator answers 400 with bankml's reason. Other non-2xx and a mid-stream `{"error": …}` → `BankmlError`. - Health: `/health`, `/readyz` and `/coach/api/health` probe `GET /bankml` (the identity endpoint) and list the model from `/v1/models`. - **Auto-detect** (no `MINDXTRAIN_BACKEND`): ollama → vllm → bankml → vllm. bankml is chosen only when ollama is down, `GET /bankml` answers and vLLM does not, so a host that resolved to ollama or vllm before keeps doing so. - Governance panels (`resolve_chat_base_url`): `MINDXTRAIN_BACKEND=bankml` uses bankml's URL; else `MINDXTRAIN_BANKML_BASE_URL` is appended *last* to the OPENAI → VLLM → OLLAMA chain, and only when it is set. ## Serving a trained run — `mindxtrain serve --to bankml` ```bash uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \ [--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \ [--bankml-system "You are mindX, generation 80."] \ [--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>'] ``` A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured on gen39, the bare pin answers runs of `,` with or without a per-request penalty, and the layered tag answers in words. 1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella). 2. Checks the binary: `bankml version` and the verbs in `bankml --help`. `bankml create` and `bankml convert` arrive in **bankml 0.3.5**; an older binary is reported as `bankml_too_old` (exit 2), never as a crash. The verbs are the truth — an unreleased build may still say 0.3.4 and already carry them. 3. Merges the LoRA (`merge_lora_adapter`, needs `uv sync --extra ml`), or takes the checkpoint as an already-merged directory when it holds `config.json` and no `adapter_config.json`. 4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts. 5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create -f Modelfile`, which converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it. The directory is passed through a link named after the tag (`//`): llama.cpp names a model after the directory it reads, so `merged/` would give `general.name` "Merged" and a different sha256. Through the link gen39 converts to `6b64c748…`, the pin mindX serves. Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before. `--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM `. 6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest. 7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: }` (best-effort, as for ollama). ### The Modelfile subset (`bankml_sanitize`) | instruction | bankml | |---|---| | `FROM` merged dir / pinned GGUF / registry name | taken | | `SYSTEM`, `MESSAGE`, `LICENSE`, `REQUIRES` | taken (recorded) | | `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken | | `ADAPTER` | **refused** — merge first (`push_to_bankml` does) | | `TEMPLATE` | **refused** unless equal to the base's own chat template | | penalties, `repeat_last_n` | taken on bankml **0.3.6+**; **refused** before (checked against `bankml version`) | | mirostat*, `typical_p` | **refused** — not reproduced | | `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** — a resource option is not part of a model | Each refusal is returned with its reason; nothing is dropped silently. Python API: `mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult` (never raises; `status` is one of `created`, `refused`, `bankml_missing`, `bankml_too_old`, `merge_failed`, `failed`, `error`). ## A second imprint instrument — `mindxtrain imprint-bankml` ```bash uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \ [--seed 0] [--num-predict 48] [--system "…"] [--base-url http://127.0.0.1:18093/v1] ``` Poses the script's user-turns to two tags on bankml's `/api/chat` with `temperature 0`, a fixed `seed`, `num_predict 48` and **no penalties**, then scores with the existing `score_imprint`. The report (`BankmlImprintReport`) carries the decoding, every utterance's receipt and the distinct `model_sha256` values, and `report.method` is tagged `/bankml-greedy`. **It is not comparable with the canonical gate.** `mindxtrain imprint` decodes with transformers greedy, `repetition_penalty 1.3` and `no_repeat_ngram_size 3`; every number in an ascent log comes from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so (`canonical_gate: false`, `comparable_with: "bankml-greedy only"`). What it buys: - **reproducible** — the same seed and weights give the same tokens, identical to llama.cpp; - **auditable** — a score is tied to the exact weights by sha256, not to a tag name; - **cheap** — a 135M F16 actor answers on one CPU core, with no torch in the probing process. Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes without a system turn (runs of `,` and `?||`), which is the very behaviour the 1.3 penalty in the canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops repeating; pass the persona's `--system` as the coach does. ## Console and published Modelfiles - `mindxtrain/ui/console.py` no longer sends a penalty or mirostat option left at the engine's own default (`repeat_penalty 1.1`, presence / frequency `0`, `typical_p 1`, `mirostat 0` and its tau / eta while it is off), so bankml accepts a console request with default settings. A deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means its own default, except that a Modelfile's `PARAMETER repeat_penalty` now applies where the console used to override it with 1.1. - `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the penalty line, which bankml can load. The default stays 1.3. ## The chat and judge models on a bankml node The governance panel and the LLM judges used to name `llama3.2` when no model was given; bankml serves no such model. Both now read the environment at call time: | variable | for | example on the VPS | |---|---|---| | `MINDXTRAIN_CHAT_MODEL` | boardroom members and dojo judges with no `model` | `bonsai-8b-q1_0` | | `MINDXTRAIN_JUDGE_MODEL` | `CorrectnessEvaluator`, `PairwiseEvaluator`, `GuidelineEvaluator`, classroom (falls back to the chat model) | `bonsai-8b-q1_0` | | `MINDXTRAIN_CHAT_OPTIONS` | extra fields on every `chat_once` body | `{"repeat_penalty": 1.3}` | `classroom(use_judge=True)` with no model now uses that judge instead of silently skipping it. Speeds (bankml's `docs/PERFORMANCE.md`, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5 (the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge). The VPS runs bankml on one thread: expect about a third of that. Deploying all of this to the mindX VPS, step by step: [install.md](install.md). ## Tests `tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`, `tests/test_bankml_extras.py` — no network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`. `tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml cannot change what the other auto-detect tests resolve to.