File size: 10,908 Bytes
730c5bb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 | # bankml β the verified CPU engine
[bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime for 1-bit (`Q1_0`),
ternary (`Q2_0_g64`) and F16 GGUF models, **token-identical to llama.cpp b11192** on every oracle in
its release gate. `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*`
from its own forward pass on `127.0.0.1:18093`, and puts a **receipt** on every answer. mindXtrain
uses it as a serve target, an operator backend and a second imprint instrument.
mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml code is vendored
(the clean-room policy in [`CLAUDE.md`](../CLAUDE.md)). bankml's own docs:
[README](https://github.com/cryptoAGI/bankml#readme) Β·
[usage](https://github.com/cryptoAGI/bankml/blob/main/docs/usage.md) Β·
[bankML as mindX's Ollama](https://github.com/cryptoAGI/bankml/blob/main/docs/OLLAMA.md).
## What bankml runs, and what it refuses
| | bankml |
|---|---|
| **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
| **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop`; from **0.3.6** also `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`, token for token as llama-server b11192 applies them |
| **Refused with HTTP 400 and a reason** | penalties on bankml before 0.3.6, `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` |
| **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |
A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
mindXtrain mirrors that β it **never retries a refused request with altered parameters**, and never
drops a Modelfile instruction to make bankml accept it.
## Operator backend β `MINDXTRAIN_BACKEND=bankml`
`mindxtrain/operator/backends/bankml.py` registers `bankml` (a subclass of `openai_compat`).
```bash
bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry # 127.0.0.1:18093
MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
```
- `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
- `MINDXTRAIN_BANKML_OPTIONS`: a JSON object of sampling fields to send on every request (or
`BankmlBackend(options=...)`), e.g. `{"repeat_penalty": 1.3}` for a small generation on 0.3.6+.
Nothing extra is sent unless it is set.
- `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
Streamed answers parse bankml's final `data: {"bankml_receipt": β¦}` event.
- HTTP 400 β `BankmlRefusal(reason)` β the operator answers 400 with bankml's reason. Other non-2xx
and a mid-stream `{"error": β¦}` β `BankmlError`.
- Health: `/health`, `/readyz` and `/coach/api/health` probe `GET /bankml` (the identity endpoint) and
list the model from `/v1/models`.
- **Auto-detect** (no `MINDXTRAIN_BACKEND`): ollama β vllm β bankml β vllm. bankml is chosen only
when ollama is down, `GET /bankml` answers and vLLM does not, so a host that resolved to ollama or
vllm before keeps doing so.
- Governance panels (`resolve_chat_base_url`): `MINDXTRAIN_BACKEND=bankml` uses bankml's URL; else
`MINDXTRAIN_BANKML_BASE_URL` is appended *last* to the OPENAI β VLLM β OLLAMA chain, and only
when it is set.
## Serving a trained run β `mindxtrain serve --to bankml`
```bash
uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
[--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \
[--bankml-system "You are mindX, generation 80."] \
[--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>']
```
A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured
on gen39, the bare pin answers runs of `,` with or without a per-request penalty, and the layered
tag answers in words.
1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
2. Checks the binary: `bankml version` and the verbs in `bankml --help`. `bankml create` and
`bankml convert` arrive in **bankml 0.3.5**; an older binary is reported as
`bankml_too_old` (exit 2), never as a crash. The verbs are the truth β an unreleased build may
still say 0.3.4 and already carry them.
3. Merges the LoRA (`merge_lora_adapter`, needs `uv sync --extra ml`), or takes the checkpoint as an
already-merged directory when it holds `config.json` and no `adapter_config.json`.
4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
The directory is passed through a link named after the tag (`<work>/<tag>/<tag>`): llama.cpp
names a model after the directory it reads, so `merged/` would give `general.name` "Merged"
and a different sha256. Through the link gen39 converts to `6b64c748β¦`, the pin mindX serves.
Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before.
`--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
(best-effort, as for ollama).
### The Modelfile subset (`bankml_sanitize`)
| instruction | bankml |
|---|---|
| `FROM` merged dir / pinned GGUF / registry name | taken |
| `SYSTEM`, `MESSAGE`, `LICENSE`, `REQUIRES` | taken (recorded) |
| `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
| `ADAPTER` | **refused** β merge first (`push_to_bankml` does) |
| `TEMPLATE` | **refused** unless equal to the base's own chat template |
| penalties, `repeat_last_n` | taken on bankml **0.3.6+**; **refused** before (checked against `bankml version`) |
| mirostat*, `typical_p` | **refused** β not reproduced |
| `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** β a resource option is not part of a model |
Each refusal is returned with its reason; nothing is dropped silently. Python API:
`mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult` (never raises; `status` is
one of `created`, `refused`, `bankml_missing`, `bankml_too_old`, `merge_failed`, `failed`, `error`).
## A second imprint instrument β `mindxtrain imprint-bankml`
```bash
uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \
[--seed 0] [--num-predict 48] [--system "β¦"] [--base-url http://127.0.0.1:18093/v1]
```
Poses the script's user-turns to two tags on bankml's `/api/chat` with `temperature 0`, a fixed
`seed`, `num_predict 48` and **no penalties**, then scores with the existing `score_imprint`. The
report (`BankmlImprintReport`) carries the decoding, every utterance's receipt and the distinct
`model_sha256` values, and `report.method` is tagged `<scorer>/bankml-greedy`.
**It is not comparable with the canonical gate.** `mindxtrain imprint` decodes with transformers
greedy, `repetition_penalty 1.3` and `no_repeat_ngram_size 3`; every number in an ascent log comes
from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so
(`canonical_gate: false`, `comparable_with: "bankml-greedy only"`). What it buys:
- **reproducible** β the same seed and weights give the same tokens, identical to llama.cpp;
- **auditable** β a score is tied to the exact weights by sha256, not to a tag name;
- **cheap** β a 135M F16 actor answers on one CPU core, with no torch in the probing process.
Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes
without a system turn (runs of `,` and `?||`), which is the very behaviour the 1.3 penalty in the
canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops
repeating; pass the persona's `--system` as the coach does.
## Console and published Modelfiles
- `mindxtrain/ui/console.py` no longer sends a penalty or mirostat option left at the engine's own
default (`repeat_penalty 1.1`, presence / frequency `0`, `typical_p 1`, `mirostat 0` and its
tau / eta while it is off), so bankml accepts a console request with default settings. A
deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means
its own default, except that a Modelfile's `PARAMETER repeat_penalty` now applies where the
console used to override it with 1.1.
- `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
penalty line, which bankml can load. The default stays 1.3.
## The chat and judge models on a bankml node
The governance panel and the LLM judges used to name `llama3.2` when no model was given; bankml
serves no such model. Both now read the environment at call time:
| variable | for | example on the VPS |
|---|---|---|
| `MINDXTRAIN_CHAT_MODEL` | boardroom members and dojo judges with no `model` | `bonsai-8b-q1_0` |
| `MINDXTRAIN_JUDGE_MODEL` | `CorrectnessEvaluator`, `PairwiseEvaluator`, `GuidelineEvaluator`, classroom (falls back to the chat model) | `bonsai-8b-q1_0` |
| `MINDXTRAIN_CHAT_OPTIONS` | extra fields on every `chat_once` body | `{"repeat_penalty": 1.3}` |
`classroom(use_judge=True)` with no model now uses that judge instead of silently skipping it.
Speeds (bankml's `docs/PERFORMANCE.md`, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M
F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5
(the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge).
The VPS runs bankml on one thread: expect about a third of that.
Deploying all of this to the mindX VPS, step by step: [install.md](install.md).
## Tests
`tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`,
`tests/test_bankml_extras.py` β no
network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
`tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
cannot change what the other auto-detect tests resolve to.
|