File size: 10,908 Bytes
730c5bb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
# bankml β€” the verified CPU engine

[bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime for 1-bit (`Q1_0`),
ternary (`Q2_0_g64`) and F16 GGUF models, **token-identical to llama.cpp b11192** on every oracle in
its release gate. `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*`
from its own forward pass on `127.0.0.1:18093`, and puts a **receipt** on every answer. mindXtrain
uses it as a serve target, an operator backend and a second imprint instrument.

mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml code is vendored
(the clean-room policy in [`CLAUDE.md`](../CLAUDE.md)). bankml's own docs:
[README](https://github.com/cryptoAGI/bankml#readme) Β·
[usage](https://github.com/cryptoAGI/bankml/blob/main/docs/usage.md) Β·
[bankML as mindX's Ollama](https://github.com/cryptoAGI/bankml/blob/main/docs/OLLAMA.md).

## What bankml runs, and what it refuses

| | bankml |
|---|---|
| **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
| **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop`; from **0.3.6** also `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`, token for token as llama-server b11192 applies them |
| **Refused with HTTP 400 and a reason** | penalties on bankml before 0.3.6, `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` |
| **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |

A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
mindXtrain mirrors that β€” it **never retries a refused request with altered parameters**, and never
drops a Modelfile instruction to make bankml accept it.

## Operator backend β€” `MINDXTRAIN_BACKEND=bankml`

`mindxtrain/operator/backends/bankml.py` registers `bankml` (a subclass of `openai_compat`).

```bash
bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry   # 127.0.0.1:18093
MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
```

- `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
- `MINDXTRAIN_BANKML_OPTIONS`: a JSON object of sampling fields to send on every request (or
  `BankmlBackend(options=...)`), e.g. `{"repeat_penalty": 1.3}` for a small generation on 0.3.6+.
  Nothing extra is sent unless it is set.
- `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
  Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event.
- HTTP 400 β†’ `BankmlRefusal(reason)` β†’ the operator answers 400 with bankml's reason. Other non-2xx
  and a mid-stream `{"error": …}` β†’ `BankmlError`.
- Health: `/health`, `/readyz` and `/coach/api/health` probe `GET /bankml` (the identity endpoint) and
  list the model from `/v1/models`.
- **Auto-detect** (no `MINDXTRAIN_BACKEND`): ollama β†’ vllm β†’ bankml β†’ vllm. bankml is chosen only
  when ollama is down, `GET /bankml` answers and vLLM does not, so a host that resolved to ollama or
  vllm before keeps doing so.
- Governance panels (`resolve_chat_base_url`): `MINDXTRAIN_BACKEND=bankml` uses bankml's URL; else
  `MINDXTRAIN_BANKML_BASE_URL` is appended *last* to the OPENAI β†’ VLLM β†’ OLLAMA chain, and only
  when it is set.

## Serving a trained run β€” `mindxtrain serve --to bankml`

```bash
uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
    [--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \
    [--bankml-system "You are mindX, generation 80."] \
    [--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>']
```

A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured
on gen39, the bare pin answers runs of `,` with or without a per-request penalty, and the layered
tag answers in words.

1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
   the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
   families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
2. Checks the binary: `bankml version` and the verbs in `bankml --help`. `bankml create` and
   `bankml convert` arrive in **bankml 0.3.5**; an older binary is reported as
   `bankml_too_old` (exit 2), never as a crash. The verbs are the truth β€” an unreleased build may
   still say 0.3.4 and already carry them.
3. Merges the LoRA (`merge_lora_adapter`, needs `uv sync --extra ml`), or takes the checkpoint as an
   already-merged directory when it holds `config.json` and no `adapter_config.json`.
4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
   converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
   The directory is passed through a link named after the tag (`<work>/<tag>/<tag>`): llama.cpp
   names a model after the directory it reads, so `merged/` would give `general.name` "Merged"
   and a different sha256. Through the link gen39 converts to `6b64c748…`, the pin mindX serves.
   Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before.
   `--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
   (best-effort, as for ollama).

### The Modelfile subset (`bankml_sanitize`)

| instruction | bankml |
|---|---|
| `FROM` merged dir / pinned GGUF / registry name | taken |
| `SYSTEM`, `MESSAGE`, `LICENSE`, `REQUIRES` | taken (recorded) |
| `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
| `ADAPTER` | **refused** β€” merge first (`push_to_bankml` does) |
| `TEMPLATE` | **refused** unless equal to the base's own chat template |
| penalties, `repeat_last_n` | taken on bankml **0.3.6+**; **refused** before (checked against `bankml version`) |
| mirostat*, `typical_p` | **refused** β€” not reproduced |
| `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** β€” a resource option is not part of a model |

Each refusal is returned with its reason; nothing is dropped silently. Python API:
`mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult` (never raises; `status` is
one of `created`, `refused`, `bankml_missing`, `bankml_too_old`, `merge_failed`, `failed`, `error`).

## A second imprint instrument β€” `mindxtrain imprint-bankml`

```bash
uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \
    [--seed 0] [--num-predict 48] [--system "…"] [--base-url http://127.0.0.1:18093/v1]
```

Poses the script's user-turns to two tags on bankml's `/api/chat` with `temperature 0`, a fixed
`seed`, `num_predict 48` and **no penalties**, then scores with the existing `score_imprint`. The
report (`BankmlImprintReport`) carries the decoding, every utterance's receipt and the distinct
`model_sha256` values, and `report.method` is tagged `<scorer>/bankml-greedy`.

**It is not comparable with the canonical gate.** `mindxtrain imprint` decodes with transformers
greedy, `repetition_penalty 1.3` and `no_repeat_ngram_size 3`; every number in an ascent log comes
from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so
(`canonical_gate: false`, `comparable_with: "bankml-greedy only"`). What it buys:

- **reproducible** β€” the same seed and weights give the same tokens, identical to llama.cpp;
- **auditable** β€” a score is tied to the exact weights by sha256, not to a tag name;
- **cheap** β€” a 135M F16 actor answers on one CPU core, with no torch in the probing process.

Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes
without a system turn (runs of `,` and `?||`), which is the very behaviour the 1.3 penalty in the
canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops
repeating; pass the persona's `--system` as the coach does.

## Console and published Modelfiles

- `mindxtrain/ui/console.py` no longer sends a penalty or mirostat option left at the engine's own
  default (`repeat_penalty 1.1`, presence / frequency `0`, `typical_p 1`, `mirostat 0` and its
  tau / eta while it is off), so bankml accepts a console request with default settings. A
  deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means
  its own default, except that a Modelfile's `PARAMETER repeat_penalty` now applies where the
  console used to override it with 1.1.
- `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
  penalty line, which bankml can load. The default stays 1.3.

## The chat and judge models on a bankml node

The governance panel and the LLM judges used to name `llama3.2` when no model was given; bankml
serves no such model. Both now read the environment at call time:

| variable | for | example on the VPS |
|---|---|---|
| `MINDXTRAIN_CHAT_MODEL` | boardroom members and dojo judges with no `model` | `bonsai-8b-q1_0` |
| `MINDXTRAIN_JUDGE_MODEL` | `CorrectnessEvaluator`, `PairwiseEvaluator`, `GuidelineEvaluator`, classroom (falls back to the chat model) | `bonsai-8b-q1_0` |
| `MINDXTRAIN_CHAT_OPTIONS` | extra fields on every `chat_once` body | `{"repeat_penalty": 1.3}` |

`classroom(use_judge=True)` with no model now uses that judge instead of silently skipping it.

Speeds (bankml's `docs/PERFORMANCE.md`, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M
F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5
(the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge).
The VPS runs bankml on one thread: expect about a third of that.

Deploying all of this to the mindX VPS, step by step: [install.md](install.md).

## Tests

`tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`,
`tests/test_bankml_extras.py` β€” no
network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
`tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
cannot change what the other auto-detect tests resolve to.