mindXtrain / docs /bankml.md
Gregory-L's picture
bankml: the verified CPU engine as backend, serve target and imprint probe (#1)
730c5bb
|
Raw History Blame Contribute Delete
10.9 kB

bankml β€” the verified CPU engine

bankml is a zero-dependency Rust runtime for 1-bit (Q1_0), ternary (Q2_0_g64) and F16 GGUF models, token-identical to llama.cpp b11192 on every oracle in its release gate. bankml serve --native answers OpenAI /v1/chat/completions and Ollama /api/* from its own forward pass on 127.0.0.1:18093, and puts a receipt on every answer. mindXtrain uses it as a serve target, an operator backend and a second imprint instrument.

mindXtrain reaches bankml only over HTTP and as a CLI subprocess. No bankml code is vendored (the clean-room policy in CLAUDE.md). bankml's own docs: README Β· usage Β· bankML as mindX's Ollama.

What bankml runs, and what it refuses

bankml
Architectures Qwen3 (Q1_0, Q2_0_g64: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's mindx-genN)
Sampling it reproduces temperature, top_k, top_p, min_p, seed, JSON mode; num_ctx, num_predict, stop; from 0.3.6 also repeat_penalty, repeat_last_n, presence_penalty, frequency_penalty, token for token as llama-server b11192 applies them
Refused with HTTP 400 and a reason penalties on bankml before 0.3.6, mirostat, typical_p, tools, images, a replacement template, unknown architectures, Q8_0 / Q4_K / BF16
Receipt (bankml_receipt) bankml version, engine, model_sha256, guard, prompt_tokens, completion_tokens, ttft_ms, wall_ms, response_sha256, request_sha256, signed: false

A refusal is the product, not a defect: bankml answers only what its verified forward pass does. mindXtrain mirrors that β€” it never retries a refused request with altered parameters, and never drops a Modelfile instruction to make bankml accept it.

Operator backend β€” MINDXTRAIN_BACKEND=bankml

mindxtrain/operator/backends/bankml.py registers bankml (a subclass of openai_compat).

bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry   # 127.0.0.1:18093
MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
  • MINDXTRAIN_BANKML_BASE_URL (default http://127.0.0.1:18093/v1).
  • MINDXTRAIN_BANKML_OPTIONS: a JSON object of sampling fields to send on every request (or BankmlBackend(options=...)), e.g. {"repeat_penalty": 1.3} for a small generation on 0.3.6+. Nothing extra is sent unless it is set.
  • POST /v1/chat/completions returns ChatResponse.receipt; the backend also keeps last_receipt. Streamed answers parse bankml's final data: {"bankml_receipt": …} event.
  • HTTP 400 β†’ BankmlRefusal(reason) β†’ the operator answers 400 with bankml's reason. Other non-2xx and a mid-stream {"error": …} β†’ BankmlError.
  • Health: /health, /readyz and /coach/api/health probe GET /bankml (the identity endpoint) and list the model from /v1/models.
  • Auto-detect (no MINDXTRAIN_BACKEND): ollama β†’ vllm β†’ bankml β†’ vllm. bankml is chosen only when ollama is down, GET /bankml answers and vLLM does not, so a host that resolved to ollama or vllm before keeps doing so.
  • Governance panels (resolve_chat_base_url): MINDXTRAIN_BACKEND=bankml uses bankml's URL; else MINDXTRAIN_BANKML_BASE_URL is appended last to the OPENAI β†’ VLLM β†’ OLLAMA chain, and only when it is set.

Serving a trained run β€” mindxtrain serve --to bankml

uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
    [--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \
    [--bankml-system "You are mindX, generation 80."] \
    [--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>']

A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured on gen39, the bare pin answers runs of , with or without a per-request penalty, and the layered tag answers in words.

  1. Refuses up front (exit 2): quantize.enabled with a scheme other than none (bankml serves the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
  2. Checks the binary: bankml version and the verbs in bankml --help. bankml create and bankml convert arrive in bankml 0.3.5; an older binary is reported as bankml_too_old (exit 2), never as a crash. The verbs are the truth β€” an unreleased build may still say 0.3.4 and already carry them.
  3. Merges the LoRA (merge_lora_adapter, needs uv sync --extra ml), or takes the checkpoint as an already-merged directory when it holds config.json and no adapter_config.json.
  4. Re-checks the merged config.json: only LlamaForCausalLM converts.
  5. Writes a Modelfile through bankml_sanitize and runs bankml create <tag> -f Modelfile, which converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it. The directory is passed through a link named after the tag (<work>/<tag>/<tag>): llama.cpp names a model after the directory it reads, so merged/ would give general.name "Merged" and a different sha256. Through the link gen39 converts to 6b64c748…, the pin mindX serves. Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before. --bankml-convert runs bankml convert (GGUF + FORK.json) first and writes FROM <gguf>.
  6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
  7. --register-as-fallback PATCHes mindX's fallback model to {provider: "bankml", model: <tag>} (best-effort, as for ollama).

The Modelfile subset (bankml_sanitize)

instruction bankml
FROM merged dir / pinned GGUF / registry name taken
SYSTEM, MESSAGE, LICENSE, REQUIRES taken (recorded)
PARAMETER temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; stop taken
ADAPTER refused β€” merge first (push_to_bankml does)
TEMPLATE refused unless equal to the base's own chat template
penalties, repeat_last_n taken on bankml 0.3.6+; refused before (checked against bankml version)
mirostat*, typical_p refused β€” not reproduced
num_gpu, num_thread, num_batch, num_keep, draft_num_predict refused β€” a resource option is not part of a model

Each refusal is returned with its reason; nothing is dropped silently. Python API: mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult (never raises; status is one of created, refused, bankml_missing, bankml_too_old, merge_failed, failed, error).

A second imprint instrument β€” mindxtrain imprint-bankml

uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \
    [--seed 0] [--num-predict 48] [--system "…"] [--base-url http://127.0.0.1:18093/v1]

Poses the script's user-turns to two tags on bankml's /api/chat with temperature 0, a fixed seed, num_predict 48 and no penalties, then scores with the existing score_imprint. The report (BankmlImprintReport) carries the decoding, every utterance's receipt and the distinct model_sha256 values, and report.method is tagged <scorer>/bankml-greedy.

It is not comparable with the canonical gate. mindxtrain imprint decodes with transformers greedy, repetition_penalty 1.3 and no_repeat_ngram_size 3; every number in an ascent log comes from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so (canonical_gate: false, comparable_with: "bankml-greedy only"). What it buys:

  • reproducible β€” the same seed and weights give the same tokens, identical to llama.cpp;
  • auditable β€” a score is tied to the exact weights by sha256, not to a tag name;
  • cheap β€” a 135M F16 actor answers on one CPU core, with no torch in the probing process.

Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes without a system turn (runs of , and ?||), which is the very behaviour the 1.3 penalty in the canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops repeating; pass the persona's --system as the coach does.

Console and published Modelfiles

  • mindxtrain/ui/console.py no longer sends a penalty or mirostat option left at the engine's own default (repeat_penalty 1.1, presence / frequency 0, typical_p 1, mirostat 0 and its tau / eta while it is off), so bankml accepts a console request with default settings. A deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means its own default, except that a Modelfile's PARAMETER repeat_penalty now applies where the console used to override it with 1.1.
  • hf.extension.publish_generation(..., repeat_penalty=None) publishes a Modelfile without the penalty line, which bankml can load. The default stays 1.3.

The chat and judge models on a bankml node

The governance panel and the LLM judges used to name llama3.2 when no model was given; bankml serves no such model. Both now read the environment at call time:

variable for example on the VPS
MINDXTRAIN_CHAT_MODEL boardroom members and dojo judges with no model bonsai-8b-q1_0
MINDXTRAIN_JUDGE_MODEL CorrectnessEvaluator, PairwiseEvaluator, GuidelineEvaluator, classroom (falls back to the chat model) bonsai-8b-q1_0
MINDXTRAIN_CHAT_OPTIONS extra fields on every chat_once body {"repeat_penalty": 1.3}

classroom(use_judge=True) with no model now uses that judge instead of silently skipping it.

Speeds (bankml's docs/PERFORMANCE.md, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5 (the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge). The VPS runs bankml on one thread: expect about a third of that.

Deploying all of this to the mindX VPS, step by step: install.md.

Tests

tests/test_bankml_backend.py, tests/test_bankml_push.py, tests/test_imprint_bankml.py, tests/test_bankml_extras.py β€” no network and no binary: httpx.MockTransport, monkeypatched subprocess.run / shutil.which. tests/conftest.py pins the bankml auto-detect probe to "absent" so a developer box running bankml cannot change what the other auto-detect tests resolve to.