bankml 0.3.6 on PR #1: penalties, byte-identical pins, a bankml node's own chat and judge models, and the VPS install guide
Browse files- serve --to bankml takes repeat_penalty / repeat_last_n / presence_penalty /
frequency_penalty when `bankml version` is 0.3.6+ (bankml's O2) and refuses them,
with the reason, before; --bankml-system, --bankml-param KEY=VALUE and
--bankml-stop layer a persona and its sampling over the pin. Measured on gen39:
the bare pin answers ",,,,"; the layered tag answers in words.
- The merged directory is converted through a link named after the tag: llama.cpp
names a model after its directory, so `merged/` gave general.name "Merged" and a
different sha256. Through the link gen39 converts to 6b64c748…, the pin mindX
serves (verified against bankml 0.3.6 with the Hub's gen39 weights).
- BankmlBackend sends sampling fields from options= / MINDXTRAIN_BANKML_OPTIONS,
only when set.
- MINDXTRAIN_CHAT_MODEL / MINDXTRAIN_JUDGE_MODEL / MINDXTRAIN_CHAT_OPTIONS: the
panel and the judges named llama3.2, which a bankml node does not serve.
classroom(use_judge=True) with no model uses the default judge instead of skipping.
- docs/install.md: bankml 0.3.6 and this lane on the mindX VPS, in copy-paste
steps with a check after each and a rollback; steps 3–4 replayed locally.
852 passed, 3 skipped (PR #1's 842 + 10); ruff clean on the changed files; mypy clean.
Co-Authored-By: Professor Codephreak <codephreak@pythai.net>
- docs/CHANGELOG.md +13 -0
- docs/NAV.md +2 -0
- docs/bankml.md +40 -5
- docs/install.md +247 -0
- mindxtrain/cli/main.py +44 -0
- mindxtrain/deploy/bankml_push.py +59 -10
- mindxtrain/eval/llama_evals.py +16 -6
- mindxtrain/governance/classroom.py +1 -1
- mindxtrain/governance/panel.py +38 -5
- mindxtrain/operator/backends/bankml.py +26 -4
- tests/test_bankml_extras.py +114 -0
- tests/test_bankml_push.py +41 -5
|
@@ -8,6 +8,19 @@ project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
| 8 |
|
| 9 |
### Added
|
| 10 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
- **bankml, the verified CPU engine** ([github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml),
|
| 12 |
[`docs/bankml.md`](bankml.md)), reached over HTTP and its CLI only — nothing vendored.
|
| 13 |
- `operator/backends/bankml.py`: `@register_backend("bankml")`, `MINDXTRAIN_BANKML_BASE_URL`
|
|
|
|
| 8 |
|
| 9 |
### Added
|
| 10 |
|
| 11 |
+
- **bankml 0.3.6 penalties.** `serve --to bankml` takes `repeat_penalty`, `repeat_last_n`,
|
| 12 |
+
`presence_penalty` and `frequency_penalty` in the Modelfile when `bankml version` is 0.3.6 or
|
| 13 |
+
later (bankml's O2), and refuses them with the reason before; `--bankml-system`,
|
| 14 |
+
`--bankml-param KEY=VALUE` and `--bankml-stop` layer a persona and its sampling over the pin.
|
| 15 |
+
The backend sends sampling fields from `options=` / `MINDXTRAIN_BANKML_OPTIONS`, only when set.
|
| 16 |
+
- **Byte-identical conversion.** The merged directory is converted through a link named after the
|
| 17 |
+
tag, so llama.cpp's `general.name` is the tag's and gen39 converts to `6b64c748…`, the GGUF mindX
|
| 18 |
+
pins (through `merged/` it was "Merged" and a different sha256).
|
| 19 |
+
- **`MINDXTRAIN_CHAT_MODEL`, `MINDXTRAIN_JUDGE_MODEL`, `MINDXTRAIN_CHAT_OPTIONS`.** The panel and
|
| 20 |
+
the judges named `llama3.2` when given no model; a bankml node serves none. Read at call time.
|
| 21 |
+
`classroom(use_judge=True)` without a model now uses the default judge instead of skipping it.
|
| 22 |
+
- **`docs/install.md`**: deploying bankml 0.3.6 and mindXtrain's bankml lane to the mindX VPS.
|
| 23 |
+
|
| 24 |
- **bankml, the verified CPU engine** ([github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml),
|
| 25 |
[`docs/bankml.md`](bankml.md)), reached over HTTP and its CLI only — nothing vendored.
|
| 26 |
- `operator/backends/bankml.py`: `@register_backend("bankml")`, `MINDXTRAIN_BANKML_BASE_URL`
|
|
@@ -13,6 +13,7 @@ reading paths first, then jump straight to the section you need.
|
|
| 13 |
- **Running the demo / operating a real run?** → [HANDOFF.md](HANDOFF.md) → [CLI](cli.md) → [YAML schema](yaml_schema.md).
|
| 14 |
- **Contributing code?** → [Development workflow](development.md) → [Architecture](architecture.md) → [Actualization status](actualization_status.md).
|
| 15 |
- **Driving the UI?** → [Coach UI](coach.md) → [dcoach](dcoach.md) → [Governance](governance.md).
|
|
|
|
| 16 |
- **Researching the space?** → [Decentralized training deep-dive](decentralized-training-deep-dive-2026.md) → [LLM training-stack landscape](mindxtrain-llm-training-landscape-2026.md) → [Autoresearch strategy](mindXtrain_autoresearch_strategy.md).
|
| 17 |
|
| 18 |
---
|
|
@@ -132,6 +133,7 @@ The interactive `/coach/` operator UI: create-script, live-training diagnostics,
|
|
| 132 |
- [What bankml runs, and what it refuses](bankml.md#what-bankml-runs-and-what-it-refuses)
|
| 133 |
- [Operator backend](bankml.md#operator-backend--mindxtrain_backendbankml) · [`serve --to bankml`](bankml.md#serving-a-trained-run--mindxtrain-serve---to-bankml) · [The Modelfile subset](bankml.md#the-modelfile-subset-bankml_sanitize)
|
| 134 |
- [`imprint-bankml` (not comparable with the canonical gate)](bankml.md#a-second-imprint-instrument--mindxtrain-imprint-bankml)
|
|
|
|
| 135 |
|
| 136 |
### [Governance](governance.md)
|
| 137 |
classroom (graduation) / boardroom (any-N consensus) / dojo (prime-N dispute settlement), model-backed deliberation.
|
|
|
|
| 13 |
- **Running the demo / operating a real run?** → [HANDOFF.md](HANDOFF.md) → [CLI](cli.md) → [YAML schema](yaml_schema.md).
|
| 14 |
- **Contributing code?** → [Development workflow](development.md) → [Architecture](architecture.md) → [Actualization status](actualization_status.md).
|
| 15 |
- **Driving the UI?** → [Coach UI](coach.md) → [dcoach](dcoach.md) → [Governance](governance.md).
|
| 16 |
+
- **Deploying to the mindX VPS?** → [install.md](install.md) (bankml 0.3.6 + the bankml lane, step by step) → [bankml](bankml.md).
|
| 17 |
- **Researching the space?** → [Decentralized training deep-dive](decentralized-training-deep-dive-2026.md) → [LLM training-stack landscape](mindxtrain-llm-training-landscape-2026.md) → [Autoresearch strategy](mindXtrain_autoresearch_strategy.md).
|
| 18 |
|
| 19 |
---
|
|
|
|
| 133 |
- [What bankml runs, and what it refuses](bankml.md#what-bankml-runs-and-what-it-refuses)
|
| 134 |
- [Operator backend](bankml.md#operator-backend--mindxtrain_backendbankml) · [`serve --to bankml`](bankml.md#serving-a-trained-run--mindxtrain-serve---to-bankml) · [The Modelfile subset](bankml.md#the-modelfile-subset-bankml_sanitize)
|
| 135 |
- [`imprint-bankml` (not comparable with the canonical gate)](bankml.md#a-second-imprint-instrument--mindxtrain-imprint-bankml)
|
| 136 |
+
- [The chat and judge models on a bankml node](bankml.md#the-chat-and-judge-models-on-a-bankml-node) · [Install on the mindX VPS](install.md)
|
| 137 |
|
| 138 |
### [Governance](governance.md)
|
| 139 |
classroom (graduation) / boardroom (any-N consensus) / dojo (prime-N dispute settlement), model-backed deliberation.
|
|
@@ -17,8 +17,8 @@ mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml
|
|
| 17 |
| | bankml |
|
| 18 |
|---|---|
|
| 19 |
| **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
|
| 20 |
-
| **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop` |
|
| 21 |
-
| **Refused with HTTP 400 and a reason** |
|
| 22 |
| **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |
|
| 23 |
|
| 24 |
A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
|
|
@@ -35,6 +35,9 @@ MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
|
|
| 35 |
```
|
| 36 |
|
| 37 |
- `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
|
|
|
|
|
|
|
|
|
|
| 38 |
- `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
|
| 39 |
Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event.
|
| 40 |
- HTTP 400 → `BankmlRefusal(reason)` → the operator answers 400 with bankml's reason. Other non-2xx
|
|
@@ -52,9 +55,15 @@ MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
|
|
| 52 |
|
| 53 |
```bash
|
| 54 |
uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
|
| 55 |
-
[--bankml-bin PATH] [--bankml-convert] [--register-as-fallback]
|
|
|
|
|
|
|
| 56 |
```
|
| 57 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
|
| 59 |
the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
|
| 60 |
families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
|
|
@@ -67,6 +76,10 @@ uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint D
|
|
| 67 |
4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
|
| 68 |
5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
|
| 69 |
converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 70 |
`--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
|
| 71 |
6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
|
| 72 |
7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
|
|
@@ -81,7 +94,8 @@ uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint D
|
|
| 81 |
| `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
|
| 82 |
| `ADAPTER` | **refused** — merge first (`push_to_bankml` does) |
|
| 83 |
| `TEMPLATE` | **refused** unless equal to the base's own chat template |
|
| 84 |
-
| penalties, `repeat_last_n`
|
|
|
|
| 85 |
| `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** — a resource option is not part of a model |
|
| 86 |
|
| 87 |
Each refusal is returned with its reason; nothing is dropped silently. Python API:
|
|
@@ -125,9 +139,30 @@ repeating; pass the persona's `--system` as the coach does.
|
|
| 125 |
- `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
|
| 126 |
penalty line, which bankml can load. The default stays 1.3.
|
| 127 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
## Tests
|
| 129 |
|
| 130 |
-
`tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`
|
|
|
|
| 131 |
network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
|
| 132 |
`tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
|
| 133 |
cannot change what the other auto-detect tests resolve to.
|
|
|
|
| 17 |
| | bankml |
|
| 18 |
|---|---|
|
| 19 |
| **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
|
| 20 |
+
| **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop`; from **0.3.6** also `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`, token for token as llama-server b11192 applies them |
|
| 21 |
+
| **Refused with HTTP 400 and a reason** | penalties on bankml before 0.3.6, `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` |
|
| 22 |
| **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |
|
| 23 |
|
| 24 |
A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
|
|
|
|
| 35 |
```
|
| 36 |
|
| 37 |
- `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
|
| 38 |
+
- `MINDXTRAIN_BANKML_OPTIONS`: a JSON object of sampling fields to send on every request (or
|
| 39 |
+
`BankmlBackend(options=...)`), e.g. `{"repeat_penalty": 1.3}` for a small generation on 0.3.6+.
|
| 40 |
+
Nothing extra is sent unless it is set.
|
| 41 |
- `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
|
| 42 |
Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event.
|
| 43 |
- HTTP 400 → `BankmlRefusal(reason)` → the operator answers 400 with bankml's reason. Other non-2xx
|
|
|
|
| 55 |
|
| 56 |
```bash
|
| 57 |
uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
|
| 58 |
+
[--bankml-bin PATH] [--bankml-convert] [--register-as-fallback] \
|
| 59 |
+
[--bankml-system "You are mindX, generation 80."] \
|
| 60 |
+
[--bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>']
|
| 61 |
```
|
| 62 |
|
| 63 |
+
A SmolLM2 / mindx-genN generation wants all of the bracketed last line on bankml 0.3.6+: measured
|
| 64 |
+
on gen39, the bare pin answers runs of `,` with or without a per-request penalty, and the layered
|
| 65 |
+
tag answers in words.
|
| 66 |
+
|
| 67 |
1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
|
| 68 |
the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
|
| 69 |
families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
|
|
|
|
| 76 |
4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
|
| 77 |
5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
|
| 78 |
converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
|
| 79 |
+
The directory is passed through a link named after the tag (`<work>/<tag>/<tag>`): llama.cpp
|
| 80 |
+
names a model after the directory it reads, so `merged/` would give `general.name` "Merged"
|
| 81 |
+
and a different sha256. Through the link gen39 converts to `6b64c748…`, the pin mindX serves.
|
| 82 |
+
Penalties in the Modelfile are taken on bankml 0.3.6+ and refused, with the reason, before.
|
| 83 |
`--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
|
| 84 |
6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
|
| 85 |
7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
|
|
|
|
| 94 |
| `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
|
| 95 |
| `ADAPTER` | **refused** — merge first (`push_to_bankml` does) |
|
| 96 |
| `TEMPLATE` | **refused** unless equal to the base's own chat template |
|
| 97 |
+
| penalties, `repeat_last_n` | taken on bankml **0.3.6+**; **refused** before (checked against `bankml version`) |
|
| 98 |
+
| mirostat*, `typical_p` | **refused** — not reproduced |
|
| 99 |
| `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** — a resource option is not part of a model |
|
| 100 |
|
| 101 |
Each refusal is returned with its reason; nothing is dropped silently. Python API:
|
|
|
|
| 139 |
- `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
|
| 140 |
penalty line, which bankml can load. The default stays 1.3.
|
| 141 |
|
| 142 |
+
## The chat and judge models on a bankml node
|
| 143 |
+
|
| 144 |
+
The governance panel and the LLM judges used to name `llama3.2` when no model was given; bankml
|
| 145 |
+
serves no such model. Both now read the environment at call time:
|
| 146 |
+
|
| 147 |
+
| variable | for | example on the VPS |
|
| 148 |
+
|---|---|---|
|
| 149 |
+
| `MINDXTRAIN_CHAT_MODEL` | boardroom members and dojo judges with no `model` | `bonsai-8b-q1_0` |
|
| 150 |
+
| `MINDXTRAIN_JUDGE_MODEL` | `CorrectnessEvaluator`, `PairwiseEvaluator`, `GuidelineEvaluator`, classroom (falls back to the chat model) | `bonsai-8b-q1_0` |
|
| 151 |
+
| `MINDXTRAIN_CHAT_OPTIONS` | extra fields on every `chat_once` body | `{"repeat_penalty": 1.3}` |
|
| 152 |
+
|
| 153 |
+
`classroom(use_judge=True)` with no model now uses that judge instead of silently skipping it.
|
| 154 |
+
|
| 155 |
+
Speeds (bankml's `docs/PERFORMANCE.md`, a Ryzen 3 3200U at 3 threads): mindx-genN / SmolLM2-135M
|
| 156 |
+
F16 ~38 tokens/s (the fastest; serving and voice probes, too weak to judge); Bonsai-1.7B Q1_0 ~8.5
|
| 157 |
+
(the fastest that judges); Bonsai-8B Q1_0 and Ternary-Bonsai-8B Q2_0_g64 ~2.4 (teacher and judge).
|
| 158 |
+
The VPS runs bankml on one thread: expect about a third of that.
|
| 159 |
+
|
| 160 |
+
Deploying all of this to the mindX VPS, step by step: [install.md](install.md).
|
| 161 |
+
|
| 162 |
## Tests
|
| 163 |
|
| 164 |
+
`tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py`,
|
| 165 |
+
`tests/test_bankml_extras.py` — no
|
| 166 |
network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
|
| 167 |
`tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
|
| 168 |
cannot change what the other auto-detect tests resolve to.
|
|
@@ -0,0 +1,247 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Install on the mindX VPS: bankml 0.3.6 and mindXtrain's bankml lane
|
| 2 |
+
|
| 3 |
+
This guide upgrades the mindX node (mindx.pythai.net, `168.231.126.58`) in eight steps:
|
| 4 |
+
|
| 5 |
+
1. bankml 0.3.5 → 0.3.6 (repetition penalties, token-identical to llama-server)
|
| 6 |
+
2. gen39 re-layered so it answers in words
|
| 7 |
+
3. mindXtrain → the Hub's `main` (the bankml serve target, backend and judges)
|
| 8 |
+
4. mindX pointed at bankml for mindXtrain's chat, judges and teacher
|
| 9 |
+
|
| 10 |
+
Every block is meant to be copied whole and pasted into a root shell on the VPS. Each step ends
|
| 11 |
+
with a check; do not go on until it passes. [Roll back](#rollback) undoes any step.
|
| 12 |
+
|
| 13 |
+
**What you need:** root SSH to the VPS, about 15 minutes, and 1 GB free on `/`. Restarting
|
| 14 |
+
`bankml` makes mindX's default model unavailable for a few seconds; restarting `mindx` takes
|
| 15 |
+
3–4 minutes before `/health` answers again.
|
| 16 |
+
|
| 17 |
+
---
|
| 18 |
+
|
| 19 |
+
## 0. Log in and look before you touch
|
| 20 |
+
|
| 21 |
+
```bash
|
| 22 |
+
ssh root@168.231.126.58
|
| 23 |
+
```
|
| 24 |
+
|
| 25 |
+
```bash
|
| 26 |
+
df -h / | tail -1
|
| 27 |
+
systemctl is-active bankml mindx
|
| 28 |
+
curl -s 127.0.0.1:18093/bankml | head -c 200; echo
|
| 29 |
+
/home/mindx/bankml-0.3.5/bankml --version
|
| 30 |
+
```
|
| 31 |
+
|
| 32 |
+
You should see both services `active`, a `"verdict": "play"` from bankml, and `bankml 0.3.5`.
|
| 33 |
+
|
| 34 |
+
Check that no training run is going (a mindX restart in step 7 would end it):
|
| 35 |
+
|
| 36 |
+
```bash
|
| 37 |
+
pgrep -af 'mindxtrain (train|ascend)' || echo "no training running"
|
| 38 |
+
```
|
| 39 |
+
|
| 40 |
+
## 1. Back up what you will change
|
| 41 |
+
|
| 42 |
+
```bash
|
| 43 |
+
TS=$(date -u +%Y%m%dT%H%M%SZ); mkdir -p /root/deploy_backup_bankml036_$TS
|
| 44 |
+
cp /etc/systemd/system/bankml.service /home/mindx/mindX/.env /root/deploy_backup_bankml036_$TS/
|
| 45 |
+
echo "/root/deploy_backup_bankml036_$TS" | tee /root/LAST_BANKML036_BACKUP
|
| 46 |
+
ls -la /root/deploy_backup_bankml036_$TS
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
## 2. Build bankml 0.3.6
|
| 50 |
+
|
| 51 |
+
bankml has no dependencies to fetch; it builds with the `mindx` user's cargo in a few minutes on
|
| 52 |
+
the VPS's two cores. Build from the `v0.3.6` tag once it is published; until then, from `main`.
|
| 53 |
+
|
| 54 |
+
```bash
|
| 55 |
+
sudo -u mindx bash -lc 'cd /home/mindx && rm -rf bankml-src-0.3.6 && git clone --depth 1 --branch v0.3.6 https://github.com/cryptoAGI/bankml bankml-src-0.3.6 || git clone --depth 1 https://github.com/cryptoAGI/bankml bankml-src-0.3.6'
|
| 56 |
+
```
|
| 57 |
+
|
| 58 |
+
```bash
|
| 59 |
+
sudo -u mindx bash -lc 'cd /home/mindx/bankml-src-0.3.6 && nice -n 19 cargo build --release 2>&1 | tail -3'
|
| 60 |
+
```
|
| 61 |
+
|
| 62 |
+
```bash
|
| 63 |
+
sudo -u mindx mkdir -p /home/mindx/bankml-0.3.6
|
| 64 |
+
sudo -u mindx cp /home/mindx/bankml-src-0.3.6/target/release/bankml /home/mindx/bankml-0.3.6/bankml
|
| 65 |
+
/home/mindx/bankml-0.3.6/bankml --version
|
| 66 |
+
```
|
| 67 |
+
|
| 68 |
+
**Check:** it prints `bankml 0.3.6` (or later). If it prints 0.3.5, the tag is not out yet and
|
| 69 |
+
`main` has not moved past it either: stop here, nothing has changed on the running node.
|
| 70 |
+
|
| 71 |
+
## 3. Point the service at 0.3.6 and restart bankml
|
| 72 |
+
|
| 73 |
+
```bash
|
| 74 |
+
sed -i 's#/home/mindx/bankml-0\.3\.5/bankml#/home/mindx/bankml-0.3.6/bankml#' /etc/systemd/system/bankml.service
|
| 75 |
+
grep ExecStart /etc/systemd/system/bankml.service
|
| 76 |
+
```
|
| 77 |
+
|
| 78 |
+
```bash
|
| 79 |
+
systemctl daemon-reload && systemctl restart bankml && sleep 5 && systemctl is-active bankml
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
```bash
|
| 83 |
+
curl -s 127.0.0.1:18093/bankml | grep -o '"bankml": *"[^"]*"'
|
| 84 |
+
```
|
| 85 |
+
|
| 86 |
+
**Check:** `"bankml": "0.3.6"`. Now prove the penalties are honoured (0.3.5 answered HTTP 400 to
|
| 87 |
+
this request):
|
| 88 |
+
|
| 89 |
+
```bash
|
| 90 |
+
curl -s -o /dev/null -w '%{http_code}\n' 127.0.0.1:18093/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"mindx-gen39-f16","messages":[{"role":"user","content":"Who are you?"}],"max_tokens":16,"repeat_penalty":1.3}'
|
| 91 |
+
```
|
| 92 |
+
|
| 93 |
+
**Check:** `200`.
|
| 94 |
+
|
| 95 |
+
## 4. Layer gen39 so it answers in words
|
| 96 |
+
|
| 97 |
+
A 135M generation on its own repeats itself (`,,,,`). A layer over the pinned GGUF adds its
|
| 98 |
+
persona, its stop string and a 1.3 repeat penalty; the weights are not copied.
|
| 99 |
+
|
| 100 |
+
```bash
|
| 101 |
+
cat > /home/mindx/models/Modelfile.mindx-gen39 <<'MF'
|
| 102 |
+
FROM mindx-gen39-f16
|
| 103 |
+
SYSTEM """You are mindX, generation 39."""
|
| 104 |
+
PARAMETER stop "<|im_end|>"
|
| 105 |
+
PARAMETER num_ctx 2048
|
| 106 |
+
PARAMETER repeat_penalty 1.3
|
| 107 |
+
MF
|
| 108 |
+
chown mindx:mindx /home/mindx/models/Modelfile.mindx-gen39
|
| 109 |
+
```
|
| 110 |
+
|
| 111 |
+
```bash
|
| 112 |
+
sudo -u mindx /home/mindx/bankml-0.3.6/bankml create mindx-gen39 -f /home/mindx/models/Modelfile.mindx-gen39 --registry /home/mindx/models/registry
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
```bash
|
| 116 |
+
curl -s 127.0.0.1:18093/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"mindx-gen39","messages":[{"role":"user","content":"Who are you?"}],"max_tokens":48}' | python3 -c 'import json,sys; print(json.load(sys.stdin)["choices"][0]["message"]["content"])'
|
| 117 |
+
```
|
| 118 |
+
|
| 119 |
+
**Check:** an answer in words, not a run of commas. (`mindx-gen39-f16` is still the bare pin;
|
| 120 |
+
`mindx-gen39` is now the layer. One thread on the VPS: give it 10–30 seconds.)
|
| 121 |
+
|
| 122 |
+
## 5. Bring mindXtrain to the Hub's `main`
|
| 123 |
+
|
| 124 |
+
mindXtrain moved from GitHub (archived) to `huggingface.co/PYTHAI/mindXtrain`. The checkout on
|
| 125 |
+
the VPS still tracks GitHub at v1.0.2 and has a local `uv.lock` edit, which is set aside first.
|
| 126 |
+
|
| 127 |
+
```bash
|
| 128 |
+
cd /home/mindx/mindXtrain
|
| 129 |
+
sudo -u mindx git stash push -m "pre-bankml uv.lock" -- uv.lock || true
|
| 130 |
+
sudo -u mindx git remote add hf https://huggingface.co/PYTHAI/mindXtrain 2>/dev/null || true
|
| 131 |
+
sudo -u mindx GIT_LFS_SKIP_SMUDGE=1 git fetch hf main
|
| 132 |
+
sudo -u mindx git checkout -B main hf/main
|
| 133 |
+
sudo -u mindx git log --oneline -1
|
| 134 |
+
```
|
| 135 |
+
|
| 136 |
+
Install it into the existing virtualenv. `--inexact` keeps the extras already installed there
|
| 137 |
+
(torch, peft, trl), which a plain `uv sync` would remove:
|
| 138 |
+
|
| 139 |
+
```bash
|
| 140 |
+
cd /home/mindx/mindXtrain && sudo -u mindx /home/mindx/.local/bin/uv sync --inexact --extra ml 2>&1 | tail -2
|
| 141 |
+
```
|
| 142 |
+
|
| 143 |
+
```bash
|
| 144 |
+
cd /home/mindx/mindXtrain && sudo -u mindx .venv/bin/mindxtrain serve --help | grep -c 'bankml'
|
| 145 |
+
```
|
| 146 |
+
|
| 147 |
+
**Check:** a number above 0 (the `--to bankml` options are there).
|
| 148 |
+
|
| 149 |
+
## 6. Tell mindX's mindXtrain to use bankml
|
| 150 |
+
|
| 151 |
+
mindX starts mindXtrain with its own environment (`/home/mindx/mindX/.env`). Append the bankml
|
| 152 |
+
lane (Bonsai-8B is the model on the VPS that is good enough to judge and teach; gen39 is too
|
| 153 |
+
small for that):
|
| 154 |
+
|
| 155 |
+
```bash
|
| 156 |
+
cat >> /home/mindx/mindX/.env <<'ENV'
|
| 157 |
+
|
| 158 |
+
# mindXtrain on bankml (docs/install.md in PYTHAI/mindXtrain)
|
| 159 |
+
MINDXTRAIN_BANKML_BASE_URL=http://127.0.0.1:18093/v1
|
| 160 |
+
MINDXTRAIN_CHAT_MODEL=bonsai-8b-q1_0
|
| 161 |
+
MINDXTRAIN_JUDGE_MODEL=bonsai-8b-q1_0
|
| 162 |
+
MINDXTRAIN_TEACHER_BASE_URL=http://127.0.0.1:18093/v1
|
| 163 |
+
MINDXTRAIN_TEACHER_MODEL=bonsai-8b-q1_0
|
| 164 |
+
ENV
|
| 165 |
+
grep -n 'MINDXTRAIN_' /home/mindx/mindX/.env
|
| 166 |
+
```
|
| 167 |
+
|
| 168 |
+
`MINDXTRAIN_BACKEND=bankml` is deliberately not set: with it, the panel would use bankml even
|
| 169 |
+
where an `MINDXTRAIN_OPENAI_BASE_URL` is configured. Add it only if bankml should be the only
|
| 170 |
+
chat endpoint.
|
| 171 |
+
|
| 172 |
+
## 7. Restart mindX and confirm
|
| 173 |
+
|
| 174 |
+
Look at what is inside mindX's cgroup first; a restart ends everything listed there:
|
| 175 |
+
|
| 176 |
+
```bash
|
| 177 |
+
for p in $(cat /sys/fs/cgroup/system.slice/mindx.service/cgroup.procs); do ps -o pid=,args= -p $p; done | cut -c1-120
|
| 178 |
+
```
|
| 179 |
+
|
| 180 |
+
```bash
|
| 181 |
+
systemctl restart mindx && echo "restarted at $(date -u +%T); /health answers in 3-4 minutes"
|
| 182 |
+
```
|
| 183 |
+
|
| 184 |
+
```bash
|
| 185 |
+
until curl -s -o /dev/null -w '%{http_code}' 127.0.0.1:8000/health | grep -q 200; do sleep 15; done; echo "mindX is up"
|
| 186 |
+
```
|
| 187 |
+
|
| 188 |
+
```bash
|
| 189 |
+
tr '\0' '\n' < /proc/$(systemctl show -p MainPID --value mindx)/environ | grep MINDXTRAIN_
|
| 190 |
+
```
|
| 191 |
+
|
| 192 |
+
**Check:** the five `MINDXTRAIN_` lines from step 6.
|
| 193 |
+
|
| 194 |
+
## 8. Serving the next generation
|
| 195 |
+
|
| 196 |
+
When a run finishes, put it on bankml (merge, pin byte-identically, layer) in one command. Use
|
| 197 |
+
the run's config and tag:
|
| 198 |
+
|
| 199 |
+
```bash
|
| 200 |
+
cd /home/mindx/mindXtrain && sudo -u mindx BANKML_FORKS=/home/mindx/models/registry .venv/bin/mindxtrain serve /path/to/run.yaml --to bankml --tag mindx-gen82 --bankml-bin /home/mindx/bankml-0.3.6/bankml --bankml-system "You are mindX, generation 82." --bankml-param repeat_penalty=1.3 --bankml-param num_ctx=2048 --bankml-stop '<|im_end|>'
|
| 201 |
+
```
|
| 202 |
+
|
| 203 |
+
It prints the model's sha256; `curl -s 127.0.0.1:18093/api/tags` then lists `mindx-gen82`.
|
| 204 |
+
|
| 205 |
+
---
|
| 206 |
+
|
| 207 |
+
## Rollback
|
| 208 |
+
|
| 209 |
+
Each line undoes one step; run only the ones you need.
|
| 210 |
+
|
| 211 |
+
```bash
|
| 212 |
+
B=$(cat /root/LAST_BANKML036_BACKUP); echo "restoring from $B"
|
| 213 |
+
```
|
| 214 |
+
|
| 215 |
+
bankml back to 0.3.5 (steps 2–3):
|
| 216 |
+
|
| 217 |
+
```bash
|
| 218 |
+
cp $B/bankml.service /etc/systemd/system/bankml.service && systemctl daemon-reload && systemctl restart bankml && /home/mindx/bankml-0.3.5/bankml --version
|
| 219 |
+
```
|
| 220 |
+
|
| 221 |
+
gen39's layer (step 4), so `mindx-gen39` is the bare pin's alias again:
|
| 222 |
+
|
| 223 |
+
```bash
|
| 224 |
+
rm -f /home/mindx/models/registry/mindx-gen39.MODEL.json
|
| 225 |
+
```
|
| 226 |
+
|
| 227 |
+
mindX's environment (step 6), then restart mindX as in step 7:
|
| 228 |
+
|
| 229 |
+
```bash
|
| 230 |
+
cp $B/.env /home/mindx/mindX/.env && chown mindx:mindx /home/mindx/mindX/.env
|
| 231 |
+
```
|
| 232 |
+
|
| 233 |
+
mindXtrain back to the GitHub checkout (step 5):
|
| 234 |
+
|
| 235 |
+
```bash
|
| 236 |
+
cd /home/mindx/mindXtrain && sudo -u mindx git checkout -B main github/main && sudo -u mindx git stash pop || true
|
| 237 |
+
```
|
| 238 |
+
|
| 239 |
+
## Troubleshooting
|
| 240 |
+
|
| 241 |
+
| you see | it means | do |
|
| 242 |
+
|---|---|---|
|
| 243 |
+
| `cargo: command not found` | the build ran outside `mindx`'s login shell | keep the `sudo -u mindx bash -lc '…'` form |
|
| 244 |
+
| step 3 answers `400` | bankml is still 0.3.5 | `grep ExecStart /etc/systemd/system/bankml.service`, then `systemctl daemon-reload && systemctl restart bankml` |
|
| 245 |
+
| `bankml create … refuse` | the Modelfile asks for something bankml does not reproduce | the message names it; penalties need 0.3.6 |
|
| 246 |
+
| `git fetch hf` fails | the Hub was unreachable | `curl -sI https://huggingface.co/PYTHAI/mindXtrain`, then retry |
|
| 247 |
+
| `/health` never turns 200 | mindX did not start | `journalctl -u mindx -n 50 --no-pager` |
|
|
@@ -394,6 +394,25 @@ def _serve_openai_server(
|
|
| 394 |
)
|
| 395 |
|
| 396 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 397 |
def _serve_bankml(
|
| 398 |
cfg: XTrainConfig,
|
| 399 |
checkpoint: Path | None,
|
|
@@ -403,6 +422,9 @@ def _serve_bankml(
|
|
| 403 |
convert: bool,
|
| 404 |
register_as_fallback: bool,
|
| 405 |
mindx_base_url: str | None,
|
|
|
|
|
|
|
|
|
|
| 406 |
) -> None:
|
| 407 |
"""`serve --to bankml`: refuse what bankml cannot serve, else merge + `bankml create`.
|
| 408 |
|
|
@@ -432,6 +454,11 @@ def _serve_bankml(
|
|
| 432 |
console.print(f"[red]checkpoint not found:[/red] {ckpt}")
|
| 433 |
raise typer.Exit(code=1)
|
| 434 |
is_adapter = (ckpt / "adapter_config.json").exists() or not (ckpt / "config.json").exists()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 435 |
result = push_to_bankml(
|
| 436 |
base_model,
|
| 437 |
tag or run_name,
|
|
@@ -439,6 +466,9 @@ def _serve_bankml(
|
|
| 439 |
merged_dir=None if is_adapter else ckpt,
|
| 440 |
bankml_bin=bankml_bin,
|
| 441 |
convert=convert,
|
|
|
|
|
|
|
|
|
|
| 442 |
register_with_mindx=register_as_fallback,
|
| 443 |
mindx_base_url=mindx_base_url,
|
| 444 |
sink=lambda line: console.print(line, markup=False, highlight=False),
|
|
@@ -530,6 +560,19 @@ def serve(
|
|
| 530 |
None, "--bankml-bin",
|
| 531 |
help="Override the bankml binary path for `--to bankml` (defaults to PATH lookup).",
|
| 532 |
),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 533 |
bankml_convert: bool = typer.Option(
|
| 534 |
False, "--bankml-convert",
|
| 535 |
help="With `--to bankml`: run `bankml convert` (GGUF F16 + FORK.json pin) before "
|
|
@@ -579,6 +622,7 @@ def serve(
|
|
| 579 |
_serve_bankml(
|
| 580 |
cfg, checkpoint, tag=tag, bankml_bin=bankml_bin, convert=bankml_convert,
|
| 581 |
register_as_fallback=register_as_fallback, mindx_base_url=mindx_base_url,
|
|
|
|
| 582 |
)
|
| 583 |
return
|
| 584 |
|
|
|
|
| 394 |
)
|
| 395 |
|
| 396 |
|
| 397 |
+
def _parse_params(items: list[str]) -> dict[str, float | int | str]:
|
| 398 |
+
"""['repeat_penalty=1.3', 'num_ctx=2048'] → {'repeat_penalty': 1.3, 'num_ctx': 2048}."""
|
| 399 |
+
out: dict[str, float | int | str] = {}
|
| 400 |
+
for item in items:
|
| 401 |
+
key, sep, raw = item.partition("=")
|
| 402 |
+
if not sep or not key.strip():
|
| 403 |
+
msg = f"{item!r} is not KEY=VALUE"
|
| 404 |
+
raise ValueError(msg)
|
| 405 |
+
val: float | int | str = raw.strip()
|
| 406 |
+
for cast in (int, float):
|
| 407 |
+
try:
|
| 408 |
+
val = cast(raw)
|
| 409 |
+
break
|
| 410 |
+
except ValueError:
|
| 411 |
+
continue
|
| 412 |
+
out[key.strip()] = val
|
| 413 |
+
return out
|
| 414 |
+
|
| 415 |
+
|
| 416 |
def _serve_bankml(
|
| 417 |
cfg: XTrainConfig,
|
| 418 |
checkpoint: Path | None,
|
|
|
|
| 422 |
convert: bool,
|
| 423 |
register_as_fallback: bool,
|
| 424 |
mindx_base_url: str | None,
|
| 425 |
+
system: str | None = None,
|
| 426 |
+
params: list[str] | None = None,
|
| 427 |
+
stop: list[str] | None = None,
|
| 428 |
) -> None:
|
| 429 |
"""`serve --to bankml`: refuse what bankml cannot serve, else merge + `bankml create`.
|
| 430 |
|
|
|
|
| 454 |
console.print(f"[red]checkpoint not found:[/red] {ckpt}")
|
| 455 |
raise typer.Exit(code=1)
|
| 456 |
is_adapter = (ckpt / "adapter_config.json").exists() or not (ckpt / "config.json").exists()
|
| 457 |
+
try:
|
| 458 |
+
parsed = _parse_params(params or [])
|
| 459 |
+
except ValueError as exc:
|
| 460 |
+
console.print(f"[red]--bankml-param:[/red] {exc}")
|
| 461 |
+
raise typer.Exit(code=2) from exc
|
| 462 |
result = push_to_bankml(
|
| 463 |
base_model,
|
| 464 |
tag or run_name,
|
|
|
|
| 466 |
merged_dir=None if is_adapter else ckpt,
|
| 467 |
bankml_bin=bankml_bin,
|
| 468 |
convert=convert,
|
| 469 |
+
system=system,
|
| 470 |
+
params=parsed,
|
| 471 |
+
stop=list(stop or []),
|
| 472 |
register_with_mindx=register_as_fallback,
|
| 473 |
mindx_base_url=mindx_base_url,
|
| 474 |
sink=lambda line: console.print(line, markup=False, highlight=False),
|
|
|
|
| 560 |
None, "--bankml-bin",
|
| 561 |
help="Override the bankml binary path for `--to bankml` (defaults to PATH lookup).",
|
| 562 |
),
|
| 563 |
+
bankml_system: str = typer.Option(
|
| 564 |
+
None, "--bankml-system",
|
| 565 |
+
help="With `--to bankml`: the SYSTEM prompt layered over the pinned base (the persona).",
|
| 566 |
+
),
|
| 567 |
+
bankml_param: list[str] = typer.Option(
|
| 568 |
+
None, "--bankml-param",
|
| 569 |
+
help="With `--to bankml`: a Modelfile PARAMETER as KEY=VALUE, repeatable. A SmolLM2 / "
|
| 570 |
+
"mindx-genN generation wants `repeat_penalty=1.3` (bankml 0.3.6+) and `num_ctx=2048`.",
|
| 571 |
+
),
|
| 572 |
+
bankml_stop: list[str] = typer.Option(
|
| 573 |
+
None, "--bankml-stop",
|
| 574 |
+
help="With `--to bankml`: a stop string, repeatable (SmolLM2: '<|im_end|>').",
|
| 575 |
+
),
|
| 576 |
bankml_convert: bool = typer.Option(
|
| 577 |
False, "--bankml-convert",
|
| 578 |
help="With `--to bankml`: run `bankml convert` (GGUF F16 + FORK.json pin) before "
|
|
|
|
| 622 |
_serve_bankml(
|
| 623 |
cfg, checkpoint, tag=tag, bankml_bin=bankml_bin, convert=bankml_convert,
|
| 624 |
register_as_fallback=register_as_fallback, mindx_base_url=mindx_base_url,
|
| 625 |
+
system=bankml_system, params=bankml_param, stop=bankml_stop,
|
| 626 |
)
|
| 627 |
return
|
| 628 |
|
|
@@ -14,13 +14,20 @@ from the one described):
|
|
| 14 |
- **Taken:** `FROM` (the merged safetensors directory, which `bankml create` converts to GGUF F16
|
| 15 |
byte-identically to llama.cpp b11192, or a GGUF `bankml convert` already pinned), `SYSTEM`,
|
| 16 |
`MESSAGE`, `LICENSE`, `REQUIRES`, `PARAMETER` temperature / top_k / top_p / min_p / seed /
|
| 17 |
-
num_ctx / num_predict, and `stop`.
|
|
|
|
|
|
|
|
|
|
| 18 |
- **Refused:** `ADAPTER` (merge first — this module does it for you given `adapter_dir`), a
|
| 19 |
-
`TEMPLATE` other than the base's own, the penalties
|
| 20 |
-
|
| 21 |
num_keep, draft_num_predict), and any architecture other than Llama (SmolLM2 = mindx-genN):
|
| 22 |
bankml converts Llama safetensors only; Qwen3 it serves only as a pinned Q1_0/Q2_0_g64 GGUF.
|
| 23 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
`bankml create` / `bankml convert` arrive in bankml 0.3.5. An older binary is detected from
|
| 25 |
`bankml --help` (the verbs it lists) and reported as `status="bankml_too_old"`, never raised.
|
| 26 |
|
|
@@ -49,6 +56,7 @@ BANKML_PARAMS: frozenset[str] = frozenset(
|
|
| 49 |
)
|
| 50 |
|
| 51 |
_PENALTY = frozenset({"repeat_penalty", "presence_penalty", "frequency_penalty", "repeat_last_n"})
|
|
|
|
| 52 |
_MIROSTAT = frozenset({"mirostat", "mirostat_tau", "mirostat_eta"})
|
| 53 |
_RESOURCE = frozenset({"num_gpu", "num_thread", "num_batch", "num_keep", "draft_num_predict"})
|
| 54 |
|
|
@@ -63,10 +71,21 @@ PushStatus = Literal[
|
|
| 63 |
]
|
| 64 |
|
| 65 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
def _param_refusal(name: str) -> str:
|
| 67 |
if name in _PENALTY:
|
| 68 |
-
return (f"PARAMETER {name}: penalties are
|
| 69 |
-
"
|
|
|
|
| 70 |
if name in _MIROSTAT:
|
| 71 |
return f"PARAMETER {name}: mirostat is not reproduced by bankml; leave it out"
|
| 72 |
if name == "typical_p":
|
|
@@ -142,12 +161,15 @@ class BankmlSanitizeResult:
|
|
| 142 |
return not self.refusals
|
| 143 |
|
| 144 |
|
| 145 |
-
def bankml_sanitize(
|
|
|
|
|
|
|
| 146 |
"""Check a ModelfileSpec against the subset `bankml create` reproduces.
|
| 147 |
|
| 148 |
Refuses (and names) every instruction bankml would not honour; it never drops one, because a
|
| 149 |
silently dropped PARAMETER serves a model that answers differently from the spec. A
|
| 150 |
`TEMPLATE` passes only when it equals `base_template` (the base GGUF's own chat template).
|
|
|
|
| 151 |
"""
|
| 152 |
refusals: list[str] = []
|
| 153 |
if spec.adapter:
|
|
@@ -157,8 +179,9 @@ def bankml_sanitize(spec: ModelfileSpec, *, base_template: str | None = None) ->
|
|
| 157 |
refusals.append("TEMPLATE: bankml renders the base model's own chat template "
|
| 158 |
"(byte-identical to llama.cpp); a different template is not reproduced — "
|
| 159 |
"leave TEMPLATE out")
|
|
|
|
| 160 |
for name in sorted(spec.parameters):
|
| 161 |
-
if name not in
|
| 162 |
refusals.append(_param_refusal(name))
|
| 163 |
for label, value in (("SYSTEM", spec.system), ("TEMPLATE", spec.template),
|
| 164 |
("LICENSE", spec.license)):
|
|
@@ -296,7 +319,7 @@ def push_to_bankml(
|
|
| 296 |
# 1. the subset — checked before anything expensive runs
|
| 297 |
spec = ModelfileSpec(from_model="<merged>", system=system or "", parameters=dict(params or {}),
|
| 298 |
stop=list(stop or []))
|
| 299 |
-
checked = bankml_sanitize(spec)
|
| 300 |
if not checked.ok:
|
| 301 |
for r in checked.refusals:
|
| 302 |
emit(f"[push-bankml] refuse: {r}")
|
|
@@ -315,6 +338,13 @@ def push_to_bankml(
|
|
| 315 |
return done("bankml_too_old", bankml_version=caps.version, reason=(
|
| 316 |
f"bankml {caps.version or '(unknown version)'} has no `{'`/`'.join(missing)}` verb "
|
| 317 |
f"(they arrive in bankml 0.3.5); upgrade from {BANKML_URL}"))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 318 |
emit(f"[push-bankml] bankml {caps.version} at {caps.binary}")
|
| 319 |
|
| 320 |
work = Path(work_dir) if work_dir else (
|
|
@@ -349,8 +379,9 @@ def push_to_bankml(
|
|
| 349 |
gguf: Path | None = None
|
| 350 |
model_sha = ""
|
| 351 |
try:
|
| 352 |
-
# 4. optional explicit conversion
|
| 353 |
-
|
|
|
|
| 354 |
if convert:
|
| 355 |
registry.mkdir(parents=True, exist_ok=True)
|
| 356 |
gguf = registry / f"{name}-base-F16.gguf"
|
|
@@ -419,6 +450,21 @@ def push_to_bankml(
|
|
| 419 |
mindx_fallback_swap=swap)
|
| 420 |
|
| 421 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 422 |
def _last_line(text: str) -> str:
|
| 423 |
rows = [r.strip() for r in text.strip().splitlines() if r.strip()]
|
| 424 |
return rows[-1] if rows else ""
|
|
@@ -427,6 +473,7 @@ def _last_line(text: str) -> str:
|
|
| 427 |
__all__ = [
|
| 428 |
"BANKML_CONVERT_ARCHS",
|
| 429 |
"BANKML_PARAMS",
|
|
|
|
| 430 |
"BankmlCapabilities",
|
| 431 |
"BankmlPushResult",
|
| 432 |
"BankmlSanitizeResult",
|
|
@@ -435,5 +482,7 @@ __all__ = [
|
|
| 435 |
"bankml_sanitize",
|
| 436 |
"base_family_refusal",
|
| 437 |
"check_merged_arch",
|
|
|
|
| 438 |
"push_to_bankml",
|
|
|
|
| 439 |
]
|
|
|
|
| 14 |
- **Taken:** `FROM` (the merged safetensors directory, which `bankml create` converts to GGUF F16
|
| 15 |
byte-identically to llama.cpp b11192, or a GGUF `bankml convert` already pinned), `SYSTEM`,
|
| 16 |
`MESSAGE`, `LICENSE`, `REQUIRES`, `PARAMETER` temperature / top_k / top_p / min_p / seed /
|
| 17 |
+
num_ctx / num_predict, and `stop`; from bankml 0.3.6 (its O2) also the penalties —
|
| 18 |
+
repeat_penalty, repeat_last_n, presence_penalty, frequency_penalty — which it applies token
|
| 19 |
+
for token as llama-server b11192 does. A small generation (SmolLM2 / mindx-genN) needs
|
| 20 |
+
`repeat_penalty` to answer in words rather than runs of `,`.
|
| 21 |
- **Refused:** `ADAPTER` (merge first — this module does it for you given `adapter_dir`), a
|
| 22 |
+
`TEMPLATE` other than the base's own, the penalties on a bankml older than 0.3.6,
|
| 23 |
+
mirostat*, typical_p, resource options (num_gpu, num_thread, num_batch,
|
| 24 |
num_keep, draft_num_predict), and any architecture other than Llama (SmolLM2 = mindx-genN):
|
| 25 |
bankml converts Llama safetensors only; Qwen3 it serves only as a pinned Q1_0/Q2_0_g64 GGUF.
|
| 26 |
|
| 27 |
+
The merged directory is converted through a link named after the tag: llama.cpp's converter
|
| 28 |
+
names a model after the directory it reads (`merged` → general.name "Merged"), and only the
|
| 29 |
+
tag's name gives the GGUF byte-identical to the one mindX pins (gen39 → 6b64c748…).
|
| 30 |
+
|
| 31 |
`bankml create` / `bankml convert` arrive in bankml 0.3.5. An older binary is detected from
|
| 32 |
`bankml --help` (the verbs it lists) and reported as `status="bankml_too_old"`, never raised.
|
| 33 |
|
|
|
|
| 56 |
)
|
| 57 |
|
| 58 |
_PENALTY = frozenset({"repeat_penalty", "presence_penalty", "frequency_penalty", "repeat_last_n"})
|
| 59 |
+
PENALTIES_SINCE = (0, 3, 6) # bankml O2: penalties token-identical to llama-server b11192
|
| 60 |
_MIROSTAT = frozenset({"mirostat", "mirostat_tau", "mirostat_eta"})
|
| 61 |
_RESOURCE = frozenset({"num_gpu", "num_thread", "num_batch", "num_keep", "draft_num_predict"})
|
| 62 |
|
|
|
|
| 71 |
]
|
| 72 |
|
| 73 |
|
| 74 |
+
def version_tuple(version: str) -> tuple[int, ...]:
|
| 75 |
+
"""'0.3.6' → (0, 3, 6); '' or unparseable → () (treated as older than any release)."""
|
| 76 |
+
m = re.match(r"(\d+)\.(\d+)\.(\d+)", version or "")
|
| 77 |
+
return tuple(int(g) for g in m.groups()) if m else ()
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
def penalties_supported(version: str) -> bool:
|
| 81 |
+
return version_tuple(version) >= PENALTIES_SINCE
|
| 82 |
+
|
| 83 |
+
|
| 84 |
def _param_refusal(name: str) -> str:
|
| 85 |
if name in _PENALTY:
|
| 86 |
+
return (f"PARAMETER {name}: penalties are reproduced from bankml "
|
| 87 |
+
f"{'.'.join(map(str, PENALTIES_SINCE))} (O2); this bankml predates it — upgrade, "
|
| 88 |
+
"or leave it out")
|
| 89 |
if name in _MIROSTAT:
|
| 90 |
return f"PARAMETER {name}: mirostat is not reproduced by bankml; leave it out"
|
| 91 |
if name == "typical_p":
|
|
|
|
| 161 |
return not self.refusals
|
| 162 |
|
| 163 |
|
| 164 |
+
def bankml_sanitize(
|
| 165 |
+
spec: ModelfileSpec, *, base_template: str | None = None, penalties: bool = False,
|
| 166 |
+
) -> BankmlSanitizeResult:
|
| 167 |
"""Check a ModelfileSpec against the subset `bankml create` reproduces.
|
| 168 |
|
| 169 |
Refuses (and names) every instruction bankml would not honour; it never drops one, because a
|
| 170 |
silently dropped PARAMETER serves a model that answers differently from the spec. A
|
| 171 |
`TEMPLATE` passes only when it equals `base_template` (the base GGUF's own chat template).
|
| 172 |
+
`penalties=True` admits the penalties (bankml 0.3.6+; see `penalties_supported`).
|
| 173 |
"""
|
| 174 |
refusals: list[str] = []
|
| 175 |
if spec.adapter:
|
|
|
|
| 179 |
refusals.append("TEMPLATE: bankml renders the base model's own chat template "
|
| 180 |
"(byte-identical to llama.cpp); a different template is not reproduced — "
|
| 181 |
"leave TEMPLATE out")
|
| 182 |
+
allowed = BANKML_PARAMS | _PENALTY if penalties else BANKML_PARAMS
|
| 183 |
for name in sorted(spec.parameters):
|
| 184 |
+
if name not in allowed:
|
| 185 |
refusals.append(_param_refusal(name))
|
| 186 |
for label, value in (("SYSTEM", spec.system), ("TEMPLATE", spec.template),
|
| 187 |
("LICENSE", spec.license)):
|
|
|
|
| 319 |
# 1. the subset — checked before anything expensive runs
|
| 320 |
spec = ModelfileSpec(from_model="<merged>", system=system or "", parameters=dict(params or {}),
|
| 321 |
stop=list(stop or []))
|
| 322 |
+
checked = bankml_sanitize(spec, penalties=True) # penalties are gated on the version below
|
| 323 |
if not checked.ok:
|
| 324 |
for r in checked.refusals:
|
| 325 |
emit(f"[push-bankml] refuse: {r}")
|
|
|
|
| 338 |
return done("bankml_too_old", bankml_version=caps.version, reason=(
|
| 339 |
f"bankml {caps.version or '(unknown version)'} has no `{'`/`'.join(missing)}` verb "
|
| 340 |
f"(they arrive in bankml 0.3.5); upgrade from {BANKML_URL}"))
|
| 341 |
+
asked = sorted(_PENALTY & set(spec.parameters))
|
| 342 |
+
if asked and not penalties_supported(caps.version):
|
| 343 |
+
refusals = tuple(_param_refusal(n) for n in asked)
|
| 344 |
+
for r in refusals:
|
| 345 |
+
emit(f"[push-bankml] refuse: {r}")
|
| 346 |
+
return done("refused", bankml_version=caps.version, refusals=refusals,
|
| 347 |
+
reason=f"bankml {caps.version or '(unknown version)'} does not reproduce penalties")
|
| 348 |
emit(f"[push-bankml] bankml {caps.version} at {caps.binary}")
|
| 349 |
|
| 350 |
work = Path(work_dir) if work_dir else (
|
|
|
|
| 379 |
gguf: Path | None = None
|
| 380 |
model_sha = ""
|
| 381 |
try:
|
| 382 |
+
# 4. optional explicit conversion — through a link named after the tag, so the GGUF's
|
| 383 |
+
# general.name (and so its bytes and sha256) is the one mindX pins
|
| 384 |
+
source = str(_named_source(merged, work / name, name))
|
| 385 |
if convert:
|
| 386 |
registry.mkdir(parents=True, exist_ok=True)
|
| 387 |
gguf = registry / f"{name}-base-F16.gguf"
|
|
|
|
| 450 |
mindx_fallback_swap=swap)
|
| 451 |
|
| 452 |
|
| 453 |
+
def _named_source(merged: Path, at: Path, name: str) -> Path:
|
| 454 |
+
"""A link `at/<name>` → `merged`, unless `merged` is already called `name`."""
|
| 455 |
+
target = merged.resolve()
|
| 456 |
+
if target.name == name:
|
| 457 |
+
return target
|
| 458 |
+
at.mkdir(parents=True, exist_ok=True)
|
| 459 |
+
link = at / name
|
| 460 |
+
if link.is_symlink():
|
| 461 |
+
link.unlink()
|
| 462 |
+
elif link.exists():
|
| 463 |
+
return link # a real directory by that name: the caller put it there
|
| 464 |
+
link.symlink_to(target, target_is_directory=True)
|
| 465 |
+
return link
|
| 466 |
+
|
| 467 |
+
|
| 468 |
def _last_line(text: str) -> str:
|
| 469 |
rows = [r.strip() for r in text.strip().splitlines() if r.strip()]
|
| 470 |
return rows[-1] if rows else ""
|
|
|
|
| 473 |
__all__ = [
|
| 474 |
"BANKML_CONVERT_ARCHS",
|
| 475 |
"BANKML_PARAMS",
|
| 476 |
+
"PENALTIES_SINCE",
|
| 477 |
"BankmlCapabilities",
|
| 478 |
"BankmlPushResult",
|
| 479 |
"BankmlSanitizeResult",
|
|
|
|
| 482 |
"bankml_sanitize",
|
| 483 |
"base_family_refusal",
|
| 484 |
"check_merged_arch",
|
| 485 |
+
"penalties_supported",
|
| 486 |
"push_to_bankml",
|
| 487 |
+
"version_tuple",
|
| 488 |
]
|
|
@@ -24,6 +24,16 @@ _CHOICE_RE = re.compile(r"\b(A|B|TIE)\b", re.IGNORECASE)
|
|
| 24 |
_DEFAULT_JUDGE = "llama3.2"
|
| 25 |
|
| 26 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
class EvalScore(BaseModel):
|
| 28 |
"""A single evaluation result, score normalized to [0, 1]."""
|
| 29 |
|
|
@@ -79,8 +89,8 @@ class SemanticSimilarityEvaluator:
|
|
| 79 |
class CorrectnessEvaluator:
|
| 80 |
"""LLM-as-judge: is RESPONSE correct/faithful vs REFERENCE for the QUERY?"""
|
| 81 |
|
| 82 |
-
def __init__(self, model: str =
|
| 83 |
-
self.model, self.base_url, self.threshold = model, base_url, threshold
|
| 84 |
|
| 85 |
def evaluate(self, query: str, response: str, reference: str) -> EvalScore:
|
| 86 |
system = (
|
|
@@ -97,8 +107,8 @@ class CorrectnessEvaluator:
|
|
| 97 |
class PairwiseEvaluator:
|
| 98 |
"""Which of two responses better matches the persona/reference? B (after) vs A (before)."""
|
| 99 |
|
| 100 |
-
def __init__(self, model: str =
|
| 101 |
-
self.model, self.base_url = model, base_url
|
| 102 |
|
| 103 |
def evaluate(self, query: str, response_a: str, response_b: str, *, reference: str = "") -> EvalScore:
|
| 104 |
ref = f" The target voice/reference is: {reference}." if reference else ""
|
|
@@ -131,8 +141,8 @@ class PairwiseEvaluator:
|
|
| 131 |
class GuidelineEvaluator:
|
| 132 |
"""LLM-as-judge: does RESPONSE comply with the GUIDELINES/rubric?"""
|
| 133 |
|
| 134 |
-
def __init__(self, model: str =
|
| 135 |
-
self.model, self.base_url, self.threshold = model, base_url, threshold
|
| 136 |
|
| 137 |
def evaluate(self, response: str, guidelines: str) -> EvalScore:
|
| 138 |
system = (
|
|
|
|
| 24 |
_DEFAULT_JUDGE = "llama3.2"
|
| 25 |
|
| 26 |
|
| 27 |
+
def default_judge_model() -> str:
|
| 28 |
+
"""`MINDXTRAIN_JUDGE_MODEL`, else the panel's chat model, else llama3.2 — read at call
|
| 29 |
+
time so a bankml node judges with a model it actually serves."""
|
| 30 |
+
import os
|
| 31 |
+
|
| 32 |
+
from mindxtrain.governance.panel import default_chat_model
|
| 33 |
+
|
| 34 |
+
return os.environ.get("MINDXTRAIN_JUDGE_MODEL") or default_chat_model()
|
| 35 |
+
|
| 36 |
+
|
| 37 |
class EvalScore(BaseModel):
|
| 38 |
"""A single evaluation result, score normalized to [0, 1]."""
|
| 39 |
|
|
|
|
| 89 |
class CorrectnessEvaluator:
|
| 90 |
"""LLM-as-judge: is RESPONSE correct/faithful vs REFERENCE for the QUERY?"""
|
| 91 |
|
| 92 |
+
def __init__(self, model: str | None = None, base_url: str | None = None, threshold: float = 0.6) -> None:
|
| 93 |
+
self.model, self.base_url, self.threshold = model or default_judge_model(), base_url, threshold
|
| 94 |
|
| 95 |
def evaluate(self, query: str, response: str, reference: str) -> EvalScore:
|
| 96 |
system = (
|
|
|
|
| 107 |
class PairwiseEvaluator:
|
| 108 |
"""Which of two responses better matches the persona/reference? B (after) vs A (before)."""
|
| 109 |
|
| 110 |
+
def __init__(self, model: str | None = None, base_url: str | None = None) -> None:
|
| 111 |
+
self.model, self.base_url = model or default_judge_model(), base_url
|
| 112 |
|
| 113 |
def evaluate(self, query: str, response_a: str, response_b: str, *, reference: str = "") -> EvalScore:
|
| 114 |
ref = f" The target voice/reference is: {reference}." if reference else ""
|
|
|
|
| 141 |
class GuidelineEvaluator:
|
| 142 |
"""LLM-as-judge: does RESPONSE comply with the GUIDELINES/rubric?"""
|
| 143 |
|
| 144 |
+
def __init__(self, model: str | None = None, base_url: str | None = None, threshold: float = 0.6) -> None:
|
| 145 |
+
self.model, self.base_url, self.threshold = model or default_judge_model(), base_url, threshold
|
| 146 |
|
| 147 |
def evaluate(self, response: str, guidelines: str) -> EvalScore:
|
| 148 |
system = (
|
|
@@ -100,7 +100,7 @@ def evaluate_classroom(
|
|
| 100 |
|
| 101 |
# Pairwise: judge each inquiry (before vs after) toward the persona, else use the
|
| 102 |
# imprint delta sign as the signal.
|
| 103 |
-
if use_judge
|
| 104 |
from mindxtrain.eval.llama_evals import PairwiseEvaluator
|
| 105 |
|
| 106 |
ev = PairwiseEvaluator(model=model, base_url=base_url)
|
|
|
|
| 100 |
|
| 101 |
# Pairwise: judge each inquiry (before vs after) toward the persona, else use the
|
| 102 |
# imprint delta sign as the signal.
|
| 103 |
+
if use_judge: # no model named → MINDXTRAIN_JUDGE_MODEL / MINDXTRAIN_CHAT_MODEL / llama3.2
|
| 104 |
from mindxtrain.eval.llama_evals import PairwiseEvaluator
|
| 105 |
|
| 106 |
ev = PairwiseEvaluator(model=model, base_url=base_url)
|
|
@@ -12,6 +12,7 @@ or is recorded as a reject (dojo, which forbids abstention). Pure stdlib + httpx
|
|
| 12 |
|
| 13 |
from __future__ import annotations
|
| 14 |
|
|
|
|
| 15 |
import os
|
| 16 |
import re
|
| 17 |
from collections.abc import Callable
|
|
@@ -33,6 +34,31 @@ ROLE_STANCE: dict[Role, str] = {
|
|
| 33 |
}
|
| 34 |
|
| 35 |
_DEFAULT_MODEL = "llama3.2"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
_VERDICT_RE = re.compile(r"verdict\s*[:\-]?\s*(approve|reject|abstain)", re.IGNORECASE)
|
| 37 |
|
| 38 |
|
|
@@ -71,8 +97,12 @@ def chat_once(
|
|
| 71 |
temperature: float = 0.0,
|
| 72 |
max_tokens: int = 256,
|
| 73 |
timeout_s: float = 60.0,
|
|
|
|
| 74 |
) -> str:
|
| 75 |
-
"""One non-streaming OpenAI-compatible chat completion; returns the content.
|
|
|
|
|
|
|
|
|
|
| 76 |
base = resolve_chat_base_url(base_url)
|
| 77 |
key = api_key or os.environ.get("MINDXTRAIN_OPENAI_API_KEY", "")
|
| 78 |
headers = {"Content-Type": "application/json"}
|
|
@@ -82,6 +112,7 @@ def chat_once(
|
|
| 82 |
resp = client.post(
|
| 83 |
f"{base}/chat/completions",
|
| 84 |
json={
|
|
|
|
| 85 |
"model": model,
|
| 86 |
"messages": messages,
|
| 87 |
"temperature": temperature,
|
|
@@ -132,7 +163,7 @@ def deliberate(
|
|
| 132 |
*,
|
| 133 |
base_url: str | None = None,
|
| 134 |
allow_abstain: bool = True,
|
| 135 |
-
default_model: str =
|
| 136 |
**chat_kw: object,
|
| 137 |
) -> Deliberation:
|
| 138 |
"""Query `member`'s model from its role stance and parse a vote + rationale."""
|
|
@@ -143,7 +174,7 @@ def deliberate(
|
|
| 143 |
f"motion. {stance} Give a one-sentence rationale, then on the final line write "
|
| 144 |
f"exactly 'VERDICT: <{verdicts}>'."
|
| 145 |
)
|
| 146 |
-
model = member.model or default_model
|
| 147 |
# Reasoning ("thinking") models spend tokens before the verdict — give enough
|
| 148 |
# room and time by default so the `VERDICT:` line is actually reached (concurrent
|
| 149 |
# members also queue inside a single-GPU backend). Callers can override.
|
|
@@ -168,7 +199,7 @@ def deliberate(
|
|
| 168 |
|
| 169 |
|
| 170 |
def model_ballot(
|
| 171 |
-
*, base_url: str | None = None, default_model: str =
|
| 172 |
) -> Callable[[Member, str], Vote]:
|
| 173 |
"""A boardroom ballot backed by real models (members may abstain)."""
|
| 174 |
|
|
@@ -182,7 +213,7 @@ def model_ballot(
|
|
| 182 |
|
| 183 |
|
| 184 |
def model_judge_ballot(
|
| 185 |
-
*, base_url: str | None = None, default_model: str =
|
| 186 |
) -> Callable[[Member, str], JudgeVote]:
|
| 187 |
"""A dojo ballot backed by real models (judges must approve/reject)."""
|
| 188 |
|
|
@@ -200,6 +231,8 @@ __all__ = [
|
|
| 200 |
"ROLE_STANCE",
|
| 201 |
"Deliberation",
|
| 202 |
"chat_once",
|
|
|
|
|
|
|
| 203 |
"deliberate",
|
| 204 |
"model_ballot",
|
| 205 |
"model_judge_ballot",
|
|
|
|
| 12 |
|
| 13 |
from __future__ import annotations
|
| 14 |
|
| 15 |
+
import json
|
| 16 |
import os
|
| 17 |
import re
|
| 18 |
from collections.abc import Callable
|
|
|
|
| 34 |
}
|
| 35 |
|
| 36 |
_DEFAULT_MODEL = "llama3.2"
|
| 37 |
+
|
| 38 |
+
|
| 39 |
+
def default_chat_model() -> str:
|
| 40 |
+
"""The model a member or judge uses when none is named: `MINDXTRAIN_CHAT_MODEL`, else
|
| 41 |
+
llama3.2. Read at call time, so a node pointed at bankml (which serves no llama3.2)
|
| 42 |
+
names its own model without a code change."""
|
| 43 |
+
return os.environ.get("MINDXTRAIN_CHAT_MODEL") or _DEFAULT_MODEL
|
| 44 |
+
|
| 45 |
+
|
| 46 |
+
def chat_options() -> dict[str, object]:
|
| 47 |
+
"""Extra sampling fields for every chat body, from `MINDXTRAIN_CHAT_OPTIONS` (a JSON
|
| 48 |
+
object), e.g. `{"repeat_penalty": 1.3}` for a small generation on bankml 0.3.6+. Empty
|
| 49 |
+
when unset: a server that refuses fields it does not know is never sent any."""
|
| 50 |
+
raw = os.environ.get("MINDXTRAIN_CHAT_OPTIONS")
|
| 51 |
+
if not raw:
|
| 52 |
+
return {}
|
| 53 |
+
try:
|
| 54 |
+
val = json.loads(raw)
|
| 55 |
+
except json.JSONDecodeError as exc:
|
| 56 |
+
msg = f"MINDXTRAIN_CHAT_OPTIONS is not JSON: {exc}"
|
| 57 |
+
raise ValueError(msg) from exc
|
| 58 |
+
if not isinstance(val, dict):
|
| 59 |
+
msg = "MINDXTRAIN_CHAT_OPTIONS must be a JSON object"
|
| 60 |
+
raise ValueError(msg)
|
| 61 |
+
return val
|
| 62 |
_VERDICT_RE = re.compile(r"verdict\s*[:\-]?\s*(approve|reject|abstain)", re.IGNORECASE)
|
| 63 |
|
| 64 |
|
|
|
|
| 97 |
temperature: float = 0.0,
|
| 98 |
max_tokens: int = 256,
|
| 99 |
timeout_s: float = 60.0,
|
| 100 |
+
options: dict[str, object] | None = None,
|
| 101 |
) -> str:
|
| 102 |
+
"""One non-streaming OpenAI-compatible chat completion; returns the content.
|
| 103 |
+
|
| 104 |
+
`options` (else `MINDXTRAIN_CHAT_OPTIONS`) adds sampling fields beyond OpenAI's — bankml
|
| 105 |
+
0.3.6+ honours `repeat_penalty`, `repeat_last_n`, `presence_penalty`, `frequency_penalty`."""
|
| 106 |
base = resolve_chat_base_url(base_url)
|
| 107 |
key = api_key or os.environ.get("MINDXTRAIN_OPENAI_API_KEY", "")
|
| 108 |
headers = {"Content-Type": "application/json"}
|
|
|
|
| 112 |
resp = client.post(
|
| 113 |
f"{base}/chat/completions",
|
| 114 |
json={
|
| 115 |
+
**(chat_options() if options is None else options),
|
| 116 |
"model": model,
|
| 117 |
"messages": messages,
|
| 118 |
"temperature": temperature,
|
|
|
|
| 163 |
*,
|
| 164 |
base_url: str | None = None,
|
| 165 |
allow_abstain: bool = True,
|
| 166 |
+
default_model: str | None = None,
|
| 167 |
**chat_kw: object,
|
| 168 |
) -> Deliberation:
|
| 169 |
"""Query `member`'s model from its role stance and parse a vote + rationale."""
|
|
|
|
| 174 |
f"motion. {stance} Give a one-sentence rationale, then on the final line write "
|
| 175 |
f"exactly 'VERDICT: <{verdicts}>'."
|
| 176 |
)
|
| 177 |
+
model = member.model or default_model or default_chat_model()
|
| 178 |
# Reasoning ("thinking") models spend tokens before the verdict — give enough
|
| 179 |
# room and time by default so the `VERDICT:` line is actually reached (concurrent
|
| 180 |
# members also queue inside a single-GPU backend). Callers can override.
|
|
|
|
| 199 |
|
| 200 |
|
| 201 |
def model_ballot(
|
| 202 |
+
*, base_url: str | None = None, default_model: str | None = None, **chat_kw: object,
|
| 203 |
) -> Callable[[Member, str], Vote]:
|
| 204 |
"""A boardroom ballot backed by real models (members may abstain)."""
|
| 205 |
|
|
|
|
| 213 |
|
| 214 |
|
| 215 |
def model_judge_ballot(
|
| 216 |
+
*, base_url: str | None = None, default_model: str | None = None, **chat_kw: object,
|
| 217 |
) -> Callable[[Member, str], JudgeVote]:
|
| 218 |
"""A dojo ballot backed by real models (judges must approve/reject)."""
|
| 219 |
|
|
|
|
| 231 |
"ROLE_STANCE",
|
| 232 |
"Deliberation",
|
| 233 |
"chat_once",
|
| 234 |
+
"chat_options",
|
| 235 |
+
"default_chat_model",
|
| 236 |
"deliberate",
|
| 237 |
"model_ballot",
|
| 238 |
"model_judge_ballot",
|
|
@@ -13,12 +13,19 @@ What this backend adds over `openai_compat`:
|
|
| 13 |
in `ChatResponse.receipt`; streamed it arrives as one extra `data: {"bankml_receipt": ...}`
|
| 14 |
event before `data: [DONE]`. Either way the latest receipt is also kept on `last_receipt`.
|
| 15 |
- **Refusals are typed, never retried.** bankml answers HTTP 400 with a plain-text reason when a
|
| 16 |
-
request asks for something its verified forward pass does not reproduce (
|
| 17 |
-
|
| 18 |
That becomes `BankmlRefusal(reason)`. The request is never re-sent with altered parameters:
|
| 19 |
a changed request would be a different, unreceipted question.
|
| 20 |
|
| 21 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
"""
|
| 23 |
|
| 24 |
from __future__ import annotations
|
|
@@ -91,6 +98,17 @@ def raise_for_bankml(resp: httpx.Response) -> None:
|
|
| 91 |
raise BankmlError(refusal_reason(resp), resp.status_code)
|
| 92 |
|
| 93 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 94 |
@register_backend("bankml")
|
| 95 |
class BankmlBackend(OpenAICompatBackend):
|
| 96 |
"""OpenAI-compatible client for `bankml serve --native`, keeping receipts and refusals."""
|
|
@@ -103,12 +121,16 @@ class BankmlBackend(OpenAICompatBackend):
|
|
| 103 |
timeout_s: float = 300.0,
|
| 104 |
*,
|
| 105 |
seed: int | None = None,
|
|
|
|
| 106 |
transport: httpx.AsyncBaseTransport | None = None,
|
| 107 |
) -> None:
|
| 108 |
super().__init__(base_url=base_url or bankml_base_url(), api_key=None, timeout_s=timeout_s)
|
| 109 |
# `api_key or env` in the parent would pick up an OpenAI key; bankml is loopback, keyless.
|
| 110 |
self.api_key = ""
|
| 111 |
self.seed = seed
|
|
|
|
|
|
|
|
|
|
| 112 |
self._transport = transport
|
| 113 |
self.last_receipt: dict[str, Any] | None = None
|
| 114 |
|
|
@@ -116,7 +138,7 @@ class BankmlBackend(OpenAICompatBackend):
|
|
| 116 |
return httpx.AsyncClient(timeout=self.timeout_s, transport=self._transport)
|
| 117 |
|
| 118 |
def _payload(self, request: ChatRequest, *, stream: bool) -> dict[str, object]:
|
| 119 |
-
payload = super()._payload(request, stream=stream)
|
| 120 |
if self.seed is not None:
|
| 121 |
payload["seed"] = self.seed
|
| 122 |
return payload
|
|
|
|
| 13 |
in `ChatResponse.receipt`; streamed it arrives as one extra `data: {"bankml_receipt": ...}`
|
| 14 |
event before `data: [DONE]`. Either way the latest receipt is also kept on `last_receipt`.
|
| 15 |
- **Refusals are typed, never retried.** bankml answers HTTP 400 with a plain-text reason when a
|
| 16 |
+
request asks for something its verified forward pass does not reproduce (penalties before
|
| 17 |
+
0.3.6, mirostat, typical_p, tools, an unknown architecture, Q8_0 / Q4_K / BF16).
|
| 18 |
That becomes `BankmlRefusal(reason)`. The request is never re-sent with altered parameters:
|
| 19 |
a changed request would be a different, unreceipted question.
|
| 20 |
|
| 21 |
+
- **Sampling options, only when set.** bankml 0.3.6 reproduces `repeat_penalty`,
|
| 22 |
+
`repeat_last_n`, `presence_penalty` and `frequency_penalty` as llama-server b11192 does; a
|
| 23 |
+
small generation (mindx-genN) needs `repeat_penalty` to answer in words. They are sent from
|
| 24 |
+
`options=` or `MINDXTRAIN_BANKML_OPTIONS`, never by default, so an older bankml is never sent
|
| 25 |
+
a field it refuses.
|
| 26 |
+
|
| 27 |
+
Env: `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`),
|
| 28 |
+
`MINDXTRAIN_BANKML_OPTIONS` (a JSON object, e.g. `{"repeat_penalty": 1.3}`).
|
| 29 |
"""
|
| 30 |
|
| 31 |
from __future__ import annotations
|
|
|
|
| 98 |
raise BankmlError(refusal_reason(resp), resp.status_code)
|
| 99 |
|
| 100 |
|
| 101 |
+
def _env_options() -> dict[str, object]:
|
| 102 |
+
raw = os.environ.get("MINDXTRAIN_BANKML_OPTIONS")
|
| 103 |
+
if not raw:
|
| 104 |
+
return {}
|
| 105 |
+
val = json.loads(raw)
|
| 106 |
+
if not isinstance(val, dict):
|
| 107 |
+
msg = "MINDXTRAIN_BANKML_OPTIONS must be a JSON object"
|
| 108 |
+
raise ValueError(msg)
|
| 109 |
+
return val
|
| 110 |
+
|
| 111 |
+
|
| 112 |
@register_backend("bankml")
|
| 113 |
class BankmlBackend(OpenAICompatBackend):
|
| 114 |
"""OpenAI-compatible client for `bankml serve --native`, keeping receipts and refusals."""
|
|
|
|
| 121 |
timeout_s: float = 300.0,
|
| 122 |
*,
|
| 123 |
seed: int | None = None,
|
| 124 |
+
options: dict[str, object] | None = None,
|
| 125 |
transport: httpx.AsyncBaseTransport | None = None,
|
| 126 |
) -> None:
|
| 127 |
super().__init__(base_url=base_url or bankml_base_url(), api_key=None, timeout_s=timeout_s)
|
| 128 |
# `api_key or env` in the parent would pick up an OpenAI key; bankml is loopback, keyless.
|
| 129 |
self.api_key = ""
|
| 130 |
self.seed = seed
|
| 131 |
+
# Sampling fields beyond OpenAI's that bankml honours (repeat_penalty & co. from 0.3.6),
|
| 132 |
+
# sent only when set: from MINDXTRAIN_BANKML_OPTIONS (a JSON object), then `options`.
|
| 133 |
+
self.options: dict[str, object] = {**_env_options(), **(options or {})}
|
| 134 |
self._transport = transport
|
| 135 |
self.last_receipt: dict[str, Any] | None = None
|
| 136 |
|
|
|
|
| 138 |
return httpx.AsyncClient(timeout=self.timeout_s, transport=self._transport)
|
| 139 |
|
| 140 |
def _payload(self, request: ChatRequest, *, stream: bool) -> dict[str, object]:
|
| 141 |
+
payload = {**self.options, **super()._payload(request, stream=stream)}
|
| 142 |
if self.seed is not None:
|
| 143 |
payload["seed"] = self.seed
|
| 144 |
return payload
|
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""bankml 0.3.6 on top of PR #1: sampling options, the chat/judge models a bankml node names,
|
| 2 |
+
and the serve flags that layer a persona and its parameters."""
|
| 3 |
+
|
| 4 |
+
from __future__ import annotations
|
| 5 |
+
|
| 6 |
+
import asyncio
|
| 7 |
+
import json
|
| 8 |
+
|
| 9 |
+
import httpx
|
| 10 |
+
import pytest
|
| 11 |
+
|
| 12 |
+
from mindxtrain.models.registry import ChatMessage, ChatRequest
|
| 13 |
+
from mindxtrain.operator.backends.bankml import BankmlBackend
|
| 14 |
+
|
| 15 |
+
_ENV = ("MINDXTRAIN_CHAT_MODEL", "MINDXTRAIN_JUDGE_MODEL", "MINDXTRAIN_CHAT_OPTIONS",
|
| 16 |
+
"MINDXTRAIN_BANKML_OPTIONS")
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
@pytest.fixture(autouse=True)
|
| 20 |
+
def _clean_env(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 21 |
+
for k in _ENV:
|
| 22 |
+
monkeypatch.delenv(k, raising=False)
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def _backend(seen: list[dict[str, object]], **kw: object) -> BankmlBackend:
|
| 26 |
+
def handler(request: httpx.Request) -> httpx.Response:
|
| 27 |
+
seen.append(json.loads(request.content))
|
| 28 |
+
return httpx.Response(200, json={
|
| 29 |
+
"model": "mindx-gen39", "choices": [{"message": {"content": "I am mindX."}, "finish_reason": "stop"}],
|
| 30 |
+
"usage": {"prompt_tokens": 9, "completion_tokens": 4},
|
| 31 |
+
"bankml_receipt": {"model_sha256": "6b64c748"},
|
| 32 |
+
})
|
| 33 |
+
|
| 34 |
+
return BankmlBackend(transport=httpx.MockTransport(handler), **kw) # type: ignore[arg-type]
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def _ask(b: BankmlBackend) -> None:
|
| 38 |
+
asyncio.run(b.chat(ChatRequest(model="mindx-gen39", messages=[ChatMessage(role="user", content="hi")],
|
| 39 |
+
max_tokens=16)))
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def test_backend_sends_options_only_when_set(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 43 |
+
seen: list[dict[str, object]] = []
|
| 44 |
+
_ask(_backend(seen))
|
| 45 |
+
assert not {"repeat_penalty", "repeat_last_n"} & set(seen[0])
|
| 46 |
+
monkeypatch.setenv("MINDXTRAIN_BANKML_OPTIONS", '{"repeat_penalty": 1.3, "repeat_last_n": 64}')
|
| 47 |
+
_ask(_backend(seen, options={"repeat_last_n": 32}))
|
| 48 |
+
assert seen[1]["repeat_penalty"] == 1.3 and seen[1]["repeat_last_n"] == 32 # ctor over env
|
| 49 |
+
assert seen[1]["model"] == "mindx-gen39" and seen[1]["max_tokens"] == 16 # request wins
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
def test_backend_options_must_be_an_object(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 53 |
+
monkeypatch.setenv("MINDXTRAIN_BANKML_OPTIONS", "[1.3]")
|
| 54 |
+
with pytest.raises(ValueError, match="JSON object"):
|
| 55 |
+
BankmlBackend()
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
def test_chat_and_judge_models_follow_env(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 59 |
+
from mindxtrain.eval.llama_evals import (
|
| 60 |
+
CorrectnessEvaluator,
|
| 61 |
+
PairwiseEvaluator,
|
| 62 |
+
default_judge_model,
|
| 63 |
+
)
|
| 64 |
+
from mindxtrain.governance.panel import default_chat_model
|
| 65 |
+
|
| 66 |
+
assert default_chat_model() == "llama3.2" and default_judge_model() == "llama3.2"
|
| 67 |
+
monkeypatch.setenv("MINDXTRAIN_CHAT_MODEL", "bonsai-8b-q1_0")
|
| 68 |
+
assert default_judge_model() == "bonsai-8b-q1_0" # judges follow the chat model
|
| 69 |
+
monkeypatch.setenv("MINDXTRAIN_JUDGE_MODEL", "bonsai-1.7b")
|
| 70 |
+
assert CorrectnessEvaluator().model == "bonsai-1.7b"
|
| 71 |
+
assert PairwiseEvaluator(model="explicit").model == "explicit"
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def test_chat_once_options(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 75 |
+
from mindxtrain.governance import panel
|
| 76 |
+
|
| 77 |
+
bodies: list[dict[str, object]] = []
|
| 78 |
+
|
| 79 |
+
def handler(request: httpx.Request) -> httpx.Response:
|
| 80 |
+
bodies.append(json.loads(request.content))
|
| 81 |
+
return httpx.Response(200, json={"choices": [{"message": {"content": "VERDICT: APPROVE"}}]})
|
| 82 |
+
|
| 83 |
+
real = httpx.Client
|
| 84 |
+
monkeypatch.setattr(httpx, "Client", lambda **k: real(transport=httpx.MockTransport(handler), **k))
|
| 85 |
+
url = "http://127.0.0.1:18093/v1"
|
| 86 |
+
panel.chat_once("m", [{"role": "user", "content": "x"}], base_url=url)
|
| 87 |
+
monkeypatch.setenv("MINDXTRAIN_CHAT_OPTIONS", '{"repeat_penalty": 1.3}')
|
| 88 |
+
panel.chat_once("m", [{"role": "user", "content": "x"}], base_url=url)
|
| 89 |
+
panel.chat_once("m", [{"role": "user", "content": "x"}], base_url=url, options={})
|
| 90 |
+
assert "repeat_penalty" not in bodies[0] and bodies[1]["repeat_penalty"] == 1.3
|
| 91 |
+
assert "repeat_penalty" not in bodies[2] # explicit {} wins
|
| 92 |
+
monkeypatch.setenv("MINDXTRAIN_CHAT_OPTIONS", "nope")
|
| 93 |
+
with pytest.raises(ValueError, match="not JSON"):
|
| 94 |
+
panel.chat_options()
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def test_deliberate_uses_the_env_chat_model(monkeypatch: pytest.MonkeyPatch) -> None:
|
| 98 |
+
from mindxtrain.governance import panel
|
| 99 |
+
from mindxtrain.governance.boardroom import Member
|
| 100 |
+
|
| 101 |
+
asked: list[str] = []
|
| 102 |
+
monkeypatch.setattr(panel, "chat_once", lambda model, *a, **k: asked.append(model) or "VERDICT: APPROVE")
|
| 103 |
+
monkeypatch.setenv("MINDXTRAIN_CHAT_MODEL", "bonsai-8b-q1_0")
|
| 104 |
+
d = panel.deliberate(Member(id="m1", role="critic"), "ship it")
|
| 105 |
+
assert asked == ["bonsai-8b-q1_0"] and d.vote == "approve"
|
| 106 |
+
|
| 107 |
+
|
| 108 |
+
def test_serve_param_parsing() -> None:
|
| 109 |
+
from mindxtrain.cli.main import _parse_params
|
| 110 |
+
|
| 111 |
+
assert _parse_params(["repeat_penalty=1.3", "num_ctx=2048", "seed=0"]) == {
|
| 112 |
+
"repeat_penalty": 1.3, "num_ctx": 2048, "seed": 0}
|
| 113 |
+
with pytest.raises(ValueError):
|
| 114 |
+
_parse_params(["repeat_penalty"])
|
|
@@ -30,8 +30,10 @@ USAGE_034 = "usage: bankml usage [PID …]\n bankml serve FILE --fork FORK
|
|
| 30 |
class FakeBankml:
|
| 31 |
"""Records argv; answers `version`, `--help`, `convert`, `create`, `sha256` like bankml."""
|
| 32 |
|
| 33 |
-
def __init__(self, usage: str = USAGE_035, create_rc: int = 0, create_err: str = ""
|
|
|
|
| 34 |
self.calls: list[list[str]] = []
|
|
|
|
| 35 |
self.usage = usage
|
| 36 |
self.create_rc = create_rc
|
| 37 |
self.create_err = create_err
|
|
@@ -41,7 +43,7 @@ class FakeBankml:
|
|
| 41 |
self.calls.append(list(cmd))
|
| 42 |
verb = cmd[1]
|
| 43 |
if verb == "version":
|
| 44 |
-
return subprocess.CompletedProcess(cmd, 0, "bankml
|
| 45 |
if verb == "--help":
|
| 46 |
return subprocess.CompletedProcess(cmd, 1, "", self.usage)
|
| 47 |
if verb == "convert":
|
|
@@ -172,7 +174,9 @@ def test_push_from_merged_dir(fake: FakeBankml, tmp_path: Path) -> None:
|
|
| 172 |
assert res.model_sha256 == SHA and res.digest == DIGEST
|
| 173 |
create = next(c for c in fake.calls if c[1] == "create")
|
| 174 |
assert create[2] == "mindx-gen99" and "--registry" in create
|
| 175 |
-
|
|
|
|
|
|
|
| 176 |
assert 'SYSTEM """You are mindX."""' in fake.modelfile_text
|
| 177 |
assert "repeat_penalty" not in fake.modelfile_text
|
| 178 |
|
|
@@ -188,12 +192,44 @@ def test_push_with_convert_first(fake: FakeBankml, tmp_path: Path) -> None:
|
|
| 188 |
|
| 189 |
def test_refused_params_never_reach_bankml(fake: FakeBankml, tmp_path: Path) -> None:
|
| 190 |
res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path),
|
| 191 |
-
params={"
|
| 192 |
assert res.status == "refused"
|
| 193 |
-
assert any("
|
| 194 |
assert fake.calls == [] # refused before even probing the binary
|
| 195 |
|
| 196 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 197 |
def test_non_llama_merged_dir_refused(fake: FakeBankml, tmp_path: Path) -> None:
|
| 198 |
res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path, "Qwen3ForCausalLM"))
|
| 199 |
assert res.status == "refused" and "Llama-architecture" in res.reason
|
|
|
|
| 30 |
class FakeBankml:
|
| 31 |
"""Records argv; answers `version`, `--help`, `convert`, `create`, `sha256` like bankml."""
|
| 32 |
|
| 33 |
+
def __init__(self, usage: str = USAGE_035, create_rc: int = 0, create_err: str = "",
|
| 34 |
+
version: str = "0.3.5") -> None:
|
| 35 |
self.calls: list[list[str]] = []
|
| 36 |
+
self.version = version
|
| 37 |
self.usage = usage
|
| 38 |
self.create_rc = create_rc
|
| 39 |
self.create_err = create_err
|
|
|
|
| 43 |
self.calls.append(list(cmd))
|
| 44 |
verb = cmd[1]
|
| 45 |
if verb == "version":
|
| 46 |
+
return subprocess.CompletedProcess(cmd, 0, f"bankml {self.version}\n", "")
|
| 47 |
if verb == "--help":
|
| 48 |
return subprocess.CompletedProcess(cmd, 1, "", self.usage)
|
| 49 |
if verb == "convert":
|
|
|
|
| 174 |
assert res.model_sha256 == SHA and res.digest == DIGEST
|
| 175 |
create = next(c for c in fake.calls if c[1] == "create")
|
| 176 |
assert create[2] == "mindx-gen99" and "--registry" in create
|
| 177 |
+
# converted through a link named after the tag: the GGUF's general.name is "Mindx Gen99"
|
| 178 |
+
link = tmp_path / "work" / "mindx-gen99" / "mindx-gen99"
|
| 179 |
+
assert fake.modelfile_text.startswith(f"FROM {link}\n") and link.resolve() == merged.resolve()
|
| 180 |
assert 'SYSTEM """You are mindX."""' in fake.modelfile_text
|
| 181 |
assert "repeat_penalty" not in fake.modelfile_text
|
| 182 |
|
|
|
|
| 192 |
|
| 193 |
def test_refused_params_never_reach_bankml(fake: FakeBankml, tmp_path: Path) -> None:
|
| 194 |
res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path),
|
| 195 |
+
params={"mirostat": 2, "temperature": 0.7})
|
| 196 |
assert res.status == "refused"
|
| 197 |
+
assert any("mirostat" in r for r in res.refusals)
|
| 198 |
assert fake.calls == [] # refused before even probing the binary
|
| 199 |
|
| 200 |
|
| 201 |
+
def test_penalties_refused_before_0_3_6(fake: FakeBankml, tmp_path: Path) -> None:
|
| 202 |
+
res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path),
|
| 203 |
+
params={"repeat_penalty": 1.3, "temperature": 0.7})
|
| 204 |
+
assert res.status == "refused" and res.bankml_version == "0.3.5"
|
| 205 |
+
assert any("repeat_penalty" in r and "0.3.6" in r for r in res.refusals)
|
| 206 |
+
assert not any(c[1] in ("create", "convert") for c in fake.calls)
|
| 207 |
+
|
| 208 |
+
|
| 209 |
+
def test_penalties_taken_from_0_3_6(fake: FakeBankml, tmp_path: Path) -> None:
|
| 210 |
+
fake.version = "0.3.6"
|
| 211 |
+
res = BP.push_to_bankml("b", "mindx-gen99", merged_dir=_merged(tmp_path), work_dir=tmp_path / "w",
|
| 212 |
+
params={"repeat_penalty": 1.3, "num_ctx": 2048}, stop=["<|im_end|>"])
|
| 213 |
+
assert res.ok, res
|
| 214 |
+
assert "PARAMETER repeat_penalty 1.3" in fake.modelfile_text
|
| 215 |
+
|
| 216 |
+
|
| 217 |
+
def test_version_gate() -> None:
|
| 218 |
+
assert BP.penalties_supported("0.3.6") and BP.penalties_supported("0.4.0")
|
| 219 |
+
assert not BP.penalties_supported("0.3.5") and not BP.penalties_supported("")
|
| 220 |
+
assert BP.bankml_sanitize(ModelfileSpec(from_model="x", parameters={"repeat_penalty": 1.3})).refusals
|
| 221 |
+
assert BP.bankml_sanitize(ModelfileSpec(from_model="x", parameters={"repeat_penalty": 1.3}),
|
| 222 |
+
penalties=True).ok
|
| 223 |
+
|
| 224 |
+
|
| 225 |
+
def test_merged_dir_already_named_after_the_tag_is_used_as_is(fake: FakeBankml, tmp_path: Path) -> None:
|
| 226 |
+
merged = tmp_path / "mindx-gen99"
|
| 227 |
+
merged.mkdir()
|
| 228 |
+
(merged / "config.json").write_text(json.dumps({"architectures": ["LlamaForCausalLM"]}))
|
| 229 |
+
res = BP.push_to_bankml("b", "mindx-gen99", merged_dir=merged, work_dir=tmp_path / "w")
|
| 230 |
+
assert res.ok and fake.modelfile_text.startswith(f"FROM {merged.resolve()}\n")
|
| 231 |
+
|
| 232 |
+
|
| 233 |
def test_non_llama_merged_dir_refused(fake: FakeBankml, tmp_path: Path) -> None:
|
| 234 |
res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path, "Qwen3ForCausalLM"))
|
| 235 |
assert res.status == "refused" and "Llama-architecture" in res.reason
|