|
Download docs/modules/native.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 19.1 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/native.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/native.md
-
curl -L -o native.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/native.md
19.1 kB
| # `bankML/native.rs` — the engine behind `bankml serve --native` | |
| ## Summary | |
| `native.rs` answers chat completions from bankML's own forward pass. It was added in 0.3.0. It is built to be | |
| token-identical to llama-server b11192 on its oracle: the same template, tokenizer, forward pass (every kernel | |
| llama.cpp chooses), sampler chain and, across requests, the same prompt cache. It holds one conversation slot, as | |
| llama-server's `-np 1`. | |
| 0.3.8 (unreleased) adds what llama-server does at the slot's edges: the context limit, saving and restoring the slot | |
| to a file, llama-server's host prompt cache for conversations that take turns, and each token's logprobs, reported | |
| step by step for streaming. 0.3.9 (unreleased, in progress) adds an optional q8_0 KV cache | |
| (`BANKML_CACHE_TYPE=q8_0`). | |
| Since 0.3.1 it also owns the model's lifecycle. `Registry` names every GGUF pinned in the forks directory. | |
| `Residency` keeps one model resident at a time, runs the full `verify` (the guard, then the sha256 pin) on every | |
| load, and drops an idle model, with its memory map, when its keep-alive runs out. | |
| `serve.rs` routes `/v1/chat/completions` and llama-server's endpoints here; `ollama.rs` routes `/api/*` here; the | |
| C API's `bankml_chat` uses the same engine. All of them share one `Residency`, so they share one engine and one slot. | |
| ## Technical usage | |
| ### `Native`: one loaded model, one slot | |
| | item | what it does | | |
| |---|---| | |
| | `Native::open(model: &Path, n_ctx: usize) -> Result<Native, String>` | reads the template, tokenizer, weights, the GGUF's sampling defaults and the end-of-generation set; allocates the KV cache in the type `BANKML_CACHE_TYPE` names (below) | | |
| | `prompt(&self, messages: &Json) -> Result<Vec<u32>, String>` | the template, then tokenization with special tokens | | |
| | `prompt_fit(&self, messages, num_ctx: Option<usize>)` | the same, fitted under a `num_ctx` (a request's option or a derived model's `PARAMETER num_ctx`) as Ollama fits it (below; 0.3.5) | | |
| | `params(&self, req: &Json) -> Result<Params, String>` | `sampling(self.defaults.clone(), req)` | | |
| | `grammar(&self, c: &Constraint)` | JSON mode's or a schema's grammar with the template's generation prompt prefilled, or a user GBNF as given | | |
| | `grammar_vocab(&self)` | every token's piece and the end set, as the grammar and DRY read them; built once, on first use | | |
| | `complete(&self, prompt, params, max_tokens: Option<usize>, grammar, emit: impl FnMut(&str) -> bool) -> Result<Done, String>` | one completion; `emit` gets whole UTF-8 pieces and returns false to stop; under a grammar each token is drawn as llama.cpp's `common_sampler_sample` draws it | | |
| | `complete_with(&self, prompt, params, max_tokens, grammar, emit: impl FnMut(Step) -> bool) -> Result<Done, String>` | the same, reporting every step (0.3.8); `complete` is built on it | | |
| | `save_slot(&self, path, model_sha256) -> Result<(usize, u64), String>` | writes the slot to a file; returns (tokens, bytes) (0.3.8) | | |
| | `restore_slot(&self, path, model_sha256) -> Result<(usize, u64), String>` | reads one back; on any error empties the slot and refuses with `Unable to restore slot: …` (0.3.8) | | |
| | `erase_slot(&self) -> usize` | empties the slot; returns the tokens it held (0.3.8) | | |
| | `reset(&self)` | empties the slot, as a fresh llama-server or `cache_prompt: false` | | |
| | `props_json(&self)` | `/props` in the shape Savante reads | | |
| `Done` carries `prompt_tokens`, `cached_tokens`, `completion_tokens`, `finish_reason` (`stop` or `length`), `text`, | |
| `prompt_ns`, `eval_ns`, the generated `tokens` (the end token included when one ended the answer), under a grammar | |
| `grammar_ns` and `resampled`, since 0.3.7 `ttft_ns` and `energy_j` (the CPU package's energy when RAPL is readable, | |
| else `None`), and since 0.3.8 `probs`: one `TokenLogprob` per generated token when `Params::n_probs` > 0. | |
| `Step` (0.3.8) is one step of `complete_with`: | |
| | field | meaning | | |
| |---|---| | |
| | `piece: &str` | the text this step released; may be empty | | |
| | `complete: bool` | the token left no incomplete UTF-8 behind (llama-server sends a token only then) | | |
| | `entry: Option<&TokenLogprob>` | its logprobs, when asked for and complete | | |
| | `n: usize` | tokens generated so far | | |
| | `eog: bool` | the end-of-turn token, which ends the answer and releases no text | | |
| A last step with `n` unchanged and `complete` false flushes bytes left incomplete at the end. `TokenLogprob` holds | |
| the token's `id`, its probability `p`, its `piece`, the `top` tokens as (id, p, bytes), and whether it was `emitted` | |
| and `complete`. The probabilities come from `sampler::token_probs` on the raw logits. `serve.rs` turns the steps into | |
| llama-server's stream (`NativeChat::run_steps`, see [serve.md](serve.md)). | |
| ```rust | |
| let d = eng.complete_with(&prompt, params, Some(64), None, |s| { | |
| if s.complete && !s.piece.is_empty() { print!("{}", s.piece); } | |
| true // false stops the answer | |
| })?; | |
| ``` | |
| ### The prompt cache rule | |
| `complete` reproduces llama-server's slot reuse (`server-context.cpp`): | |
| - first, llama-server's host prompt cache ([prompt_cache.md](prompt_cache.md), 0.3.8): when the slot serves the | |
| prompt poorly, its state goes to RAM and a cached state that serves it better comes back; | |
| - the longest common prefix of the tokens already in the slot and the new prompt is kept; | |
| - when the whole prompt is cached, one token less is kept (llama-server evaluates at least one); | |
| - the KV cache is truncated there, and the rest is computed in micro-batches of 512 (`Weights::prefill`). | |
| Starting prefill where llama-server starts it means each row takes the same kernels as llama-server's rows, which is | |
| why the reuse rule is part of token-identity. | |
| Before the first draw, every prompt token is accepted into the sampler, cached or not. That is how llama-server | |
| fills the penalties' window (O2) and DRY's. The end-of-turn token ends the answer, is not part of the text, and is | |
| counted among the completion tokens, as llama-server counts it. | |
| The context limit (0.3.8), as llama-server with context shift off: a prompt as long as the context or longer is | |
| refused (`the prompt has N tokens; the context is M`; `serve.rs` refuses it earlier, with llama-server's 400 body), | |
| and a generation that fills the context stops there with `finish_reason` `length`. | |
| ### `sampling()`: the request over the model's defaults | |
| `pub fn sampling(defaults: Params, req: &Json) -> Result<Params, String>` resolves the request as llama-server does. | |
| It reads these numeric fields: | |
| - `temperature`, `top_k`, `top_p`, `min_p`, `min_keep`, `seed`; | |
| - since 0.3.6, the penalties: `repeat_last_n`, `repeat_penalty`, `frequency_penalty`, `presence_penalty`; | |
| - since 0.3.7, the rest of llama-server's default chain: `typical_p`, `top_n_sigma`, `xtc_probability`, | |
| `xtc_threshold`, `dynatemp_range`, `dynatemp_exponent`, `dry_multiplier`, `dry_base`, `dry_allowed_length`, | |
| `dry_penalty_last_n`, and the array `dry_sequence_breakers`; | |
| - since 0.3.8, `n_probs` (`serve.rs` sets it from OpenAI's `logprobs` and `top_logprobs`). | |
| llama-server's soft limits clamp: `top_p`, `min_p` and the two XTC fields to [0, 1], `temperature` to ≥ 0, and a | |
| `dry_base` below 1 falls back to 1.75. Refused, because llama-server would act on them and bankML does not reproduce | |
| them: `mirostat` (non-zero) and a custom `samplers` order. `Sampler::new` (in `sampler.rs`) then refuses what | |
| llama-server refuses, with its message: a negative `repeat_last_n`, `dry_allowed_length` or `dry_penalty_last_n`, a | |
| repeat penalty of 0 or below, an empty `dry_sequence_breakers`; and top-k outside 1–128, which bankML does not | |
| reproduce. DRY's breakers are built from the vocabulary's pieces once per breaker list and kept on the engine. | |
| ### Slot files and the cache type (0.3.8, 0.3.9) | |
| `save_slot` writes bankML's own format: a magic line (`bankML slot v1`), the model's sha256, the context, token and | |
| layer counts, each layer's cache, the tokens, then a sha256 of all of it. It is written to a temporary name and | |
| renamed, so a crash leaves no half file. `restore_slot` refuses a file for another model, another context or layer | |
| shape, a damaged file (the sha256 does not match), trailing data, or tokens outside the vocabulary. The slot is then | |
| emptied, never half-filled, and the next answer computes its whole prompt. `serve.rs` exposes these as | |
| `POST /slots/0?action=save|restore|erase` (see [serve.md](serve.md)). | |
| `BANKML_CACHE_TYPE` picks the KV cache's type when the model opens: `f16` (the default, llama.cpp's) or `q8_0`, as | |
| llama.cpp's `--cache-type-k q8_0 --cache-type-v q8_0`. A q8_0 cache takes 53 % of the f16 cache's bytes (CHANGELOG | |
| 0.3.9) and needs a head size that is a multiple of 32; anything else is refused. The attention over it, with | |
| llama.cpp's Hadamard rotation, is in `forward.rs` ([forward.md](forward.md)). A slot file records its cache type, and | |
| one saved with the other type is refused with the reason. The host prompt cache counts each saved state's real | |
| bytes, so a q8_0 cache keeps more states within `BANKML_CACHE_RAM`. | |
| ### `prompt_fit` and `fit_messages`: `num_ctx` as Ollama applies it | |
| `pub fn fit_messages(msgs, num_ctx, count) -> Result<(Vec<Message>, usize), String>` is Ollama v0.13.3's `chatPrompt` | |
| truncation (`server/prompt.go`), step for step: | |
| - walking back from the last message, each earlier one is kept while the conversation from it, with the system | |
| messages before it, still fits; | |
| - the last message always stays, and so do the system messages before the first one kept; | |
| - Ollama's quirk is kept: a system message that is itself the first one cut is dropped. | |
| The tokens counted are the prompt bankML answers (the GGUF's template, as llama-server renders it). If the fitted | |
| prompt still exceeds `num_ctx`, `prompt_fit` refuses: Ollama's runner would cut tokens out of the middle (`num_keep`) | |
| and shift the cache during the answer, which bankML does not reproduce. | |
| ### `Registry`: names for pinned files | |
| - `Registry::build(model, fork_json, dir)`: the startup model first (the default), then, with `dir`, every `.gguf` | |
| named in every `*.FORK.json` there, looked for beside the startup model and in `dir`. A name pinned twice keeps its | |
| first entry. | |
| - `model_name(file)`: the file's stem, lower-cased (`bonsai-1.7b-q1_0`). | |
| - `resolve(model)`: empty means the startup model; otherwise the name, with or without `:latest`, or the file name, | |
| case-insensitively. Since 0.3.4, the name without its weight-type suffix (`base_name`: `-f16`, `-q1_0`, | |
| `-q2_0_g64`, …) is accepted when exactly one pin has that base (`mindx-gen39`). | |
| - `header_info(path)`: whether the forward pass plays a file, from its header alone (guard, architecture, weight | |
| type and graph options via `forward::plan`, tokenizer, chat template), and if not, why: the reason names what is | |
| missing and where it is on the roadmap. `/api/tags` reports it. | |
| ### `Residency`: one resident model | |
| | method | behaviour | | |
| |---|---| | |
| | `start(...)` | the startup model, already verified by `serve::run`, resident until told otherwise; refused if the forward pass does not play it | | |
| | `acquire(&run, e)` | the resident model if it is `e` and its file identity is unchanged; otherwise the resident one is dropped first, then `e` is verified and opened. Errors carry their HTTP status (404 not on this machine, 400 not playable or refused by `verify`, 503 file changed) | | |
| | `touch(name, KeepAlive)` | after a request: `Unload` drops it now, `For(d)` sets the expiry, `Forever` clears it | | |
| | `reap()` | drops the resident model once it has expired, unless a request holds the engine (`serve.rs` calls it every second) | | |
| | `current_or_default()` | the resident model, or the startup model loaded (for endpoints that name no model) | | |
| | `shown_as(tag, num_ctx, loaded)` / `shown()` | what `/api/ps` names: the model whose request loaded the weights, and its context (0.3.5). Ollama's runner keeps the name of the model whose request loaded it and reloads, under the new name, when a request's `num_ctx` differs; bankML keeps the weights and records the same: when they were (re)loaded for the request, when nothing is recorded yet, or when the context differs | | |
| `run: Mutex<()>` is held for a whole completion, load or unload, so the engine has one user and a switch never has | |
| two models in memory. `cur` is held only briefly. `derived` holds the derived models of the registry directory | |
| (O5, [create.md](create.md)), layered on the pins; `last` is the most recent verification, which `/bankml` reports | |
| when nothing is resident. | |
| ## How it is verified | |
| - `fit_messages_as_ollama` (unit test): Ollama's truncation on a synthetic count, including the system-message quirk. | |
| - `oracle_native_serve` (`#[ignore]`, in the gate): three Savante-style conversations recorded from llama-server | |
| b11192 by `testing/serve_oracle.py`, replayed in one engine; text, prompt and completion counts and cache reuse | |
| must match on every turn. docs/oracles.md records 9 of 9 turns; turn 2 reuses 49 of turn 1's 53 prompt tokens. | |
| - `oracle_native_serve_o4` (0.3.4): the same on the models O4 opens, Bonsai-1.7B (tied embeddings) and the Llama | |
| graph in F16 (SmolLM2-135M-Instruct and mindx-gen39), each against llama-server running that model, then each | |
| one's JSON mode with its template's own grammar. | |
| - `oracle_json_mode`, `oracle_json_mode_ternary`, `oracle_json_schema`, `oracle_json_schema_ternary`, | |
| `oracle_json_schema_o4`: constrained answers against `testing/json_oracle.py` and `testing/json_schema_oracle.py` | |
| records, each from an empty cache: the same tokens (end token included), raw text, message content, finish reason | |
| and counts, with the server's reported grammar and generation prompt. The schema records (O6b) cover objects with | |
| required and optional fields, enums, ranges, nested `$defs`, a pattern, and a top-level array and string, through | |
| `response_format: json_schema`, the top-level `json_schema` and `json_object` with a schema; each record carries | |
| the request as sent, read exactly (a float literal stays one). `oracle_json_schema_o4` (0.3.5) runs them on | |
| Bonsai-1.7B (the Qwen3 template) and the two ChatML templates, whose schema grammar has no reasoning block | |
| (`schema::chat_grammar`). | |
| - `oracle_penalties` and `oracle_penalties_8b` (0.3.6, O2): `testing/penalty_oracle.py` records of repeat, frequency | |
| and presence penalties over `repeat_last_n`, greedy and seeded, with the prompt in the window. Each request must | |
| resolve to the parameters llama-server read back, give its tokens, and be refused with its message where it | |
| refused. CHANGELOG 0.3.6: mindx-gen39 56 / 56, Bonsai-1.7B 56 / 56, Bonsai-8B 56 / 56; 12 / 12 refusals each. | |
| - `oracle_samplers` (0.3.7): `testing/penalty_oracle.py --kind sampler` records, 23 variants × 4 prompts (typical-p, | |
| top-n-σ, XTC, dynamic temperature, DRY, alone and together). CHANGELOG 0.3.7: mindx-gen39 76 / 76, Bonsai-1.7B | |
| 76 / 76; 16 / 16 refusals each. `oracle_samplers_8b` replays a Bonsai-8B record in the gate; | |
| CHANGELOG 0.3.7 does not yet give its count. | |
| - Live, in the gate, each driving a running `serve --native`: `serve_oracle_ollama_shape`, `penalty_oracle_live`, | |
| `sampler_oracle_live`, and from 0.3.8 `context_oracle_live` (8 / 8), `slot_oracle_live` (19 / 19: the answer after | |
| a restore, in the same server and after a restart, equals an empty slot's), `session_oracle_live` (14 / 14, the host | |
| prompt cache), `logprobs_oracle_live` (14 / 14, five streamed); from 0.3.9 `kv_oracle_live` (6 / 6 with | |
| `BANKML_CACHE_TYPE=q8_0` against llama-server `--cache-type-k/v q8_0`). Counts from CHANGELOG 0.3.8 and 0.3.9. | |
| ## Advantages and efficiency | |
| - **A verified engine, not a trusted one.** Every load runs the guard and the sha256 pin before a byte of weights | |
| is used. llama-server and Ollama load whatever file they are given. The pin uses the CPU's SHA extensions when | |
| present: docs/PERFORMANCE.md measures 0.23 s for a 248 MB model (5.5× the portable hash) and 2.9 s for the | |
| 1.16 GB 8B model. | |
| - **llama-server's prompt cache, reproduced.** A follow-up turn computes only the tokens after the shared prefix, | |
| as the server does, so a conversation does not pay for its history each turn. | |
| - **One resident model, freed on time.** As Ollama's `MAX_LOADED_MODELS=1`: the old model goes before the new one is | |
| read, and the reaper releases the memory map when the keep-alive expires. The file identity (device, inode, size, | |
| modification time) is checked before each reuse instead of rehashing. | |
| - **Streaming by whole characters.** Pieces are emitted as soon as they form valid UTF-8; the grammar vocabulary | |
| (about 10 MB) and the JSON-mode rules are built once, on first use (`OnceLock`). | |
| - **Rust practice visible here.** No external crates (Cargo.toml has an empty `[dependencies]`); errors are | |
| `Result<_, String>` that carry a reason, or `(u16, String)` with the HTTP status; poisoned mutexes are recovered | |
| with `unwrap_or_else(|e| e.into_inner())`; the toolchain is pinned to 1.99.0 in `rust-toolchain.toml`. | |
| - **A warm start that survives a restart.** A slot restore replaces the prompt's prefill with one read checked by | |
| sha256; Savante and the console pass `--slot-dir` for this (CHANGELOG 0.3.8). | |
| - **Half the KV memory, on request.** `BANKML_CACHE_TYPE=q8_0` stores the cache in 53 % of the f16 bytes, with | |
| answers token-identical to llama-server configured the same way (CHANGELOG 0.3.9). | |
| - **Logprobs cost nothing when off.** `token_probs` runs only when `n_probs` > 0: one partial sort of the vocabulary | |
| per generated token. | |
| - **Next** (docs/TODO.md 0.4.0, docs/OLLAMA.md): more than one slot with continuous batching (O8); a 4-bit KV cache | |
| (`q4_0`); Ollama's ContextShift past `num_ctx` (O2). | |
| ## Limitations | |
| - One slot, as llama-server `-np 1`. Requests are served one at a time; other conversations' states wait in the | |
| host prompt cache (`testing/session_oracle.py`). | |
| - One resident model. A request for another model unloads the current one. | |
| - Not reproduced, so refused with a reason: mirostat, a custom sampler order, top-k 0 or above 128. | |
| - A prompt that does not fit `num_ctx` after Ollama's message truncation is refused; Ollama would cut tokens out of | |
| its middle. Tokens past `num_ctx` during an answer are not shifted out (ContextShift); the answer runs within the | |
| served `--ctx` (docs/OLLAMA.md, "Still open"). | |
| - The KV cache is f16 or q8_0 only; another `BANKML_CACHE_TYPE` is refused when the model opens. The type is read | |
| from the environment, not per request, so a slot file saved under one type does not restore under the other. | |
| - Slot files are bankML's own format, not llama-server's; a file from one does not load in the other. | |
| ## See also | |
| - [../usage.md §6a, Ollama's API](../usage.md#6a-ollamas-api) · [../OLLAMA.md](../OLLAMA.md) · | |
| [../oracles.md](../oracles.md) · [../TODO.md](../TODO.md) · [../CAPI.md](../CAPI.md) | |
| - [serve.md](serve.md) · [ollama.md](ollama.md) · [create.md](create.md) · [sampler.md](sampler.md) · | |
| [prompt_cache.md](prompt_cache.md) · [grammar.md](grammar.md) · [schema.md](schema.md) · [forward.md](forward.md) · | |
| [gguf.md](gguf.md) · [metrics.md](metrics.md) | |