File size: 7,561 Bytes
28c70af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 | # `bankML/chat.rs` — chat templates, byte-identical to llama.cpp b11192
## Summary
A conversation becomes a prompt through the model's chat template, a Jinja program stored in the GGUF
(`tokenizer.chat_template`). bankml does not run Jinja. `chat.rs` writes out the rules of each template it
reproduces, and identifies the model's template by the sha256 of its text. A template it does not know is refused.
Three templates are reproduced (`TEMPLATES`):
| sha256 prefix | `Template` | carried by |
|---|---|---|
| `30a75d10e60b57e2` | `Qwen3` | Bonsai 1.7B and 8B (Qwen3), thinking off |
| `872be49dbb638044` | `SmolLm2` | SmolLM2-Instruct: ChatML with a default system message |
| `9fe579a2c222698c` | `ChatMl` | plain ChatML, no default system message (`mindx-genN`) |
On the Qwen3 template the generation prompt ends in an empty `<think>` block (thinking off). On the two ChatML
templates every message renders as `<|im_start|>{role}\n{content}<|im_end|>\n` whatever its role, and
`reasoning_content` is dropped, as llama-server renders them.
Callers: the native engine (`native.rs`: `template_of` at open, `render` for every prompt, `generation_prompt` for
the grammar's prefill); `bankml serve --native` (`POST /apply-template`, `/v1/chat/completions`, `/api/chat`); the
Ollama layer (`ollama.rs`); `bankml chat-template` and `bankml generate` (`main.rs`); and `grammar.rs` / `schema.rs`,
which choose the JSON grammar's root by template.
## Technical usage
```rust
pub const TEMPLATE_SHA256: &str = "30a75d10e60b57e2";
pub enum Template { Qwen3, SmolLm2, ChatMl }
pub const TEMPLATES: [(&str, Template, &str); 3];
impl Template {
pub fn render(self, msgs: &[Message]) -> Result<String, String>
pub fn generation_prompt(self) -> &'static str
}
pub struct Message { pub role: String, pub content: String, pub reasoning: Option<String> }
impl Message { pub fn new(role: &str, content: &str) -> Self }
pub fn template_of(gguf: &std::path::Path) -> Result<Template, String>
pub fn check_template(gguf: &std::path::Path) -> Result<(), String>
pub fn messages_from_json(v: &Json) -> Result<Vec<Message>, String>
pub fn render(msgs: &[Message]) -> Result<String, String> // the Qwen3 template
```
- `template_of` guards the file, hashes `tokenizer.chat_template` and matches the first 16 hex digits against
`TEMPLATES`. The error names the hash it found and the templates it knows.
- `messages_from_json` reads OpenAI-style `[{"role", "content", "reasoning_content"?}, …]`. An empty
`reasoning_content` is dropped before templating, as llama-server does.
- `render` (Qwen3) reproduces the template's details: a first system message; the last real user query (scanning
back, the first user message that is not a wrapped `<tool_response>`); assistant turns after it keep their
`<think>` block, those before it lose it; runs of tool messages grouped into one user turn of
`<tool_response>` blocks; Python's `split` and `strip` semantics where the template uses them.
- `generation_prompt()` is what the template appends for the assistant's turn:
`"<|im_start|>assistant\n<think>\n\n</think>\n\n"` on Qwen3, `"<|im_start|>assistant\n"` on the ChatML templates.
```sh
echo '[{"role":"system","content":"You are Savante."},{"role":"user","content":"Hi"}]' \
| bankml chat-template .models/Bonsai-8B-Q1_0.gguf
```
prints:
```text
<|im_start|>system
You are Savante.<|im_end|>
<|im_start|>user
Hi<|im_end|>
<|im_start|>assistant
<think>
</think>
```
## How it is verified
- Unit tests: `a_short_conversation`, `python_split_semantics`, `chatml_templates`.
- `oracle_chat_template` (`#[ignore]`, in the gate): `testing/template_oracle.py` records llama-server's own
`/apply-template` for 317 conversations (system prompts first, later or absent; assistant turns with and without
`<think>` blocks around the last real query; `reasoning_content`; runs of tool results; user messages that look
like tool responses; special markers and Unicode in content; 300 random conversations). Every prompt must be
byte-identical: **317 of 317** (0.3.6 gate record).
- `oracle_chat_template_chatml`: SmolLM2-Instruct's template and mindx-gen39's ChatML, each against llama-server
rendering that model's own, **317 of 317** each (docs/oracles.md §5d).
- Indirectly: the conversation oracle (`oracle_native_serve`, 9 of 9 turns, including the prompt cache's reuse) and
every greedy, sampling and JSON oracle start from a rendered prompt.
The oracle found one server behaviour the template alone would not predict: an empty `reasoning_content` is dropped.
## Advantages and efficiency
- **No template engine.** No Jinja interpreter and no crate: each template is a short, readable Rust function, and
the sha256 pin means a model whose template changed is refused instead of rendered wrongly.
- **Exact prompts.** A byte-identical prompt is what lets the prompt cache reuse what llama-server would reuse, and
lets every token-level oracle compare like with like.
- **Cheap.** Rendering is string concatenation over the messages; the template is identified once, when the engine
opens the model.
- **Rust practice.** Zero dependencies, no `unsafe`; out-of-scope input is an `Err` with the reason, not a guess.
- **Next** (docs/TODO.md, docs/OLLAMA.md): tool calls through the template (O6; JSON schemas, its first step, came in
0.3.5), and the Llama 3.x template with that architecture (0.6.0).
## Limitations
- Only the three templates above. Any other is refused by its hash.
- Refused rather than guessed: tool definitions and `tool_calls` (the template serialises them with its own
`tojson`); non-text content; a conversation ending with an assistant message (llama-server treats that as a
prefill, which is server logic beyond the template); an empty message list.
- On Qwen3, a role other than system, user, assistant and tool renders nothing, as in the template.
- Ollama's persona `SYSTEM` for `mindx-genN` lives in its Modelfile, not in the GGUF: here it is the caller's system
message (`bankml create` handles the Modelfile).
## Design notes
- History: the chat template was step two of P3 (the Bonsai / Qwen3 template only). 0.3.4 (O4) mapped templates
per model, each pinned by the sha256 of its text, adding SmolLM2-Instruct's and the plain ChatML of `mindx-genN`.
- The SmolLM2-Instruct default system message, inserted when the first message is not a system message, is
`You are a helpful AI assistant named SmolLM, trained by Hugging Face`.
- `mindx-genN` is SmolLM2-135M fine-tuned by mindXtrain; its GGUF carries plain ChatML.
- Each template's oracle is llama-server's `/apply-template` on a GGUF that carries that template
(`testing/template_oracle.py`); the oracle shows both ChatML templates render a tool result like any other role
and drop `reasoning_content`.
- Why an empty `reasoning_content` is dropped: llama-server removes it before templating, so the assistant
content's own `<think>` block is split out; a present-but-empty string would have prevented that split.
## See also
- [../oracles.md](../oracles.md) §1c, §5d — the chat-template oracle
- [../usage.md](../usage.md) §13 — `bankml chat-template`
- [../TECHNICAL.md](../TECHNICAL.md) §III.7 — from a template to a token
- [../OLLAMA.md](../OLLAMA.md) — O4 templates, O6 tool calls
- [../TODO.md](../TODO.md)
- Sibling pages: [tokenizer.md](tokenizer.md), [grammar.md](grammar.md), [schema.md](schema.md),
[native.md](native.md), [serve.md](serve.md), [ollama.md](ollama.md), [create.md](create.md)
|