File size: 7,561 Bytes
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
# `bankML/chat.rs` — chat templates, byte-identical to llama.cpp b11192

## Summary

A conversation becomes a prompt through the model's chat template, a Jinja program stored in the GGUF
(`tokenizer.chat_template`). bankml does not run Jinja. `chat.rs` writes out the rules of each template it
reproduces, and identifies the model's template by the sha256 of its text. A template it does not know is refused.

Three templates are reproduced (`TEMPLATES`):

| sha256 prefix | `Template` | carried by |
|---|---|---|
| `30a75d10e60b57e2` | `Qwen3` | Bonsai 1.7B and 8B (Qwen3), thinking off |
| `872be49dbb638044` | `SmolLm2` | SmolLM2-Instruct: ChatML with a default system message |
| `9fe579a2c222698c` | `ChatMl` | plain ChatML, no default system message (`mindx-genN`) |

On the Qwen3 template the generation prompt ends in an empty `<think>` block (thinking off). On the two ChatML
templates every message renders as `<|im_start|>{role}\n{content}<|im_end|>\n` whatever its role, and
`reasoning_content` is dropped, as llama-server renders them.

Callers: the native engine (`native.rs`: `template_of` at open, `render` for every prompt, `generation_prompt` for
the grammar's prefill); `bankml serve --native` (`POST /apply-template`, `/v1/chat/completions`, `/api/chat`); the
Ollama layer (`ollama.rs`); `bankml chat-template` and `bankml generate` (`main.rs`); and `grammar.rs` / `schema.rs`,
which choose the JSON grammar's root by template.

## Technical usage

```rust
pub const TEMPLATE_SHA256: &str = "30a75d10e60b57e2";
pub enum Template { Qwen3, SmolLm2, ChatMl }
pub const TEMPLATES: [(&str, Template, &str); 3];

impl Template {
    pub fn render(self, msgs: &[Message]) -> Result<String, String>
    pub fn generation_prompt(self) -> &'static str
}

pub struct Message { pub role: String, pub content: String, pub reasoning: Option<String> }
impl Message { pub fn new(role: &str, content: &str) -> Self }

pub fn template_of(gguf: &std::path::Path) -> Result<Template, String>
pub fn check_template(gguf: &std::path::Path) -> Result<(), String>
pub fn messages_from_json(v: &Json) -> Result<Vec<Message>, String>
pub fn render(msgs: &[Message]) -> Result<String, String>   // the Qwen3 template
```

- `template_of` guards the file, hashes `tokenizer.chat_template` and matches the first 16 hex digits against
  `TEMPLATES`. The error names the hash it found and the templates it knows.
- `messages_from_json` reads OpenAI-style `[{"role", "content", "reasoning_content"?}, …]`. An empty
  `reasoning_content` is dropped before templating, as llama-server does.
- `render` (Qwen3) reproduces the template's details: a first system message; the last real user query (scanning
  back, the first user message that is not a wrapped `<tool_response>`); assistant turns after it keep their
  `<think>` block, those before it lose it; runs of tool messages grouped into one user turn of
  `<tool_response>` blocks; Python's `split` and `strip` semantics where the template uses them.
- `generation_prompt()` is what the template appends for the assistant's turn:
  `"<|im_start|>assistant\n<think>\n\n</think>\n\n"` on Qwen3, `"<|im_start|>assistant\n"` on the ChatML templates.

```sh
echo '[{"role":"system","content":"You are Savante."},{"role":"user","content":"Hi"}]' \
  | bankml chat-template .models/Bonsai-8B-Q1_0.gguf
```

prints:

```text
<|im_start|>system
You are Savante.<|im_end|>
<|im_start|>user
Hi<|im_end|>
<|im_start|>assistant
<think>

</think>

```

## How it is verified

- Unit tests: `a_short_conversation`, `python_split_semantics`, `chatml_templates`.
- `oracle_chat_template` (`#[ignore]`, in the gate): `testing/template_oracle.py` records llama-server's own
  `/apply-template` for 317 conversations (system prompts first, later or absent; assistant turns with and without
  `<think>` blocks around the last real query; `reasoning_content`; runs of tool results; user messages that look
  like tool responses; special markers and Unicode in content; 300 random conversations). Every prompt must be
  byte-identical: **317 of 317** (0.3.6 gate record).
- `oracle_chat_template_chatml`: SmolLM2-Instruct's template and mindx-gen39's ChatML, each against llama-server
  rendering that model's own, **317 of 317** each (docs/oracles.md §5d).
- Indirectly: the conversation oracle (`oracle_native_serve`, 9 of 9 turns, including the prompt cache's reuse) and
  every greedy, sampling and JSON oracle start from a rendered prompt.

The oracle found one server behaviour the template alone would not predict: an empty `reasoning_content` is dropped.

## Advantages and efficiency

- **No template engine.** No Jinja interpreter and no crate: each template is a short, readable Rust function, and
  the sha256 pin means a model whose template changed is refused instead of rendered wrongly.
- **Exact prompts.** A byte-identical prompt is what lets the prompt cache reuse what llama-server would reuse, and
  lets every token-level oracle compare like with like.
- **Cheap.** Rendering is string concatenation over the messages; the template is identified once, when the engine
  opens the model.
- **Rust practice.** Zero dependencies, no `unsafe`; out-of-scope input is an `Err` with the reason, not a guess.
- **Next** (docs/TODO.md, docs/OLLAMA.md): tool calls through the template (O6; JSON schemas, its first step, came in
  0.3.5), and the Llama 3.x template with that architecture (0.6.0).

## Limitations

- Only the three templates above. Any other is refused by its hash.
- Refused rather than guessed: tool definitions and `tool_calls` (the template serialises them with its own
  `tojson`); non-text content; a conversation ending with an assistant message (llama-server treats that as a
  prefill, which is server logic beyond the template); an empty message list.
- On Qwen3, a role other than system, user, assistant and tool renders nothing, as in the template.
- Ollama's persona `SYSTEM` for `mindx-genN` lives in its Modelfile, not in the GGUF: here it is the caller's system
  message (`bankml create` handles the Modelfile).

## Design notes

- History: the chat template was step two of P3 (the Bonsai / Qwen3 template only). 0.3.4 (O4) mapped templates
  per model, each pinned by the sha256 of its text, adding SmolLM2-Instruct's and the plain ChatML of `mindx-genN`.
- The SmolLM2-Instruct default system message, inserted when the first message is not a system message, is
  `You are a helpful AI assistant named SmolLM, trained by Hugging Face`.
- `mindx-genN` is SmolLM2-135M fine-tuned by mindXtrain; its GGUF carries plain ChatML.
- Each template's oracle is llama-server's `/apply-template` on a GGUF that carries that template
  (`testing/template_oracle.py`); the oracle shows both ChatML templates render a tool result like any other role
  and drop `reasoning_content`.
- Why an empty `reasoning_content` is dropped: llama-server removes it before templating, so the assistant
  content's own `<think>` block is split out; a present-but-empty string would have prevented that split.

## See also

- [../oracles.md](../oracles.md) §1c, §5d — the chat-template oracle
- [../usage.md](../usage.md) §13 — `bankml chat-template`
- [../TECHNICAL.md](../TECHNICAL.md) §III.7 — from a template to a token
- [../OLLAMA.md](../OLLAMA.md) — O4 templates, O6 tool calls
- [../TODO.md](../TODO.md)
- Sibling pages: [tokenizer.md](tokenizer.md), [grammar.md](grammar.md), [schema.md](schema.md),
  [native.md](native.md), [serve.md](serve.md), [ollama.md](ollama.md), [create.md](create.md)