# `bankML/tokenizer.rs` — byte-level BPE, token-identical to llama.cpp b11192 This page also covers `bankML/unicode_letters.rs`, the tokenizer's letter table. ## Summary `tokenizer.rs` turns text into token ids exactly as llama.cpp b11192's `llama_tokenize(…, add_special = false, parse_special)` does, for GGUF vocabularies of the `gpt2` model (byte-level BPE) with one of two pre-tokenizers: `qwen2` (Qwen3, Bonsai) and `smollm` (SmolLM2 and mindX's `mindx-genN`). It also turns ids back into bytes, and collects the tokens that end generation as llama-vocab.cpp does. It follows llama.cpp's three steps: 1. **Special tokens first.** CONTROL (3) and USER_DEFINED (4) tokens are cut out of the text, longest first. Without `parse_special` only USER_DEFINED ones are. 2. **Pre-tokenize** each remaining span. `qwen2`'s pattern is written out alternative by alternative (ECMAScript semantics). `smollm` runs two passes: digits cut out one by one, then GPT-2's pattern inside each piece. 3. **Byte-level BPE** on each piece: its UTF-8 bytes become GPT-2's printable byte characters, each a token; the adjacent pair with the lowest merge rank is merged, the leftmost on ties, until none is left. The tokenizer was step one of P3 (Qwen3 / Bonsai, `qwen2`); the `smollm` pre-tokenizer came in 0.3.4 (O4) for SmolLM2 and mindX's `mindx-genN`. `unicode_letters.rs` holds Unicode general category L (Lu, Ll, Lt, Lm, Lo) as 622 sorted inclusive ranges, generated from Python's `unicodedata` (Unicode 13.0.0). It exists because `\p{L}` is category L, and Rust's `char::is_alphabetic` is the Alphabetic property, which is not the same set. Callers: the native engine (`native.rs`: prompts, grammar pieces, end-of-generation ids, since 0.3.7 DRY's sequence breakers, since 0.3.8 the bytes of each logprobs entry and its top tokens); `bankml serve --native` (`POST /tokenize`, every chat request, and a stop word's tokens, which llama-server drops from the logprobs); `bankml tokenize` and `bankml generate` (`main.rs`); and `native::header_info`, the registry's playability check, which calls `check_gguf`. ## Technical usage ```rust pub fn check_gguf(h: &crate::gguf::Header) -> Result<(), String> pub enum Pre { Qwen2, Smollm } impl Tokenizer { pub fn from_gguf(path: &Path) -> Result pub fn new(tokens: Vec, mut types: Vec, merges: &[String]) -> Result pub fn encode(&self, text: &str, parse_special: bool) -> Vec pub fn token_bytes(&self, id: u32) -> Vec pub fn piece(&self, id: u32) -> Vec pub fn eog_ids(&self, kv_ids: &[u32]) -> Vec pub fn n_tokens(&self) -> usize pub fn token_type(&self, id: u32) -> Option pub fn id(&self, text: &str) -> Option } pub fn pretokenize(text: &str) -> Vec<&str> // qwen2 pub fn pretokenize_smollm(text: &str) -> Vec<&str> // smollm pub fn is_letter(c: char) -> bool // unicode_letters.rs ``` - `from_gguf` reads `tokenizer.ggml.{model,pre,tokens,token_type,merges}` straight from the GGUF header through bankml's memory map, and refuses any vocabulary other than `gpt2/qwen2` and `gpt2/smollm`. - `new` applies llama-vocab.cpp's type overrides by text: every end-of-generation name and each fill-in-the-middle role's token become CONTROL; gpt-oss's four channel markers become USER_DEFINED. So ``, NORMAL in the Qwen3 file, is parsed as a special token. - `token_bytes(id)` is the detokenizer without special tokens: a control token gives nothing. - `piece(id)` is `token_to_piece(…, special = true)`: the bytes a grammar matches (`grammar::Vocab` is built from it). - `eog_ids(kv_ids)` is llama-vocab.cpp's end set: the FIM pad / repo / separator tokens, every name in its end-of-turn list, and the GGUF's own eos / eot / eom ids, with b11192's two exceptions. From the command line and over HTTP: ```sh echo -n "Hello world" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf # {"tokens": [...]} echo -n "<|im_start|>" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf --no-special curl -s 127.0.0.1:PORT/tokenize -d '{"content": "Hello", "parse_special": true}' # serve --native ``` `parse_special` defaults to true on `/tokenize`; `bankml tokenize` parses specials unless `--no-special` is given. ## How it is verified - Unit tests: `pretokenizer_shapes`, `smollm_pretokenizer_shapes`, `byte_chars_are_gpt2s`. - `oracle_tokenizer` (`#[ignore]`, in the gate): `testing/tokenizer_oracle.py` records llama-server b11192's own `/tokenize` on a corpus (the repository's documents and Savante's canon, the chat markers, edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with joiners, a seeded fuzz set of 2,000 strings from 17 Unicode ranges), with special tokens parsed and not. The test re-derives every case and requires the same ids in the same order: **4,346 of 4,346** in the 0.3.6 gate record (`testing/results/0.3.6.txt`). - `oracle_tokenizer_smollm`: the same corpus on SmolLM2's vocabulary against llama-server running SmolLM2, **4,346 of 4,346** (docs/oracles.md §5d; the same record). - Indirectly, every greedy, sampling, conversation, JSON and logprobs oracle: they compare token ids (and logprobs entries their bytes), so a tokenizer error shows there too. `oracle_grammar_masks` checks `piece` and the end set: 151,669 of 151,669 tokens and an end set of 6. The oracle has caught real differences: which special tokens split the text (Qwen3's `` markers are USER_DEFINED and split even without `parse_special`), the letter class, and in 0.3.4 `` taken as text in 2 of 4,346 cases before the type override was added. ## Design notes - **`qwen2`'s pattern**, with alternatives tried in order at each position (ECMAScript semantics; the backtracking results are written out in `pretokenize`): `'s|'t|'re|'ve|'m|'ll|'d` (either case) · `[^\r\n\p{L}\p{N}]?\p{L}+` · `\p{N}` · ` ?[^\s\p{L}\p{N}]+[\r\n]*` · `\s*[\r\n]+` · `\s+(?!\S)` · `\s+`. - **`smollm`**, read from llama-vocab.cpp and unicode.cpp b11192, runs two passes as llama.cpp does. First `\p{N}` cuts every digit out on its own: this is the `std::regex` pass over the collapsed text, which matches ASCII digits and non-ASCII codepoints whose only category is Number and that are not whitespace. Then GPT-2's pattern, `'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)`, runs inside each piece (`unicode_regex_split_custom_gpt2`): the end of a piece is the end of its text, so a space before a digit stays alone. Its oracle is llama-server's `/tokenize` on the SmolLM2 vocabulary. - Special tokens are cut out by llama-vocab's `tokenizer_st_partition`. ## Advantages and efficiency - **The same ids as llama.cpp, without llama.cpp.** No Python tokenizer, no regex engine, no crate. A prompt's token count, the prompt cache's reuse and every downstream oracle depend on this being exact. - **Header only.** `from_gguf` reads the vocabulary keys and skips the rest of the header; the tensors are not touched. Every length read is checked against the file's size before use (`Rd::take`, `Rd::len`), so a truncated or hostile header gives an `Err`, not a panic or an over-read. - **Precomputed tables.** Merges are keyed by token-id pairs `(left, right) → (rank, merged id)`; each byte's token is a 256-entry array; special tokens are sorted longest first once, at load. - **The letter table.** `is_letter` is a binary search over 622 ranges. - **Rust practice.** Zero dependencies, no `unsafe`, `Result` errors that say what is not reproduced. Where llama.cpp's behaviour depends on hash-map order, the vocabulary is refused rather than guessed. - **Next** (docs/TODO.md 0.6.0): the `llama3` pre-tokenizer, part of Llama 3.x support, which is refused today. ## Limitations - Only `gpt2/qwen2` and `gpt2/smollm` are reproduced. Any other model or pre-tokenizer is refused. - A header with any `tokenizer.ggml.fim_*` key is refused: llama.cpp then skips its detection by text for that role. - A vocabulary with two candidates for one fill-in-the-middle role is refused (llama.cpp picks one in hash order). - A merge naming a string outside the vocabulary is refused (llama.cpp merges by text; none of the pinned vocabularies has one). - `encode` never adds BOS or EOS (`add_special = false`). - A byte with no token of its own (SmolLM2 lacks 21) is dropped, as llama-vocab's fallback drops it. - The letter table is Unicode 13.0.0 data. ## See also - [../oracles.md](../oracles.md) §1b, §5d — the tokenizer oracle - [../usage.md](../usage.md) §13 — `bankml tokenize` - [../OLLAMA.md](../OLLAMA.md) — O4 (SmolLM2's tokenizer) and the converter's pre-tokenizer naming - [../TODO.md](../TODO.md) — 0.6.0, more models - Sibling pages: [chat.md](chat.md), [grammar.md](grammar.md), [forward.md](forward.md), [native.md](native.md), [serve.md](serve.md), [gguf.md](gguf.md)