File size: 9,033 Bytes
28c70af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 | # `bankML/tokenizer.rs` — byte-level BPE, token-identical to llama.cpp b11192
This page also covers `bankML/unicode_letters.rs`, the tokenizer's letter table.
## Summary
`tokenizer.rs` turns text into token ids exactly as llama.cpp b11192's `llama_tokenize(…, add_special = false,
parse_special)` does, for GGUF vocabularies of the `gpt2` model (byte-level BPE) with one of two pre-tokenizers:
`qwen2` (Qwen3, Bonsai) and `smollm` (SmolLM2 and mindX's `mindx-genN`). It also turns ids back into bytes, and
collects the tokens that end generation as llama-vocab.cpp does.
It follows llama.cpp's three steps:
1. **Special tokens first.** CONTROL (3) and USER_DEFINED (4) tokens are cut out of the text, longest first. Without
`parse_special` only USER_DEFINED ones are.
2. **Pre-tokenize** each remaining span. `qwen2`'s pattern is written out alternative by alternative (ECMAScript
semantics). `smollm` runs two passes: digits cut out one by one, then GPT-2's pattern inside each piece.
3. **Byte-level BPE** on each piece: its UTF-8 bytes become GPT-2's printable byte characters, each a token; the
adjacent pair with the lowest merge rank is merged, the leftmost on ties, until none is left.
The tokenizer was step one of P3 (Qwen3 / Bonsai, `qwen2`); the `smollm` pre-tokenizer came in 0.3.4 (O4) for
SmolLM2 and mindX's `mindx-genN`.
`unicode_letters.rs` holds Unicode general category L (Lu, Ll, Lt, Lm, Lo) as 622 sorted inclusive ranges, generated
from Python's `unicodedata` (Unicode 13.0.0). It exists because `\p{L}` is category L, and Rust's
`char::is_alphabetic` is the Alphabetic property, which is not the same set.
Callers: the native engine (`native.rs`: prompts, grammar pieces, end-of-generation ids, since 0.3.7 DRY's sequence
breakers, since 0.3.8 the bytes of each logprobs entry and its top tokens); `bankml serve --native` (`POST /tokenize`,
every chat request, and a stop word's tokens, which llama-server drops from the logprobs); `bankml tokenize` and
`bankml generate` (`main.rs`); and `native::header_info`, the registry's playability check, which calls `check_gguf`.
## Technical usage
```rust
pub fn check_gguf(h: &crate::gguf::Header) -> Result<(), String>
pub enum Pre { Qwen2, Smollm }
impl Tokenizer {
pub fn from_gguf(path: &Path) -> Result<Self, String>
pub fn new(tokens: Vec<String>, mut types: Vec<i32>, merges: &[String]) -> Result<Self, String>
pub fn encode(&self, text: &str, parse_special: bool) -> Vec<u32>
pub fn token_bytes(&self, id: u32) -> Vec<u8>
pub fn piece(&self, id: u32) -> Vec<u8>
pub fn eog_ids(&self, kv_ids: &[u32]) -> Vec<u32>
pub fn n_tokens(&self) -> usize
pub fn token_type(&self, id: u32) -> Option<i32>
pub fn id(&self, text: &str) -> Option<u32>
}
pub fn pretokenize(text: &str) -> Vec<&str> // qwen2
pub fn pretokenize_smollm(text: &str) -> Vec<&str> // smollm
pub fn is_letter(c: char) -> bool // unicode_letters.rs
```
- `from_gguf` reads `tokenizer.ggml.{model,pre,tokens,token_type,merges}` straight from the GGUF header through
bankml's memory map, and refuses any vocabulary other than `gpt2/qwen2` and `gpt2/smollm`.
- `new` applies llama-vocab.cpp's type overrides by text: every end-of-generation name and each fill-in-the-middle
role's token become CONTROL; gpt-oss's four channel markers become USER_DEFINED. So `</s>`, NORMAL in the Qwen3
file, is parsed as a special token.
- `token_bytes(id)` is the detokenizer without special tokens: a control token gives nothing.
- `piece(id)` is `token_to_piece(…, special = true)`: the bytes a grammar matches (`grammar::Vocab` is built from it).
- `eog_ids(kv_ids)` is llama-vocab.cpp's end set: the FIM pad / repo / separator tokens, every name in its end-of-turn
list, and the GGUF's own eos / eot / eom ids, with b11192's two exceptions.
From the command line and over HTTP:
```sh
echo -n "Hello world" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf # {"tokens": [...]}
echo -n "<|im_start|>" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf --no-special
curl -s 127.0.0.1:PORT/tokenize -d '{"content": "Hello", "parse_special": true}' # serve --native
```
`parse_special` defaults to true on `/tokenize`; `bankml tokenize` parses specials unless `--no-special` is given.
## How it is verified
- Unit tests: `pretokenizer_shapes`, `smollm_pretokenizer_shapes`, `byte_chars_are_gpt2s`.
- `oracle_tokenizer` (`#[ignore]`, in the gate): `testing/tokenizer_oracle.py` records llama-server b11192's own
`/tokenize` on a corpus (the repository's documents and Savante's canon, the chat markers, edge cases: contractions,
CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with
joiners, a seeded fuzz set of 2,000 strings from 17 Unicode ranges), with special tokens parsed and not. The test
re-derives every case and requires the same ids in the same order: **4,346 of 4,346** in the 0.3.6 gate record
(`testing/results/0.3.6.txt`).
- `oracle_tokenizer_smollm`: the same corpus on SmolLM2's vocabulary against llama-server running SmolLM2,
**4,346 of 4,346** (docs/oracles.md §5d; the same record).
- Indirectly, every greedy, sampling, conversation, JSON and logprobs oracle: they compare token ids (and logprobs
entries their bytes), so a tokenizer error shows there too. `oracle_grammar_masks` checks `piece` and the end set: 151,669 of 151,669 tokens and an end set of 6.
The oracle has caught real differences: which special tokens split the text (Qwen3's `<think>` markers are
USER_DEFINED and split even without `parse_special`), the letter class, and in 0.3.4 `</s>` taken as text in 2 of
4,346 cases before the type override was added.
## Design notes
- **`qwen2`'s pattern**, with alternatives tried in order at each position (ECMAScript semantics; the
backtracking results are written out in `pretokenize`): `'s|'t|'re|'ve|'m|'ll|'d` (either case) ·
`[^\r\n\p{L}\p{N}]?\p{L}+` · `\p{N}` · ` ?[^\s\p{L}\p{N}]+[\r\n]*` · `\s*[\r\n]+` · `\s+(?!\S)` · `\s+`.
- **`smollm`**, read from llama-vocab.cpp and unicode.cpp b11192, runs two passes as llama.cpp does. First `\p{N}`
cuts every digit out on its own: this is the `std::regex` pass over the collapsed text, which matches ASCII digits
and non-ASCII codepoints whose only category is Number and that are not whitespace. Then GPT-2's pattern,
`'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)`, runs inside each piece
(`unicode_regex_split_custom_gpt2`): the end of a piece is the end of its text, so a space before a digit stays
alone. Its oracle is llama-server's `/tokenize` on the SmolLM2 vocabulary.
- Special tokens are cut out by llama-vocab's `tokenizer_st_partition`.
## Advantages and efficiency
- **The same ids as llama.cpp, without llama.cpp.** No Python tokenizer, no regex engine, no crate. A prompt's token
count, the prompt cache's reuse and every downstream oracle depend on this being exact.
- **Header only.** `from_gguf` reads the vocabulary keys and skips the rest of the header; the tensors are not touched.
Every length read is checked against the file's size before use (`Rd::take`, `Rd::len`), so a truncated or hostile
header gives an `Err`, not a panic or an over-read.
- **Precomputed tables.** Merges are keyed by token-id pairs `(left, right) → (rank, merged id)`; each byte's token
is a 256-entry array; special tokens are sorted longest first once, at load.
- **The letter table.** `is_letter` is a binary search over 622 ranges.
- **Rust practice.** Zero dependencies, no `unsafe`, `Result` errors that say what is not reproduced. Where llama.cpp's
behaviour depends on hash-map order, the vocabulary is refused rather than guessed.
- **Next** (docs/TODO.md 0.6.0): the `llama3` pre-tokenizer, part of Llama 3.x support, which is refused today.
## Limitations
- Only `gpt2/qwen2` and `gpt2/smollm` are reproduced. Any other model or pre-tokenizer is refused.
- A header with any `tokenizer.ggml.fim_*` key is refused: llama.cpp then skips its detection by text for that role.
- A vocabulary with two candidates for one fill-in-the-middle role is refused (llama.cpp picks one in hash order).
- A merge naming a string outside the vocabulary is refused (llama.cpp merges by text; none of the pinned
vocabularies has one).
- `encode` never adds BOS or EOS (`add_special = false`).
- A byte with no token of its own (SmolLM2 lacks 21) is dropped, as llama-vocab's fallback drops it.
- The letter table is Unicode 13.0.0 data.
## See also
- [../oracles.md](../oracles.md) §1b, §5d — the tokenizer oracle
- [../usage.md](../usage.md) §13 — `bankml tokenize`
- [../OLLAMA.md](../OLLAMA.md) — O4 (SmolLM2's tokenizer) and the converter's pre-tokenizer naming
- [../TODO.md](../TODO.md) — 0.6.0, more models
- Sibling pages: [chat.md](chat.md), [grammar.md](grammar.md), [forward.md](forward.md), [native.md](native.md),
[serve.md](serve.md), [gguf.md](gguf.md)
|