|
Download docs/modules/tokenizer.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 9.03 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/tokenizer.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/tokenizer.md
-
curl -L -o tokenizer.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/tokenizer.md
9.03 kB
| # `bankML/tokenizer.rs` — byte-level BPE, token-identical to llama.cpp b11192 | |
| This page also covers `bankML/unicode_letters.rs`, the tokenizer's letter table. | |
| ## Summary | |
| `tokenizer.rs` turns text into token ids exactly as llama.cpp b11192's `llama_tokenize(…, add_special = false, | |
| parse_special)` does, for GGUF vocabularies of the `gpt2` model (byte-level BPE) with one of two pre-tokenizers: | |
| `qwen2` (Qwen3, Bonsai) and `smollm` (SmolLM2 and mindX's `mindx-genN`). It also turns ids back into bytes, and | |
| collects the tokens that end generation as llama-vocab.cpp does. | |
| It follows llama.cpp's three steps: | |
| 1. **Special tokens first.** CONTROL (3) and USER_DEFINED (4) tokens are cut out of the text, longest first. Without | |
| `parse_special` only USER_DEFINED ones are. | |
| 2. **Pre-tokenize** each remaining span. `qwen2`'s pattern is written out alternative by alternative (ECMAScript | |
| semantics). `smollm` runs two passes: digits cut out one by one, then GPT-2's pattern inside each piece. | |
| 3. **Byte-level BPE** on each piece: its UTF-8 bytes become GPT-2's printable byte characters, each a token; the | |
| adjacent pair with the lowest merge rank is merged, the leftmost on ties, until none is left. | |
| The tokenizer was step one of P3 (Qwen3 / Bonsai, `qwen2`); the `smollm` pre-tokenizer came in 0.3.4 (O4) for | |
| SmolLM2 and mindX's `mindx-genN`. | |
| `unicode_letters.rs` holds Unicode general category L (Lu, Ll, Lt, Lm, Lo) as 622 sorted inclusive ranges, generated | |
| from Python's `unicodedata` (Unicode 13.0.0). It exists because `\p{L}` is category L, and Rust's | |
| `char::is_alphabetic` is the Alphabetic property, which is not the same set. | |
| Callers: the native engine (`native.rs`: prompts, grammar pieces, end-of-generation ids, since 0.3.7 DRY's sequence | |
| breakers, since 0.3.8 the bytes of each logprobs entry and its top tokens); `bankml serve --native` (`POST /tokenize`, | |
| every chat request, and a stop word's tokens, which llama-server drops from the logprobs); `bankml tokenize` and | |
| `bankml generate` (`main.rs`); and `native::header_info`, the registry's playability check, which calls `check_gguf`. | |
| ## Technical usage | |
| ```rust | |
| pub fn check_gguf(h: &crate::gguf::Header) -> Result<(), String> | |
| pub enum Pre { Qwen2, Smollm } | |
| impl Tokenizer { | |
| pub fn from_gguf(path: &Path) -> Result<Self, String> | |
| pub fn new(tokens: Vec<String>, mut types: Vec<i32>, merges: &[String]) -> Result<Self, String> | |
| pub fn encode(&self, text: &str, parse_special: bool) -> Vec<u32> | |
| pub fn token_bytes(&self, id: u32) -> Vec<u8> | |
| pub fn piece(&self, id: u32) -> Vec<u8> | |
| pub fn eog_ids(&self, kv_ids: &[u32]) -> Vec<u32> | |
| pub fn n_tokens(&self) -> usize | |
| pub fn token_type(&self, id: u32) -> Option<i32> | |
| pub fn id(&self, text: &str) -> Option<u32> | |
| } | |
| pub fn pretokenize(text: &str) -> Vec<&str> // qwen2 | |
| pub fn pretokenize_smollm(text: &str) -> Vec<&str> // smollm | |
| pub fn is_letter(c: char) -> bool // unicode_letters.rs | |
| ``` | |
| - `from_gguf` reads `tokenizer.ggml.{model,pre,tokens,token_type,merges}` straight from the GGUF header through | |
| bankml's memory map, and refuses any vocabulary other than `gpt2/qwen2` and `gpt2/smollm`. | |
| - `new` applies llama-vocab.cpp's type overrides by text: every end-of-generation name and each fill-in-the-middle | |
| role's token become CONTROL; gpt-oss's four channel markers become USER_DEFINED. So `</s>`, NORMAL in the Qwen3 | |
| file, is parsed as a special token. | |
| - `token_bytes(id)` is the detokenizer without special tokens: a control token gives nothing. | |
| - `piece(id)` is `token_to_piece(…, special = true)`: the bytes a grammar matches (`grammar::Vocab` is built from it). | |
| - `eog_ids(kv_ids)` is llama-vocab.cpp's end set: the FIM pad / repo / separator tokens, every name in its end-of-turn | |
| list, and the GGUF's own eos / eot / eom ids, with b11192's two exceptions. | |
| From the command line and over HTTP: | |
| ```sh | |
| echo -n "Hello world" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf # {"tokens": [...]} | |
| echo -n "<|im_start|>" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf --no-special | |
| curl -s 127.0.0.1:PORT/tokenize -d '{"content": "Hello", "parse_special": true}' # serve --native | |
| ``` | |
| `parse_special` defaults to true on `/tokenize`; `bankml tokenize` parses specials unless `--no-special` is given. | |
| ## How it is verified | |
| - Unit tests: `pretokenizer_shapes`, `smollm_pretokenizer_shapes`, `byte_chars_are_gpt2s`. | |
| - `oracle_tokenizer` (`#[ignore]`, in the gate): `testing/tokenizer_oracle.py` records llama-server b11192's own | |
| `/tokenize` on a corpus (the repository's documents and Savante's canon, the chat markers, edge cases: contractions, | |
| CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with | |
| joiners, a seeded fuzz set of 2,000 strings from 17 Unicode ranges), with special tokens parsed and not. The test | |
| re-derives every case and requires the same ids in the same order: **4,346 of 4,346** in the 0.3.6 gate record | |
| (`testing/results/0.3.6.txt`). | |
| - `oracle_tokenizer_smollm`: the same corpus on SmolLM2's vocabulary against llama-server running SmolLM2, | |
| **4,346 of 4,346** (docs/oracles.md §5d; the same record). | |
| - Indirectly, every greedy, sampling, conversation, JSON and logprobs oracle: they compare token ids (and logprobs | |
| entries their bytes), so a tokenizer error shows there too. `oracle_grammar_masks` checks `piece` and the end set: 151,669 of 151,669 tokens and an end set of 6. | |
| The oracle has caught real differences: which special tokens split the text (Qwen3's `<think>` markers are | |
| USER_DEFINED and split even without `parse_special`), the letter class, and in 0.3.4 `</s>` taken as text in 2 of | |
| 4,346 cases before the type override was added. | |
| ## Design notes | |
| - **`qwen2`'s pattern**, with alternatives tried in order at each position (ECMAScript semantics; the | |
| backtracking results are written out in `pretokenize`): `'s|'t|'re|'ve|'m|'ll|'d` (either case) · | |
| `[^\r\n\p{L}\p{N}]?\p{L}+` · `\p{N}` · ` ?[^\s\p{L}\p{N}]+[\r\n]*` · `\s*[\r\n]+` · `\s+(?!\S)` · `\s+`. | |
| - **`smollm`**, read from llama-vocab.cpp and unicode.cpp b11192, runs two passes as llama.cpp does. First `\p{N}` | |
| cuts every digit out on its own: this is the `std::regex` pass over the collapsed text, which matches ASCII digits | |
| and non-ASCII codepoints whose only category is Number and that are not whitespace. Then GPT-2's pattern, | |
| `'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)`, runs inside each piece | |
| (`unicode_regex_split_custom_gpt2`): the end of a piece is the end of its text, so a space before a digit stays | |
| alone. Its oracle is llama-server's `/tokenize` on the SmolLM2 vocabulary. | |
| - Special tokens are cut out by llama-vocab's `tokenizer_st_partition`. | |
| ## Advantages and efficiency | |
| - **The same ids as llama.cpp, without llama.cpp.** No Python tokenizer, no regex engine, no crate. A prompt's token | |
| count, the prompt cache's reuse and every downstream oracle depend on this being exact. | |
| - **Header only.** `from_gguf` reads the vocabulary keys and skips the rest of the header; the tensors are not touched. | |
| Every length read is checked against the file's size before use (`Rd::take`, `Rd::len`), so a truncated or hostile | |
| header gives an `Err`, not a panic or an over-read. | |
| - **Precomputed tables.** Merges are keyed by token-id pairs `(left, right) → (rank, merged id)`; each byte's token | |
| is a 256-entry array; special tokens are sorted longest first once, at load. | |
| - **The letter table.** `is_letter` is a binary search over 622 ranges. | |
| - **Rust practice.** Zero dependencies, no `unsafe`, `Result` errors that say what is not reproduced. Where llama.cpp's | |
| behaviour depends on hash-map order, the vocabulary is refused rather than guessed. | |
| - **Next** (docs/TODO.md 0.6.0): the `llama3` pre-tokenizer, part of Llama 3.x support, which is refused today. | |
| ## Limitations | |
| - Only `gpt2/qwen2` and `gpt2/smollm` are reproduced. Any other model or pre-tokenizer is refused. | |
| - A header with any `tokenizer.ggml.fim_*` key is refused: llama.cpp then skips its detection by text for that role. | |
| - A vocabulary with two candidates for one fill-in-the-middle role is refused (llama.cpp picks one in hash order). | |
| - A merge naming a string outside the vocabulary is refused (llama.cpp merges by text; none of the pinned | |
| vocabularies has one). | |
| - `encode` never adds BOS or EOS (`add_special = false`). | |
| - A byte with no token of its own (SmolLM2 lacks 21) is dropped, as llama-vocab's fallback drops it. | |
| - The letter table is Unicode 13.0.0 data. | |
| ## See also | |
| - [../oracles.md](../oracles.md) §1b, §5d — the tokenizer oracle | |
| - [../usage.md](../usage.md) §13 — `bankml tokenize` | |
| - [../OLLAMA.md](../OLLAMA.md) — O4 (SmolLM2's tokenizer) and the converter's pre-tokenizer naming | |
| - [../TODO.md](../TODO.md) — 0.6.0, more models | |
| - Sibling pages: [chat.md](chat.md), [grammar.md](grammar.md), [forward.md](forward.md), [native.md](native.md), | |
| [serve.md](serve.md), [gguf.md](gguf.md) | |