Download docs/modules/tokenizer.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 9.03 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/tokenizer.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/tokenizer.md
-
curl -L -o tokenizer.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/tokenizer.md
bankML/tokenizer.rs — byte-level BPE, token-identical to llama.cpp b11192
This page also covers bankML/unicode_letters.rs, the tokenizer's letter table.
Summary
tokenizer.rs turns text into token ids exactly as llama.cpp b11192's llama_tokenize(…, add_special = false, parse_special) does, for GGUF vocabularies of the gpt2 model (byte-level BPE) with one of two pre-tokenizers:
qwen2 (Qwen3, Bonsai) and smollm (SmolLM2 and mindX's mindx-genN). It also turns ids back into bytes, and
collects the tokens that end generation as llama-vocab.cpp does.
It follows llama.cpp's three steps:
- Special tokens first. CONTROL (3) and USER_DEFINED (4) tokens are cut out of the text, longest first. Without
parse_specialonly USER_DEFINED ones are. - Pre-tokenize each remaining span.
qwen2's pattern is written out alternative by alternative (ECMAScript semantics).smollmruns two passes: digits cut out one by one, then GPT-2's pattern inside each piece. - Byte-level BPE on each piece: its UTF-8 bytes become GPT-2's printable byte characters, each a token; the adjacent pair with the lowest merge rank is merged, the leftmost on ties, until none is left.
The tokenizer was step one of P3 (Qwen3 / Bonsai, qwen2); the smollm pre-tokenizer came in 0.3.4 (O4) for
SmolLM2 and mindX's mindx-genN.
unicode_letters.rs holds Unicode general category L (Lu, Ll, Lt, Lm, Lo) as 622 sorted inclusive ranges, generated
from Python's unicodedata (Unicode 13.0.0). It exists because \p{L} is category L, and Rust's
char::is_alphabetic is the Alphabetic property, which is not the same set.
Callers: the native engine (native.rs: prompts, grammar pieces, end-of-generation ids, since 0.3.7 DRY's sequence
breakers, since 0.3.8 the bytes of each logprobs entry and its top tokens); bankml serve --native (POST /tokenize,
every chat request, and a stop word's tokens, which llama-server drops from the logprobs); bankml tokenize and
bankml generate (main.rs); and native::header_info, the registry's playability check, which calls check_gguf.
Technical usage
pub fn check_gguf(h: &crate::gguf::Header) -> Result<(), String>
pub enum Pre { Qwen2, Smollm }
impl Tokenizer {
pub fn from_gguf(path: &Path) -> Result<Self, String>
pub fn new(tokens: Vec<String>, mut types: Vec<i32>, merges: &[String]) -> Result<Self, String>
pub fn encode(&self, text: &str, parse_special: bool) -> Vec<u32>
pub fn token_bytes(&self, id: u32) -> Vec<u8>
pub fn piece(&self, id: u32) -> Vec<u8>
pub fn eog_ids(&self, kv_ids: &[u32]) -> Vec<u32>
pub fn n_tokens(&self) -> usize
pub fn token_type(&self, id: u32) -> Option<i32>
pub fn id(&self, text: &str) -> Option<u32>
}
pub fn pretokenize(text: &str) -> Vec<&str> // qwen2
pub fn pretokenize_smollm(text: &str) -> Vec<&str> // smollm
pub fn is_letter(c: char) -> bool // unicode_letters.rs
from_ggufreadstokenizer.ggml.{model,pre,tokens,token_type,merges}straight from the GGUF header through bankml's memory map, and refuses any vocabulary other thangpt2/qwen2andgpt2/smollm.newapplies llama-vocab.cpp's type overrides by text: every end-of-generation name and each fill-in-the-middle role's token become CONTROL; gpt-oss's four channel markers become USER_DEFINED. So</s>, NORMAL in the Qwen3 file, is parsed as a special token.token_bytes(id)is the detokenizer without special tokens: a control token gives nothing.piece(id)istoken_to_piece(…, special = true): the bytes a grammar matches (grammar::Vocabis built from it).eog_ids(kv_ids)is llama-vocab.cpp's end set: the FIM pad / repo / separator tokens, every name in its end-of-turn list, and the GGUF's own eos / eot / eom ids, with b11192's two exceptions.
From the command line and over HTTP:
echo -n "Hello world" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf # {"tokens": [...]}
echo -n "<|im_start|>" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf --no-special
curl -s 127.0.0.1:PORT/tokenize -d '{"content": "Hello", "parse_special": true}' # serve --native
parse_special defaults to true on /tokenize; bankml tokenize parses specials unless --no-special is given.
How it is verified
- Unit tests:
pretokenizer_shapes,smollm_pretokenizer_shapes,byte_chars_are_gpt2s. oracle_tokenizer(#[ignore], in the gate):testing/tokenizer_oracle.pyrecords llama-server b11192's own/tokenizeon a corpus (the repository's documents and Savante's canon, the chat markers, edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with joiners, a seeded fuzz set of 2,000 strings from 17 Unicode ranges), with special tokens parsed and not. The test re-derives every case and requires the same ids in the same order: 4,346 of 4,346 in the 0.3.6 gate record (testing/results/0.3.6.txt).oracle_tokenizer_smollm: the same corpus on SmolLM2's vocabulary against llama-server running SmolLM2, 4,346 of 4,346 (docs/oracles.md §5d; the same record).- Indirectly, every greedy, sampling, conversation, JSON and logprobs oracle: they compare token ids (and logprobs
entries their bytes), so a tokenizer error shows there too.
oracle_grammar_maskscheckspieceand the end set: 151,669 of 151,669 tokens and an end set of 6.
The oracle has caught real differences: which special tokens split the text (Qwen3's <think> markers are
USER_DEFINED and split even without parse_special), the letter class, and in 0.3.4 </s> taken as text in 2 of
4,346 cases before the type override was added.
Design notes
qwen2's pattern, with alternatives tried in order at each position (ECMAScript semantics; the backtracking results are written out inpretokenize):'s|'t|'re|'ve|'m|'ll|'d(either case) ·[^\r\n\p{L}\p{N}]?\p{L}+·\p{N}·?[^\s\p{L}\p{N}]+[\r\n]*·\s*[\r\n]+·\s+(?!\S)·\s+.smollm, read from llama-vocab.cpp and unicode.cpp b11192, runs two passes as llama.cpp does. First\p{N}cuts every digit out on its own: this is thestd::regexpass over the collapsed text, which matches ASCII digits and non-ASCII codepoints whose only category is Number and that are not whitespace. Then GPT-2's pattern,'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S), runs inside each piece (unicode_regex_split_custom_gpt2): the end of a piece is the end of its text, so a space before a digit stays alone. Its oracle is llama-server's/tokenizeon the SmolLM2 vocabulary.- Special tokens are cut out by llama-vocab's
tokenizer_st_partition.
Advantages and efficiency
- The same ids as llama.cpp, without llama.cpp. No Python tokenizer, no regex engine, no crate. A prompt's token count, the prompt cache's reuse and every downstream oracle depend on this being exact.
- Header only.
from_ggufreads the vocabulary keys and skips the rest of the header; the tensors are not touched. Every length read is checked against the file's size before use (Rd::take,Rd::len), so a truncated or hostile header gives anErr, not a panic or an over-read. - Precomputed tables. Merges are keyed by token-id pairs
(left, right) → (rank, merged id); each byte's token is a 256-entry array; special tokens are sorted longest first once, at load. - The letter table.
is_letteris a binary search over 622 ranges. - Rust practice. Zero dependencies, no
unsafe,Resulterrors that say what is not reproduced. Where llama.cpp's behaviour depends on hash-map order, the vocabulary is refused rather than guessed. - Next (docs/TODO.md 0.6.0): the
llama3pre-tokenizer, part of Llama 3.x support, which is refused today.
Limitations
- Only
gpt2/qwen2andgpt2/smollmare reproduced. Any other model or pre-tokenizer is refused. - A header with any
tokenizer.ggml.fim_*key is refused: llama.cpp then skips its detection by text for that role. - A vocabulary with two candidates for one fill-in-the-middle role is refused (llama.cpp picks one in hash order).
- A merge naming a string outside the vocabulary is refused (llama.cpp merges by text; none of the pinned vocabularies has one).
encodenever adds BOS or EOS (add_special = false).- A byte with no token of its own (SmolLM2 lacks 21) is dropped, as llama-vocab's fallback drops it.
- The letter table is Unicode 13.0.0 data.
See also
- ../oracles.md §1b, §5d — the tokenizer oracle
- ../usage.md §13 —
bankml tokenize - ../OLLAMA.md — O4 (SmolLM2's tokenizer) and the converter's pre-tokenizer naming
- ../TODO.md — 0.6.0, more models
- Sibling pages: chat.md, grammar.md, forward.md, native.md, serve.md, gguf.md