bankml / docs /modules /tokenizer.md
Gregory-L's picture
bankML: the whole source (github.com/cryptoAGI/bankml @ 12ae409) and its page, with the bankML persona; the live engine (Dockerfile, hf/start.sh) ready for Docker hardware
28c70af verified
|
Raw History Blame Contribute Delete
9.03 kB

bankML/tokenizer.rs — byte-level BPE, token-identical to llama.cpp b11192

This page also covers bankML/unicode_letters.rs, the tokenizer's letter table.

Summary

tokenizer.rs turns text into token ids exactly as llama.cpp b11192's llama_tokenize(…, add_special = false, parse_special) does, for GGUF vocabularies of the gpt2 model (byte-level BPE) with one of two pre-tokenizers: qwen2 (Qwen3, Bonsai) and smollm (SmolLM2 and mindX's mindx-genN). It also turns ids back into bytes, and collects the tokens that end generation as llama-vocab.cpp does.

It follows llama.cpp's three steps:

  1. Special tokens first. CONTROL (3) and USER_DEFINED (4) tokens are cut out of the text, longest first. Without parse_special only USER_DEFINED ones are.
  2. Pre-tokenize each remaining span. qwen2's pattern is written out alternative by alternative (ECMAScript semantics). smollm runs two passes: digits cut out one by one, then GPT-2's pattern inside each piece.
  3. Byte-level BPE on each piece: its UTF-8 bytes become GPT-2's printable byte characters, each a token; the adjacent pair with the lowest merge rank is merged, the leftmost on ties, until none is left.

The tokenizer was step one of P3 (Qwen3 / Bonsai, qwen2); the smollm pre-tokenizer came in 0.3.4 (O4) for SmolLM2 and mindX's mindx-genN.

unicode_letters.rs holds Unicode general category L (Lu, Ll, Lt, Lm, Lo) as 622 sorted inclusive ranges, generated from Python's unicodedata (Unicode 13.0.0). It exists because \p{L} is category L, and Rust's char::is_alphabetic is the Alphabetic property, which is not the same set.

Callers: the native engine (native.rs: prompts, grammar pieces, end-of-generation ids, since 0.3.7 DRY's sequence breakers, since 0.3.8 the bytes of each logprobs entry and its top tokens); bankml serve --native (POST /tokenize, every chat request, and a stop word's tokens, which llama-server drops from the logprobs); bankml tokenize and bankml generate (main.rs); and native::header_info, the registry's playability check, which calls check_gguf.

Technical usage

pub fn check_gguf(h: &crate::gguf::Header) -> Result<(), String>
pub enum Pre { Qwen2, Smollm }

impl Tokenizer {
    pub fn from_gguf(path: &Path) -> Result<Self, String>
    pub fn new(tokens: Vec<String>, mut types: Vec<i32>, merges: &[String]) -> Result<Self, String>
    pub fn encode(&self, text: &str, parse_special: bool) -> Vec<u32>
    pub fn token_bytes(&self, id: u32) -> Vec<u8>
    pub fn piece(&self, id: u32) -> Vec<u8>
    pub fn eog_ids(&self, kv_ids: &[u32]) -> Vec<u32>
    pub fn n_tokens(&self) -> usize
    pub fn token_type(&self, id: u32) -> Option<i32>
    pub fn id(&self, text: &str) -> Option<u32>
}

pub fn pretokenize(text: &str) -> Vec<&str>          // qwen2
pub fn pretokenize_smollm(text: &str) -> Vec<&str>   // smollm
pub fn is_letter(c: char) -> bool                    // unicode_letters.rs
  • from_gguf reads tokenizer.ggml.{model,pre,tokens,token_type,merges} straight from the GGUF header through bankml's memory map, and refuses any vocabulary other than gpt2/qwen2 and gpt2/smollm.
  • new applies llama-vocab.cpp's type overrides by text: every end-of-generation name and each fill-in-the-middle role's token become CONTROL; gpt-oss's four channel markers become USER_DEFINED. So </s>, NORMAL in the Qwen3 file, is parsed as a special token.
  • token_bytes(id) is the detokenizer without special tokens: a control token gives nothing.
  • piece(id) is token_to_piece(…, special = true): the bytes a grammar matches (grammar::Vocab is built from it).
  • eog_ids(kv_ids) is llama-vocab.cpp's end set: the FIM pad / repo / separator tokens, every name in its end-of-turn list, and the GGUF's own eos / eot / eom ids, with b11192's two exceptions.

From the command line and over HTTP:

echo -n "Hello world" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf               # {"tokens": [...]}
echo -n "<|im_start|>" | bankml tokenize .models/Bonsai-8B-Q1_0.gguf --no-special
curl -s 127.0.0.1:PORT/tokenize -d '{"content": "Hello", "parse_special": true}'  # serve --native

parse_special defaults to true on /tokenize; bankml tokenize parses specials unless --no-special is given.

How it is verified

  • Unit tests: pretokenizer_shapes, smollm_pretokenizer_shapes, byte_chars_are_gpt2s.
  • oracle_tokenizer (#[ignore], in the gate): testing/tokenizer_oracle.py records llama-server b11192's own /tokenize on a corpus (the repository's documents and Savante's canon, the chat markers, edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, CJK, right-to-left scripts, emoji with joiners, a seeded fuzz set of 2,000 strings from 17 Unicode ranges), with special tokens parsed and not. The test re-derives every case and requires the same ids in the same order: 4,346 of 4,346 in the 0.3.6 gate record (testing/results/0.3.6.txt).
  • oracle_tokenizer_smollm: the same corpus on SmolLM2's vocabulary against llama-server running SmolLM2, 4,346 of 4,346 (docs/oracles.md §5d; the same record).
  • Indirectly, every greedy, sampling, conversation, JSON and logprobs oracle: they compare token ids (and logprobs entries their bytes), so a tokenizer error shows there too. oracle_grammar_masks checks piece and the end set: 151,669 of 151,669 tokens and an end set of 6.

The oracle has caught real differences: which special tokens split the text (Qwen3's <think> markers are USER_DEFINED and split even without parse_special), the letter class, and in 0.3.4 </s> taken as text in 2 of 4,346 cases before the type override was added.

Design notes

  • qwen2's pattern, with alternatives tried in order at each position (ECMAScript semantics; the backtracking results are written out in pretokenize): 's|'t|'re|'ve|'m|'ll|'d (either case) · [^\r\n\p{L}\p{N}]?\p{L}+ · \p{N} · ?[^\s\p{L}\p{N}]+[\r\n]* · \s*[\r\n]+ · \s+(?!\S) · \s+.
  • smollm, read from llama-vocab.cpp and unicode.cpp b11192, runs two passes as llama.cpp does. First \p{N} cuts every digit out on its own: this is the std::regex pass over the collapsed text, which matches ASCII digits and non-ASCII codepoints whose only category is Number and that are not whitespace. Then GPT-2's pattern, 's|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S), runs inside each piece (unicode_regex_split_custom_gpt2): the end of a piece is the end of its text, so a space before a digit stays alone. Its oracle is llama-server's /tokenize on the SmolLM2 vocabulary.
  • Special tokens are cut out by llama-vocab's tokenizer_st_partition.

Advantages and efficiency

  • The same ids as llama.cpp, without llama.cpp. No Python tokenizer, no regex engine, no crate. A prompt's token count, the prompt cache's reuse and every downstream oracle depend on this being exact.
  • Header only. from_gguf reads the vocabulary keys and skips the rest of the header; the tensors are not touched. Every length read is checked against the file's size before use (Rd::take, Rd::len), so a truncated or hostile header gives an Err, not a panic or an over-read.
  • Precomputed tables. Merges are keyed by token-id pairs (left, right) → (rank, merged id); each byte's token is a 256-entry array; special tokens are sorted longest first once, at load.
  • The letter table. is_letter is a binary search over 622 ranges.
  • Rust practice. Zero dependencies, no unsafe, Result errors that say what is not reproduced. Where llama.cpp's behaviour depends on hash-map order, the vocabulary is refused rather than guessed.
  • Next (docs/TODO.md 0.6.0): the llama3 pre-tokenizer, part of Llama 3.x support, which is refused today.

Limitations

  • Only gpt2/qwen2 and gpt2/smollm are reproduced. Any other model or pre-tokenizer is refused.
  • A header with any tokenizer.ggml.fim_* key is refused: llama.cpp then skips its detection by text for that role.
  • A vocabulary with two candidates for one fill-in-the-middle role is refused (llama.cpp picks one in hash order).
  • A merge naming a string outside the vocabulary is refused (llama.cpp merges by text; none of the pinned vocabularies has one).
  • encode never adds BOS or EOS (add_special = false).
  • A byte with no token of its own (SmolLM2 lacks 21) is dropped, as llama-vocab's fallback drops it.
  • The letter table is Unicode 13.0.0 data.

See also