FineWeb-Edu BPE tokenizers
Three byte-level BPE tokenizers trained on the same ~12GB FineWeb-Edu sample
(toklens/fw_edu_tokenizer_training_data).
They differ only in the pre-tokenizer, so they isolate the effect of pre-tokenization on BPE.
| Subfolder | Pre-tokenizer | tokeval sanity check | Replaces |
|---|---|---|---|
no_pretok |
None | not yet published | fineweb_edu_12gb_tokenizer_no_pretok |
split_on_whitespace |
Split(Regex(r"[^\S\r\n]*[\n\r]+|[^\S\r\n]+"), "isolated") |
warn (0 fail, 1 warn) | fineweb_edu_12gb_tokenizer_split_on_whitespace |
sentencepiece |
Split(Regex(r"[^ ]+| [^ ]*"), "isolated") |
warn (0 fail, 1 warn) | fineweb_edu_12gb_tokenizer_sentencepiece |
Shared by all three:
- BPE, vocab 32,000, NFC normalisation,
ByteLevelpre-tokenization step and decoder - special tokens
<s>=0,</s>=1,<pad>=2,<unk>=3 - all 256 byte symbols in the base vocab, so every input round-trips and
<unk>(kept for compatibility) never occurs
Each subfolder has its own README with the exact configuration, a sample encoding and the full tokeval sanity-check results.
Usage
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("toklens/fineweb_edu_tokenizers", subfolder="sentencepiece") # or no_pretok / split_on_whitespace
Why these replace the fineweb_edu_12gb_tokenizer_* repos
The earlier repos had no <unk> token in their vocab, and the sentencepiece one had no byte-level
alphabet: characters outside its vocab were silently dropped on encode. These tokenizers are not
drop-in replacements. The ids of <s>, </s> and <pad> are unchanged, but the merges differ, so
models trained with the old tokenizers need the old repos.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support