FineWeb-Edu BPE tokenizers

Three byte-level BPE tokenizers trained on the same ~12GB FineWeb-Edu sample (toklens/fw_edu_tokenizer_training_data). They differ only in the pre-tokenizer, so they isolate the effect of pre-tokenization on BPE.

Subfolder Pre-tokenizer tokeval sanity check Replaces
no_pretok None not yet published fineweb_edu_12gb_tokenizer_no_pretok
split_on_whitespace Split(Regex(r"[^\S\r\n]*[\n\r]+|[^\S\r\n]+"), "isolated") warn (0 fail, 1 warn) fineweb_edu_12gb_tokenizer_split_on_whitespace
sentencepiece Split(Regex(r"[^ ]+| [^ ]*"), "isolated") warn (0 fail, 1 warn) fineweb_edu_12gb_tokenizer_sentencepiece

Shared by all three:

  • BPE, vocab 32,000, NFC normalisation, ByteLevel pre-tokenization step and decoder
  • special tokens <s>=0, </s>=1, <pad>=2, <unk>=3
  • all 256 byte symbols in the base vocab, so every input round-trips and <unk> (kept for compatibility) never occurs

Each subfolder has its own README with the exact configuration, a sample encoding and the full tokeval sanity-check results.

Usage

from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("toklens/fineweb_edu_tokenizers", subfolder="sentencepiece")  # or no_pretok / split_on_whitespace

Why these replace the fineweb_edu_12gb_tokenizer_* repos

The earlier repos had no <unk> token in their vocab, and the sentencepiece one had no byte-level alphabet: characters outside its vocab were silently dropped on encode. These tokenizers are not drop-in replacements. The ids of <s>, </s> and <pad> are unchanged, but the merges differ, so models trained with the old tokenizers need the old repos.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support