|
Download README.md from tachiwin/tokenizer_test2: direct link, hf CLI and curl.
- Browser
- Download file 1.32 kB
-
https://huggingface.co/tachiwin/tokenizer_test2/resolve/main/README.md
- Command line
-
hf download hf://tachiwin/tokenizer_test2/README.md
-
curl -L -o README.md https://huggingface.co/tachiwin/tokenizer_test2/resolve/main/README.md
1.32 kB
metadata
library_name: tokenizers
tags:
- tokenizer
- multilingual
- byte-level-bpe
- indigenous-languages
- tachiwin
Tachiwin multilingual tokenizer
A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.
Corpus weighting
| Component | Target |
|---|---|
| Modern + old exotic-language data | 70% |
| English | 10% |
| Spanish | 10% |
| Code | 10% |
The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.
Tokenizer
- Model: BPE
- Vocabulary target: 256,000
- Initial alphabet: complete ByteLevel alphabet
- ByteLevel GPT-2 regex: disabled
- Unicode normalizer: none
- Special tokens: 126
- Human-language tags: 92
Corpus size
Total materialized corpus:
113,921,401 bytes (0.106 GiB)
Important training note
The Hugging Face BPE trainer does not expose an internal resumable merge-state
checkpoint. The recipe therefore treats the completed tokenizer.json as the
training checkpoint:
- corpus preparation is resumable;
- recipe/statistics/checksums are stored in
recipe/; - if
tokenizer.jsonalready exists, subsequent runs skip BPE training; - an interrupted BPE computation itself must be restarted.
This avoids changing the training procedure merely to obtain artificial checkpointing.