|
Download README.md from tachiwin/tokenizer_test2: direct link, hf CLI and curl.
- Browser
- Download file 1.32 kB
-
https://huggingface.co/tachiwin/tokenizer_test2/resolve/main/README.md
- Command line
-
hf download hf://tachiwin/tokenizer_test2/README.md
-
curl -L -o README.md https://huggingface.co/tachiwin/tokenizer_test2/resolve/main/README.md
1.32 kB
| library_name: tokenizers | |
| tags: | |
| - tokenizer | |
| - multilingual | |
| - byte-level-bpe | |
| - indigenous-languages | |
| - tachiwin | |
| # Tachiwin multilingual tokenizer | |
| A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers. | |
| ## Corpus weighting | |
| | Component | Target | | |
| |---|---:| | |
| | Modern + old exotic-language data | 70% | | |
| | English | 10% | | |
| | Spanish | 10% | | |
| | Code | 10% | | |
| The complete available exotic-language corpus is used as the 70% anchor. | |
| Its existing modern/old composition is preserved. | |
| ## Tokenizer | |
| - Model: BPE | |
| - Vocabulary target: 256,000 | |
| - Initial alphabet: complete ByteLevel alphabet | |
| - ByteLevel GPT-2 regex: disabled | |
| - Unicode normalizer: none | |
| - Special tokens: 126 | |
| - Human-language tags: 92 | |
| ## Corpus size | |
| Total materialized corpus: | |
| 113,921,401 bytes | |
| (0.106 GiB) | |
| ## Important training note | |
| The Hugging Face BPE trainer does not expose an internal resumable merge-state | |
| checkpoint. The recipe therefore treats the completed `tokenizer.json` as the | |
| training checkpoint: | |
| - corpus preparation is resumable; | |
| - recipe/statistics/checksums are stored in `recipe/`; | |
| - if `tokenizer.json` already exists, subsequent runs skip BPE training; | |
| - an interrupted BPE computation itself must be restarted. | |
| This avoids changing the training procedure merely to obtain artificial | |
| checkpointing. | |