Instructions to use SpearsoftLabs/spearindic-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SpearsoftLabs/spearindic-tokenizer with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SpearsoftLabs/spearindic-tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SpearIndic Tokenizer v1
Experimental SentencePiece tokenizers for Telugu, Hindi, Tamil, Kannada, Malayalam and English, built for three ways Indians actually write:
| Style | Example |
|---|---|
| Native script | నేను రేపు హైదరాబాద్ వెళ్తున్నాను. |
| Romanized | nenu repu Hyderabad veltunnanu. |
| Code-mixed | నేను రేపు office కి వెళ్తున్నాను, meeting ఉంది. |
Status: research release (v1). Not yet used to train a language model.
Which tokenizer?
| Location | Algorithm | Vocab | Notes |
|---|---|---|---|
| repo root | Unigram | 64,000 | ✅ default: best on every Indic language |
bpe-64k/ |
BPE | 64,000 | slightly better on English |
bpe-32k/ |
BPE | 32,000 | smaller vocabulary |
bpe-16k/ |
BPE | 16,000 | smallest |
Usage
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("SpearsoftLabs/spearindic-tokenizer") # Unigram-64k
tok_bpe = AutoTokenizer.from_pretrained("SpearsoftLabs/spearindic-tokenizer", subfolder="bpe-64k")
ids = tok.encode("నేను రేపు office కి వెళ్తున్నాను", add_special_tokens=False)
print(tok.convert_ids_to_tokens(ids))
print(tok.decode(ids))
With plain SentencePiece (identical ids, verified on FLORES-200):
import sentencepiece as spm
from huggingface_hub import hf_hub_download
sp = spm.SentencePieceProcessor(
model_file=hf_hub_download("SpearsoftLabs/spearindic-tokenizer", "spearindic_unigram_64k.model")
)
If the repo is private, log in first (hf auth login) with a token that has access to the SpearsoftLabs org.
Behaviour you should know
- Special tokens:
0 <unk>,1 <s>,2 </s>,3 <pad>. Nothing is added automatically: add<s>/</s>yourself if your model needs them. - Lossless: byte fallback means unseen characters become UTF-8 byte tokens, never
<unk>.decode(encode(x)) == xon 100% of FLORES-200 sentences. Spaces are preserved exactly. - No normalization inside the tokenizer. Training text was NFC-normalized, so call
unicodedata.normalize("NFC", text)before encoding for the best compression. - Numbers are not forced into single digits.
Training data (~1M lines)
| Style | Source | Lines |
|---|---|---|
| Native script | ai4bharat/IndicCorpV2 | 120,000 per language |
| Romanized | ai4bharat/IndicCMix (romanized_casual) |
30,000 per language |
| Code-mixed | ai4bharat/IndicCMix (native_script_codemixed) |
30,000 per language |
| English | ai4bharat/IndicCMix (english) |
100,000 |
Data was streamed and sampled (deterministic hash sampling). Cleaning: Unicode NFC, whitespace normalized, blank lines removed, exact deduplication. Punctuation, case, numbers, URLs and emojis are kept.
SentencePiece settings: byte_fallback=True, character_coverage=0.9995,
normalization_rule_name=identity, remove_extra_whitespaces=False, input_sentence_size=0 (all lines).
Benchmark
FLORES-200 devtest (1,012 sentences per language; hi, te, ta, kn, ml). Bytes per token: higher = fewer tokens.
| Tokenizer | Vocab | Indic bytes/token | Round-trip |
|---|---|---|---|
| SpearIndic Unigram-64k | 64,000 | 11.61 | 100% |
| SpearIndic BPE-64k | 64,000 | 11.33 | 100% |
| SpearIndic BPE-32k | 32,000 | 9.90 | 100% |
| XLM-R base | 250,002 | 9.33 | 76% |
| SpearIndic BPE-16k | 16,000 | 8.54 | 100% |
| mT5 | 250,100 | 7.99 | 76% |
GPT-4o o200k_base |
200,019 | 7.51 | 100% |
| mBERT | 119,547 | 6.09 | 52% |
GPT-4 cl100k_base |
100,277 | 1.82 | 100% |
| IndicBERTv2* | 250,000 | 12.79 | 0.06% |
* IndicBERTv2 compresses more but does not reproduce the original text, so it is not directly comparable.
Per language, Unigram-64k (bytes/token): Hindi 9.41 · Telugu 11.16 · Tamil 13.55 · Kannada 11.74 · Malayalam 12.49 · English 3.62
Romanized and code-mixed (500 held-out IndicCMix sentences per language and style), bytes/token:
| Hindi | Telugu | Tamil | Kannada | Malayalam | |
|---|---|---|---|---|---|
| Romanized: Unigram-64k | 3.98 | 4.35 | 4.12 | 4.13 | 4.02 |
| Romanized: o200k | 3.70 | 4.04 | 3.94 | 3.69 | 3.64 |
| Code-mixed: Unigram-64k | 6.92 | 8.34 | 8.98 | 8.91 | 9.27 |
| Code-mixed: o200k | 6.95 | 6.86 | 7.22 | 7.08 | 7.30 |
Limitations
- English is weaker than GPT tokenizers (3.6 vs ~4.9 bytes/token): the only English training data was IndicCMix. v1.1 will add general English.
- Romanized and code-mixed results come from held-out rows of the same synthetic dataset used in training, so they are optimistic. Real user text may differ.
- Hindi code-mixed only ties GPT-4o's tokenizer.
- Only 5 Indic languages. Other Indic scripts (Bengali, Gujarati, …) fall back to bytes and use many tokens.
License & credits
The tokenizer files are released under the Apache License 2.0.
Built from:
- ai4bharat/IndicCorpV2 (CC-0)
- ai4bharat/IndicCMix (MIT)
Evaluated on FLORES-200 (evaluation only, not used for training).
Developed by Spearsoft AI Labs.