Transformers
tokenizer
sentencepiece
indic
multilingual
code-mixed
romanized
unigram
bpe

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

SpearIndic Tokenizer v1

Experimental SentencePiece tokenizers for Telugu, Hindi, Tamil, Kannada, Malayalam and English, built for three ways Indians actually write:

Style Example
Native script నేను రేపు హైదరాబాద్ వెళ్తున్నాను.
Romanized nenu repu Hyderabad veltunnanu.
Code-mixed నేను రేపు office కి వెళ్తున్నాను, meeting ఉంది.

Status: research release (v1). Not yet used to train a language model.

Which tokenizer?

Location Algorithm Vocab Notes
repo root Unigram 64,000 ✅ default: best on every Indic language
bpe-64k/ BPE 64,000 slightly better on English
bpe-32k/ BPE 32,000 smaller vocabulary
bpe-16k/ BPE 16,000 smallest

Usage

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("SpearsoftLabs/spearindic-tokenizer")                          # Unigram-64k
tok_bpe = AutoTokenizer.from_pretrained("SpearsoftLabs/spearindic-tokenizer", subfolder="bpe-64k")

ids = tok.encode("నేను రేపు office కి వెళ్తున్నాను", add_special_tokens=False)
print(tok.convert_ids_to_tokens(ids))
print(tok.decode(ids))

With plain SentencePiece (identical ids, verified on FLORES-200):

import sentencepiece as spm
from huggingface_hub import hf_hub_download

sp = spm.SentencePieceProcessor(
    model_file=hf_hub_download("SpearsoftLabs/spearindic-tokenizer", "spearindic_unigram_64k.model")
)

If the repo is private, log in first (hf auth login) with a token that has access to the SpearsoftLabs org.

Behaviour you should know

  • Special tokens: 0 <unk>, 1 <s>, 2 </s>, 3 <pad>. Nothing is added automatically: add <s>/</s> yourself if your model needs them.
  • Lossless: byte fallback means unseen characters become UTF-8 byte tokens, never <unk>. decode(encode(x)) == x on 100% of FLORES-200 sentences. Spaces are preserved exactly.
  • No normalization inside the tokenizer. Training text was NFC-normalized, so call unicodedata.normalize("NFC", text) before encoding for the best compression.
  • Numbers are not forced into single digits.

Training data (~1M lines)

Style Source Lines
Native script ai4bharat/IndicCorpV2 120,000 per language
Romanized ai4bharat/IndicCMix (romanized_casual) 30,000 per language
Code-mixed ai4bharat/IndicCMix (native_script_codemixed) 30,000 per language
English ai4bharat/IndicCMix (english) 100,000

Data was streamed and sampled (deterministic hash sampling). Cleaning: Unicode NFC, whitespace normalized, blank lines removed, exact deduplication. Punctuation, case, numbers, URLs and emojis are kept.

SentencePiece settings: byte_fallback=True, character_coverage=0.9995, normalization_rule_name=identity, remove_extra_whitespaces=False, input_sentence_size=0 (all lines).

Benchmark

FLORES-200 devtest (1,012 sentences per language; hi, te, ta, kn, ml). Bytes per token: higher = fewer tokens.

Tokenizer Vocab Indic bytes/token Round-trip
SpearIndic Unigram-64k 64,000 11.61 100%
SpearIndic BPE-64k 64,000 11.33 100%
SpearIndic BPE-32k 32,000 9.90 100%
XLM-R base 250,002 9.33 76%
SpearIndic BPE-16k 16,000 8.54 100%
mT5 250,100 7.99 76%
GPT-4o o200k_base 200,019 7.51 100%
mBERT 119,547 6.09 52%
GPT-4 cl100k_base 100,277 1.82 100%
IndicBERTv2* 250,000 12.79 0.06%

* IndicBERTv2 compresses more but does not reproduce the original text, so it is not directly comparable.

Per language, Unigram-64k (bytes/token): Hindi 9.41 · Telugu 11.16 · Tamil 13.55 · Kannada 11.74 · Malayalam 12.49 · English 3.62

Romanized and code-mixed (500 held-out IndicCMix sentences per language and style), bytes/token:

Hindi Telugu Tamil Kannada Malayalam
Romanized: Unigram-64k 3.98 4.35 4.12 4.13 4.02
Romanized: o200k 3.70 4.04 3.94 3.69 3.64
Code-mixed: Unigram-64k 6.92 8.34 8.98 8.91 9.27
Code-mixed: o200k 6.95 6.86 7.22 7.08 7.30

Limitations

  • English is weaker than GPT tokenizers (3.6 vs ~4.9 bytes/token): the only English training data was IndicCMix. v1.1 will add general English.
  • Romanized and code-mixed results come from held-out rows of the same synthetic dataset used in training, so they are optimistic. Real user text may differ.
  • Hindi code-mixed only ties GPT-4o's tokenizer.
  • Only 5 Indic languages. Other Indic scripts (Bengali, Gujarati, …) fall back to bytes and use many tokens.

License & credits

The tokenizer files are released under the Apache License 2.0.

Built from:

Evaluated on FLORES-200 (evaluation only, not used for training).

Developed by Spearsoft AI Labs.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train SpearsoftLabs/spearindic-tokenizer