Supernova tokenizer
Collection
Supernova tokenizer collection v5 strong in all of them implement the complete different Moe Architecture • 4 items • Updated • 1
How to use Supernova11c/Supernova-Nepali-Tokenizer with Transformers:
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Supernova11c/Supernova-Nepali-Tokenizer", device_map="auto")A high-performance, production-ready Byte-Level BPE tokenizer accuracy engineered for the Nepali language and Devanagari script. Developed as part of the Supernova project to enable efficient and accurate Nepali LLM processing.
Tested on the Supernova-teraillm dataset:
| Tokenizer | Tokens per Word | Efficiency |
|---|---|---|
| Supernova-Nepali (Ultra) | 3.79 | 2.20x Better |
| GPT-2 (Standard) | 8.21 | Baseline |
import time
from transformers import AutoTokenizer
# Load the dedicated Nepali tokenizer (Pure Tokenizer Repository)
model_id = "Supernova11c/Supernova-Nepali-Tokenizer"
print(f"Loading tokenizer for: {model_id}")
# Use clean_up_tokenization_spaces=False for BPE tokenizers to prevent warnings/corruption
tokenizer = AutoTokenizer.from_pretrained(model_id, clean_up_tokenization_spaces=False)
def stress_test_tokenizer(tokenizer):
print(f"\n--- Running Tokenizer Stress Test ---")
print(f"Tokenizer Class: {type(tokenizer).__name__}\n")
# 1. Edge Cases & Special Characters Test
edge_cases = [
"Hello, world! 🌍🚀", # Emojis & punctuation
" Multiple spaces and\nnewlines\t", # Whitespace handling
"The quick brown fox jumps over the lazy dog." * 50, # Repetition
"1234567890 -+*/=<>@#$%^&*()_[]{}|\\:;\"'.,?", # Symbols & Numbers
"नमस्ते संसार 🌟 नेपाल 🌍", # Nepali / Multi-lingual
"", # Empty string
]
print("1. Edge Case Testing:")
for i, text in enumerate(edge_cases):
try:
encoded = tokenizer.encode(text)
decoded = tokenizer.decode(encoded, skip_special_tokens=True)
match = "✓" if (text.strip() == decoded.strip() or not text) else "⚠️ (Whitespace diff)"
print(f" Test {i+1}: {match} | Length: {len(text)} chars -> {len(encoded)} tokens")
except Exception as e:
print(f" Test {i+1}: ❌ FAILED with error: {e}")
# 2. Throughput / Speed Test
print("\n2. Throughput Performance Test:")
sample_text = (
"नेपाल एक सुन्दर देश हो। यहाँ विभिन्न जातजाति र भाषाभाषीका मानिसहरू बसोबास गर्छन्। "
) * 500 # ~35,000 characters
num_iterations = 100
# Warmup
_ = tokenizer.encode(sample_text)
start_time = time.time()
for _ in range(num_iterations):
_ = tokenizer.encode(sample_text)
end_time = time.time()
total_time = end_time - start_time
total_chars = len(sample_text) * num_iterations
total_tokens = len(tokenizer.encode(sample_text)) * num_iterations
print(f" Processed {total_chars:,} characters in {total_time:.4f} seconds.")
print(f" Speed: {total_chars / total_time:,.2f} chars/sec")
print(f" Speed: {total_tokens / total_time:,.2f} tokens/sec")
# 3. Vocabulary & Configuration Check
print("\n3. Vocabulary & Configuration Check:")
print(f" Vocabulary Size: {len(tokenizer):,}")
print(f" Model Max Length: {getattr(tokenizer, 'model_max_length', 'N/A')}")
print(f" Pad Token: {tokenizer.pad_token} (ID: {tokenizer.pad_token_id})")
print(f" EOS Token: {tokenizer.eos_token} (ID: {tokenizer.eos_token_id})")
print("\n--- Stress Test Complete ---")
# Execute the test
stress_test_tokenizer(tokenizer)
[PAD], [UNK], [BOS], [EOS]Supernova text processing architecture is engineered for extreme, zero-overhead systems efficiency. Running entirely on standard CPU hardware without any GPU acceleration or heavy vector models, it delivers elite-tier throughput: