Instructions to use MuseMesh/sansar-20m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/sansar-20m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MuseMesh/sansar-20m", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-20m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuseMesh/sansar-20m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuseMesh/sansar-20m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-20m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MuseMesh/sansar-20m
- SGLang
How to use MuseMesh/sansar-20m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-20m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-20m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuseMesh/sansar-20m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuseMesh/sansar-20m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MuseMesh/sansar-20m with Docker Model Runner:
docker model run hf.co/MuseMesh/sansar-20m
Sansar 20M
A 27.1M-parameter Sanskrit language model trained from scratch, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models, their corpus and their tokenizer. Size class: the d20m preset: 6 layers, width 512 (18.9M parameters outside the embeddings; 27.1M in total with the 8k vocabulary and untied head).
- This version: v0.1.0 = training run
tokv3_v01_d20m, finished 2026-10-02 (experiment TOK-v3 screen, arm v0.1: the released tokenizer, F4 recipe at d20m). - Held-out score: 0.7177 bits per Devanagari byte (pooled over four held-out sets, excluding the Bhagavad-gītā; lower is better).
- Base model: it continues Devanagari Sanskrit text; it is not instruction-tuned.
- Why this checkpoint: The best 20M-class checkpoint on the held-out ex-Gītā number (0.7177, against 0.7537 for the best 20M run of the earlier AdamW recipe). It was trained as the control arm of the tokenizer v0.2 screen, which kept v0.1.
Quick start
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "MuseMesh/sansar-20m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt") # Devanagari in
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True)) # Devanagari out
Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs three small files from this repository: modeling_sansar.py (the network), tokenization_sansar.py and translit.py (Devanagari <-> SLP1); read them before you run them. The weights are stored in bfloat16 (transformers 5 loads them as such); pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training records were separated by </s> only.
What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 50 new tokens):
संस्कृतं नाम दैवी वाक् । तस्मिस्तु - स्वोपासकस्य मानसवतो विवाहकाले आगत्य 'महाराज' नाम्ना वसन्तों विवाहकाले आगत्य तु ऽ
and with greedy decoding:
संस्कृतं नाम दैवी वाक् । अनयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थयोरर्थ
To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.
Model details
| Architecture | decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere |
| Layers / heads / width | 6 / 8 / 512 (head dim 64); MLP 4x = 2,048 |
| Positions | rotary (base 10,000; half-split pairing), no position table |
| Attention | causal SDPA; queries and keys RMS-normalised per head (QK-norm) |
| MLP activation | squared ReLU |
| Output head | separate (untied), zero-initialised |
| Vocabulary | 8,000 (SentencePiece unigram over SLP1, MuseMesh/sansar-sanskrit-tokenizer v0.1.0) |
| Context | 512 tokens (about 3.6 kB of Devanagari text at the training data's 6.94 Devanagari bytes per token) |
| Parameters, total | 27,073,024 |
| Parameters, non-embedding | 18,881,024 |
| Token embedding | 4,096,000 |
| Output head | 4,096,000 |
| Weights in this repo | bfloat16 safetensors (trained as fp32 master weights under bf16 autocast) |
Training
| Optimizer, block matrices | Muon (18,874,368 params): momentum 0.95 (Nesterov), 5 Newton-Schulz steps in bf16, peak lr 0.02, update scaled by sqrt(max(1, rows/cols)), no weight decay |
| Optimizer, everything else | AdamW: embedding + head (8,192,000 params) with weight decay 0.1, LayerNorm gains without; peak lr 0.001, betas (0.9, 0.95), eps 1e-8 |
| Learning-rate schedule | linear warm-up over 200 steps, then cosine decay to 0.1 x peak at step 6,968 (Muon and AdamW share it) |
| Batch | 96 sequences x 1 accumulation steps x 512 tokens = 49,152 tokens per step |
| Steps | 6,968 |
| Tokens seen | 342,491,136 (1.00 passes over 342.5M training tokens) |
| Gradient clipping | global norm 1.0 |
| Dropout | 0.0 |
| Precision | bf16 autocast, fp32 master weights, torch.compile |
| Seed | 1337 |
| Data sampling | 512-token windows at uniformly random offsets of the token stream (records separated by </s>), so passes are counted in expectation |
| Hardware | 1x NVIDIA RTX 3060 12 GB, power-capped at 150 W |
| Wall time | 51 min (3,067 s), 114k tokens/s |
| Compute cost | own hardware, no cloud cost |
| Final validation | loss 3.3781 nats/token = 0.6975 bits per Devanagari byte on the run's own validation split (0.5% of records by near-duplicate cluster; not comparable across data slices) |
Training code: scripts/train/train.py, model.py and muon.py of the Sansar repository; the exact arguments are in training/summary.json.
Training data
A 23.5% record subset of train_slice_plus_clean (records whose sha1(id) falls in the band [0.50, 0.735); 3,549,167 records, 2.38 GB of cleaned Devanagari), the same text the larger plus_clean models were trained on at one quarter of the size. After held-out masking: 3,530,019 training records, 342.5M tokens. The run saw 342.5M tokens = one pass.
Held-out texts excluded by dedup key and masked inside training records (28-character windows, stride 4: 36,077 records masked, 3.9M characters removed, 5,340 records dropped). The Gītā is still partly memorised through near-copies in commentaries, so it stays out of the headline number.
The text comes from the Sansar corpus, which collects Sanskrit in Devanagari from public sources: classical e-text collections (GRETIL, SARIT, Muktabodha, the Digital Corpus of Sanskrit, DharmaNexus and others), Sanskrit Wikisource and Wikipedia, dictionaries, and the Sanskrit parts of web-crawl datasets (AI4Bharat Sangraha, IndicCorp, MADLAD-400, the sanskrit-monolingual-pretraining collection). Every record keeps its provenance and a licence tier (T0 permissive, T1 share-alike, T2 non-commercial, T3 no licence statement or all rights reserved). The training slice mixes all tiers: a large share is licensed for non-commercial use only and some sources state no licence, which is why the weights are released under CC BY-NC 4.0 (see Licence). Old archive.org OCR of printed books is left out (it measurably hurt the models).
Preparation: Unicode NFC; standalone / and // read as daṇḍa । and double daṇḍa ॥; machine reference markers (verse numbers of digital editions) stripped; transliterated to SLP1 and tokenized; one </s> after every record; a 0.5% validation split by near-duplicate cluster.
Evaluation
Metric: bits per Devanagari byte (lower is better): the model's negative log-likelihood of a text divided by the UTF-8 byte count of the same text in Devanagari, so models with different tokenizers are measured against the same denominator. Scored teacher-forced with scripts/train/eval_bpb.py: each set's records joined with </s> into one stream, 512-token windows with stride 256 (every scored token after the first window has at least 256 tokens of context), text cleaned exactly as the training text was.
Sets (E0, frozen before any model was trained and excluded from training): dcs_gold 3,000 sentences of the Digital Corpus of Sanskrit gold standard (classical), prose 2,470 prose passages, ood 2,000 web, Wikipedia and other out-of-domain texts, vedic 1,000 accented Ṛgveda pādas, gita all 700 verses of the Bhagavad-gītā. The headline is pooled excluding the Gītā (byte-weighted over the other four): the Gītā is quoted inside commentaries and epics throughout the training text, so its column measures memorisation.
| ex-Gītā (headline) | pooled, all five | dcs_gold | prose | ood | vedic | gita (memorisation) | |
|---|---|---|---|---|---|---|---|
| E0 sets (9,170 items) | 0.7177 | 0.7102 | 0.7065 | 0.6887 | 0.7225 | 0.9685 | 0.5129 |
Verse completion (600 verses: 200 each from the Bhagavad-gītā, Mahābhārata and Rāmāyaṇa; the model gets the first half-verse and greedily writes the second, stopping at ॥ or a newline): chrF 0.090, exact match 0/600. chrF credits shared character n-grams, so it rewards plausible vocabulary even when the half-verse is not the right one.
Reference points
Same metric and sets. External models were scored zero-shot with their own tokenizers and a 2,048-token window (stride 1,024), which gives them more context than our 512-token window; their training data may contain the public ood and prose texts.
| model | parameters | ex-Gītā | clean_v1 ex-Gītā | notes |
|---|---|---|---|---|
| Sansar 20M v0.1.0 (this model) | 27.1M | 0.7177 | n/a | TOK-v3 screen, arm v0.1, 0.34B tokens |
| Sansar 60M v0.1.0 | 63.2M | 0.6434 | n/a | F0, 1.34B tokens |
| Sansar 125M v0.1.0 | 97.2M | 0.6039 | 0.6254 | F6-clean-2x, 2.66B tokens |
| Sansar 350M v0.1.0 | 318.4M | 0.5720 | 0.5937 | F7, 2.66B tokens |
| Sansar 350M v0.2.0 | 318.4M | 0.5547 | 0.5773 | F9, 5.09B tokens |
| Sansar 700M v0.1.0 | 704.1M | 0.5307 | 0.5541 | F10, 4.76B tokens |
| Krutrim-2-instruct (zero-shot, nf4) | 12B | 0.512 | n/a | general model, 2,048-token window |
| Sarvam-1 (zero-shot, fp16) | 2.5B | 0.750 | n/a | general model, 2,048-token window |
| Gemma 4 E2B (zero-shot, fp16) | 4B | 0.892 | n/a | general model, 2,048-token window |
Checked before release
modeling_sansar.pyagainst the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one held-out text per set, cropped to 512 tokens, and 512 random ids): with the same bf16 weights the logits are identical (max |difference| 0.0e+00); against the fp32 training weights the bf16 storage moves logits by at most 0.159.- Bits per byte on the first 20 records of each set with
eval_bpb.py's own scoring: 0.68017 (training checkpoint, fp32) vs 0.68017 (this repo, bf16 weights) = +0.0006%. - Tokenizer: the same ids as the evaluation pipeline on 100/100 sample texts with
fence_latin=False(95/100 with the default fence; the rest contain English words); Devanagari round trip exact on 100/100. - Left-padded batches give the same logits as single sequences (max |difference| 1.4e-05).
Limitations
- Base model. It continues Sanskrit text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
- It makes things up. Output is fluent-looking Sanskrit that can be ungrammatical, mix registers and invent verses, authors and works. Do not use it as a source of quotations or facts; exact verse recall is close to zero (see verse completion).
- Other languages in the data. The training text still holds some Hindi, Marathi and Pali lines that passed the Sanskrit filters of the time (a stricter filter came after this run), so the model can drift into Hindi.
- Small and short. 27M parameters and a 512-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the last 512).
- Vedic is the weakest register: accent marks fragment the tokenization and Vedic text is a small share of the training data.
- Script. Devanagari in and out. English words inside the text are fenced by the tokenizer so they come back unchanged; the model never saw the fence marks in training, so its predictions around them are weaker (
AutoTokenizer.from_pretrained(..., fence_latin=False)reproduces the training encoding, but then Latin letters decode as Devanagari). IAST and other Indic scripts are not transliterated for you. - What the corpus says, the model says. Most of the text is religious, philosophical and classical literature, plus modern web and news Sanskrit; the model reproduces its views and its errors, including OCR errors that survived in web-crawl sources.
- The Gītā is memorised in part (see the
gitacolumn), so do not read its score as generalisation.
Versions
Each version is a git tag on this repository; main is the newest. Pin one with revision="v0.1.0". A version is one training run, named by its run id in the Sansar experiment log.
| version | date | training run | tokens seen | ex-Gītā | clean_v1 ex-Gītā |
|---|---|---|---|---|---|
| v0.1.0 | 2026-10-02 | tokv3_v01_d20m (TOK-v3 screen, arm v0.1) |
0.34B | 0.7177 | n/a |
All runs of this size (ex-Gītā on the same E0 sets):
| run | date | recipe | ex-Gītā | status |
|---|---|---|---|---|
| E8 slp1_uni8k | 2026-09-12 | 6L/8H/512, AdamW, learned positions, tied head, 0.39B tokens, key-only held-out exclusion | 0.7537 | not staged |
| E9 slp1_uni8k (+ seed 2 on a Kaggle T4) | 2026-09-12 | 11L/6H/384 (same total params), AdamW, 0.39B tokens | 0.7552 / 0.7681 | not staged (seed 2: no checkpoint) |
| VISION-OCR arms A-D (8 runs) | 2026-10-03 | 6L/8H/512, F4 recipe, 90M tokens | 0.791-0.803 | no checkpoints kept |
| TOK-v3 screen, arm v0.1 | 2026-10-02 | 6L/8H/512, F4 recipe (Muon + rotary/QK-norm/untied head/ReLU^2), 0.34B tokens of the plus_clean slice | 0.7177 | v0.1.0 |
Files
| file | what |
|---|---|
model.safetensors |
the weights, bfloat16 (no optimizer state) |
config.json, generation_config.json |
architecture and default sampling settings |
configuration_sansar.py, modeling_sansar.py |
the network for transformers (auto_map, trust_remote_code) |
tokenizer.model, tokenization_sansar.py, translit.py, tokenizer_config.json, special_tokens_map.json |
the tokenizer (MuseMesh/sansar-sanskrit-tokenizer v0.1.0) with its Devanagari <-> SLP1 wrapper |
eval/bpb.json, eval/verse_scores.json |
the evaluation outputs quoted above |
eval/verification.json, eval/smoke_test.json |
the release checks and the sample generations |
training/summary.json, training/data_meta.json |
every training argument, the loss curve's evaluation points, and the data preparation record |
LICENSE, LICENSE-CODE, CHANGELOG.md |
licences and version history |
Licence
This release is for research and non-commercial use. The training data includes texts licensed for non-commercial use only and texts without a licence statement, used here for research. A commercially licensed model, trained only on permissively licensed text, is planned as a separate release with its own version line.
- Weights (
model.safetensors) andtokenizer.model: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited". - Code (
configuration_sansar.py,modeling_sansar.py,tokenization_sansar.py,translit.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Citation
@misc{sansar_20m_2026,
title = {Sansar 20M: a Sanskrit language model},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0},
url = {https://huggingface.co/MuseMesh/sansar-20m}
}
Contact: kushal@muse-mesh.com
- Downloads last month
- 234
Datasets used to train MuseMesh/sansar-20m
chronbmm/sanskrit-monolingual-pretraining
Collection including MuseMesh/sansar-20m
Evaluation results
- bits per Devanagari byte, pooled excluding the Bhagavad-gītā (E0 held-out sets) on Sansar E0 held-out Sanskrit setsself-reported0.718
- bits per Devanagari byte, pooled over all five E0 sets on Sansar E0 held-out Sanskrit setsself-reported0.710
- bits per Devanagari byte, dcs_gold on Sansar E0 held-out Sanskrit setsself-reported0.707
- bits per Devanagari byte, prose on Sansar E0 held-out Sanskrit setsself-reported0.689
- bits per Devanagari byte, ood on Sansar E0 held-out Sanskrit setsself-reported0.723
- bits per Devanagari byte, vedic on Sansar E0 held-out Sanskrit setsself-reported0.969
- bits per Devanagari byte, gita on Sansar E0 held-out Sanskrit setsself-reported0.513