Sansar 60M

A 63.2M-parameter Sanskrit language model trained from scratch, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models, their corpus and their tokenizer. Size class: the d60m preset: 8 layers, width 768 (56.6M parameters outside the embeddings).

  • This version: v0.1.0 = training run f0_slp1_uni8k_d60m, finished 2026-09-13 (experiment F0: first floor-product run on the frozen tokenizer).
  • Held-out score: 0.6434 bits per Devanagari byte (pooled over four held-out sets, excluding the Bhagavad-gītā; lower is better).
  • Base model: it continues Devanagari Sanskrit text; it is not instruction-tuned.
  • Why this checkpoint: The best 60M-class checkpoint on the held-out ex-Gītā number. It predates held-out masking, so its held-out numbers are somewhat flattered (see Training data).

Quick start

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "MuseMesh/sansar-60m"
tok = AutoTokenizer.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, revision="v0.1.0", trust_remote_code=True)

inputs = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")      # Devanagari in
out = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=100)
print(tok.decode(out[0], skip_special_tokens=True))              # Devanagari out

Needs transformers (tested with 5.18.0), torch and sentencepiece. trust_remote_code=True runs three small files from this repository: modeling_sansar.py (the network), tokenization_sansar.py and translit.py (Devanagari <-> SLP1); read them before you run them. The weights are stored in bfloat16 (transformers 5 loads them as such); pass dtype=torch.float32 for fp32, usually the faster choice on a CPU. The tokenizer adds no BOS/EOS: training records were separated by </s> only.

What this checkpoint wrote for that prompt (CPU, torch.manual_seed(0), 50 new tokens):

संस्कृतं नाम दैवी वाक् संस्कृतवाग्विज्ञानस्य । संस्कृतमिति प्रसिद्धम् । तदप्यसत् । संस्कृते हि संस्कृता लौकिकाः । संस्कृतमिति प्रसिद्धम् । अतो लौकिकमपि 

and with greedy decoding:

संस्कृतं नाम दैवी वाक् । असंस्कृतभाषा संस्कृतभाषा । असंस्कृतभाषा संस्कृतभाषा । असंस्कृतभाषा संस्कृतभाषा । असंस्कृतभाषा संस्कृतभाषा । असंस्कृत

To score text instead of generating it, call the model with labels=input_ids; loss is the mean negative log-likelihood per token in nats.

Model details

Architecture decoder-only transformer, pre-LayerNorm (LayerNorm without bias), no biases anywhere
Layers / heads / width 8 / 12 / 768 (head dim 64); MLP 4x = 3,072
Positions learned position table (512 x 768)
Attention causal SDPA
MLP activation GELU (tanh approximation)
Output head tied to the token embedding
Vocabulary 8,000 (SentencePiece unigram over SLP1, MuseMesh/sansar-sanskrit-tokenizer v0.1.0)
Context 512 tokens (about 3.5 kB of Devanagari text at the training data's 6.89 Devanagari bytes per token)
Parameters, total 63,173,376
Parameters, non-embedding 56,636,160
Token embedding 6,144,000
Position embedding 393,216
Weights in this repo bfloat16 safetensors (trained as fp32 master weights under bf16 autocast)

Training

Optimizer AdamW (fused): peak lr 0.001, betas (0.9, 0.95), eps 1e-8, weight decay 0.1 on every 2-D tensor (matrices and embeddings), none on LayerNorm gains
Learning-rate schedule linear warm-up over 500 steps, then cosine decay to 0.1 x peak at step 27,203
Batch 48 sequences x 2 accumulation steps x 512 tokens = 49,152 tokens per step
Steps 27,203
Tokens seen 1,337,081,856 (1.00 passes over 1,337.1M training tokens)
Gradient clipping global norm 1.0
Dropout 0.0
Precision bf16 autocast, fp32 master weights, torch.compile
Seed 1337
Data sampling 512-token windows at uniformly random offsets of the token stream (records separated by </s>), so passes are counted in expectation
Hardware 1x NVIDIA RTX 3060 12 GB, power-capped at 130 W
Wall time 7.5 h (26,959 s), 49.9k tokens/s
Compute cost own hardware, no cloud cost
Final validation loss 2.9783 nats/token = 0.6261 bits per Devanagari byte on the run's own validation split (0.5% of records by near-duplicate cluster; not comparable across data slices)

Training code: scripts/train/train.py, model.py of the Sansar repository; the exact arguments are in training/summary.json.

Training data

train_slice_all of corpus rebuild #3 (about 414M words): 14,079,179 training records, 1,337.1M tokens, 9.22 GB of Devanagari text; reference markers stripped. The run saw 1,337.1M tokens = one pass.

Held-out texts were excluded by exact dedup key only (containment masking came one run later, in F1): 80 of 80 sampled held-out Gītā verses occur inside training records, so the Gītā column is memorisation, and the other sets are somewhat flattered (F1 -> F1b measured this kind of leak at about 1.6% on the ex-Gītā number). Later 60M runs trained with full held-out masking score 0.68 ex-Gītā, but on a quarter of the tokens; see the history table.

The text comes from the Sansar corpus, which collects Sanskrit in Devanagari from public sources: classical e-text collections (GRETIL, SARIT, Muktabodha, the Digital Corpus of Sanskrit, DharmaNexus and others), Sanskrit Wikisource and Wikipedia, dictionaries, and the Sanskrit parts of web-crawl datasets (AI4Bharat Sangraha, IndicCorp, MADLAD-400, the sanskrit-monolingual-pretraining collection). Every record keeps its provenance and a licence tier (T0 permissive, T1 share-alike, T2 non-commercial, T3 no licence statement or all rights reserved). The training slice mixes all tiers: a large share is licensed for non-commercial use only and some sources state no licence, which is why the weights are released under CC BY-NC 4.0 (see Licence). Old archive.org OCR of printed books is left out (it measurably hurt the models).

Preparation: Unicode NFC; standalone / and // read as daṇḍa । and double daṇḍa ॥; machine reference markers (verse numbers of digital editions) stripped; transliterated to SLP1 and tokenized; one </s> after every record; a 0.5% validation split by near-duplicate cluster.

Evaluation

Metric: bits per Devanagari byte (lower is better): the model's negative log-likelihood of a text divided by the UTF-8 byte count of the same text in Devanagari, so models with different tokenizers are measured against the same denominator. Scored teacher-forced with scripts/train/eval_bpb.py: each set's records joined with </s> into one stream, 512-token windows with stride 256 (every scored token after the first window has at least 256 tokens of context), text cleaned exactly as the training text was.

Sets (E0, frozen before any model was trained and excluded from training): dcs_gold 3,000 sentences of the Digital Corpus of Sanskrit gold standard (classical), prose 2,470 prose passages, ood 2,000 web, Wikipedia and other out-of-domain texts, vedic 1,000 accented Ṛgveda pādas, gita all 700 verses of the Bhagavad-gītā. The headline is pooled excluding the Gītā (byte-weighted over the other four): the Gītā is quoted inside commentaries and epics throughout the training text, so its column measures memorisation.

ex-Gītā (headline) pooled, all five dcs_gold prose ood vedic gita (memorisation)
E0 sets (9,170 items) 0.6434 0.6282 0.6344 0.6238 0.6390 1.1091 0.2284

Verse completion (600 verses: 200 each from the Bhagavad-gītā, Mahābhārata and Rāmāyaṇa; the model gets the first half-verse and greedily writes the second, stopping at ॥ or a newline): chrF 0.112, exact match 0/600. chrF credits shared character n-grams, so it rewards plausible vocabulary even when the half-verse is not the right one.

Reference points

Same metric and sets. External models were scored zero-shot with their own tokenizers and a 2,048-token window (stride 1,024), which gives them more context than our 512-token window; their training data may contain the public ood and prose texts.

model parameters ex-Gītā clean_v1 ex-Gītā notes
Sansar 20M v0.1.0 27.1M 0.7177 n/a TOK-v3 screen, arm v0.1, 0.34B tokens
Sansar 60M v0.1.0 (this model) 63.2M 0.6434 n/a F0, 1.34B tokens
Sansar 125M v0.1.0 97.2M 0.6039 0.6254 F6-clean-2x, 2.66B tokens
Sansar 350M v0.1.0 318.4M 0.5720 0.5937 F7, 2.66B tokens
Sansar 350M v0.2.0 318.4M 0.5547 0.5773 F9, 5.09B tokens
Sansar 700M v0.1.0 704.1M 0.5307 0.5541 F10, 4.76B tokens
Krutrim-2-instruct (zero-shot, nf4) 12B 0.512 n/a general model, 2,048-token window
Sarvam-1 (zero-shot, fp16) 2.5B 0.750 n/a general model, 2,048-token window
Gemma 4 E2B (zero-shot, fp16) 4B 0.892 n/a general model, 2,048-token window

Checked before release

  • modeling_sansar.py against the training code (scripts/train/model.py), fp32 on CPU, 6 inputs (one held-out text per set, cropped to 512 tokens, and 512 random ids): with the same bf16 weights the logits are identical (max |difference| 0.0e+00); against the fp32 training weights the bf16 storage moves logits by at most 0.151.
  • Bits per byte on the first 20 records of each set with eval_bpb.py's own scoring: 0.57871 (training checkpoint, fp32) vs 0.57873 (this repo, bf16 weights) = +0.0033%.
  • Tokenizer: the same ids as the evaluation pipeline on 100/100 sample texts with fence_latin=False (95/100 with the default fence; the rest contain English words); Devanagari round trip exact on 100/100.
  • Left-padded batches give the same logits as single sequences (max |difference| 3.9e-05).

Limitations

  • Base model. It continues Sanskrit text. It has not been instruction-tuned or aligned: it does not follow instructions, answer questions or chat, and it will not refuse anything.
  • It makes things up. Output is fluent-looking Sanskrit that can be ungrammatical, mix registers and invent verses, authors and works. Do not use it as a source of quotations or facts; exact verse recall is close to zero (see verse completion).
  • Other languages in the data. The training text still holds some Hindi, Marathi and Pali lines that passed the Sanskrit filters of the time (a stricter filter came after this run), so the model can drift into Hindi.
  • Small and short. 63M parameters and a 512-token context. This implementation has no key/value cache, so long generations are slow (each new token re-reads the last 512).
  • Vedic is the weakest register: accent marks fragment the tokenization and Vedic text is a small share of the training data.
  • Script. Devanagari in and out. English words inside the text are fenced by the tokenizer so they come back unchanged; the model never saw the fence marks in training, so its predictions around them are weaker (AutoTokenizer.from_pretrained(..., fence_latin=False) reproduces the training encoding, but then Latin letters decode as Devanagari). IAST and other Indic scripts are not transliterated for you.
  • What the corpus says, the model says. Most of the text is religious, philosophical and classical literature, plus modern web and news Sanskrit; the model reproduces its views and its errors, including OCR errors that survived in web-crawl sources.
  • The Gītā is memorised in part (see the gita column), so do not read its score as generalisation.

Versions

Each version is a git tag on this repository; main is the newest. Pin one with revision="v0.1.0". A version is one training run, named by its run id in the Sansar experiment log.

version date training run tokens seen ex-Gītā clean_v1 ex-Gītā
v0.1.0 2026-09-13 f0_slp1_uni8k_d60m (F0) 1.34B 0.6434 n/a

All runs of this size (ex-Gītā on the same E0 sets):

run date recipe ex-Gītā status
E14 slp1_uni8k 2026-09-12 8L/12H/768, AdamW, 0.39B tokens 0.7117 not staged
F0 2026-09-13 8L/12H/768, AdamW, 1.34B tokens (1 pass, rebuild #3) 0.6434 v0.1.0
E-CTX/E-VOCAB/E-VEDIC base (i9small) 2026-10-03 8L/12H/768, F4 recipe, 0.33B tokens, held-out masked 0.6825 not staged
F10-AB b (filtered) 2026-10-04 8L/12H/768, F4 recipe, 0.31B tokens, held-out masked 0.6814 not staged

Files

file what
model.safetensors the weights, bfloat16 (no optimizer state)
config.json, generation_config.json architecture and default sampling settings
configuration_sansar.py, modeling_sansar.py the network for transformers (auto_map, trust_remote_code)
tokenizer.model, tokenization_sansar.py, translit.py, tokenizer_config.json, special_tokens_map.json the tokenizer (MuseMesh/sansar-sanskrit-tokenizer v0.1.0) with its Devanagari <-> SLP1 wrapper
eval/bpb.json, eval/verse_scores.json the evaluation outputs quoted above
eval/verification.json, eval/smoke_test.json the release checks and the sample generations
training/summary.json, training/data_meta.json every training argument, the loss curve's evaluation points, and the data preparation record
LICENSE, LICENSE-CODE, CHANGELOG.md licences and version history

Licence

This release is for research and non-commercial use. The training data includes texts licensed for non-commercial use only and texts without a licence statement, used here for research. A commercially licensed model, trained only on permissively licensed text, is planned as a separate release with its own version line.

  • Weights (model.safetensors) and tokenizer.model: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited".
  • Code (configuration_sansar.py, modeling_sansar.py, tokenization_sansar.py, translit.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Citation

@misc{sansar_60m_2026,
  title  = {Sansar 60M: a Sanskrit language model},
  author = {Muse Mesh},
  year   = {2026},
  note   = {v0.1.0},
  url    = {https://huggingface.co/MuseMesh/sansar-60m}
}

Contact: kushal@muse-mesh.com

Downloads last month
242
Safetensors
Model size
63.2M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MuseMesh/sansar-60m

Collection including MuseMesh/sansar-60m

Evaluation results

  • bits per Devanagari byte, pooled excluding the Bhagavad-gītā (E0 held-out sets) on Sansar E0 held-out Sanskrit sets
    self-reported
    0.643
  • bits per Devanagari byte, pooled over all five E0 sets on Sansar E0 held-out Sanskrit sets
    self-reported
    0.628
  • bits per Devanagari byte, dcs_gold on Sansar E0 held-out Sanskrit sets
    self-reported
    0.634
  • bits per Devanagari byte, prose on Sansar E0 held-out Sanskrit sets
    self-reported
    0.624
  • bits per Devanagari byte, ood on Sansar E0 held-out Sanskrit sets
    self-reported
    0.639
  • bits per Devanagari byte, vedic on Sansar E0 held-out Sanskrit sets
    self-reported
    1.109
  • bits per Devanagari byte, gita on Sansar E0 held-out Sanskrit sets
    self-reported
    0.228