README / README.md
ks2812's picture
Card: the training log is the Sansar Lab at saansar.com
cc6a6d1 verified
|
Raw History Blame Contribute Delete
4.35 kB
---
title: README
emoji: 🪷
colorFrom: yellow
colorTo: red
sdk: static
pinned: false
---
# Muse Mesh
We build language technology for **Sanskrit** and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at the **[Sansar Lab](https://saansar.com)**.
## Sansar: Sanskrit-only language models
**Sansar** is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model.
| Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ |
|---|---|---|
| [sansar-700m](https://huggingface.co/MuseMesh/sansar-700m) | 704M | **0.5307** |
| [sansar-350m](https://huggingface.co/MuseMesh/sansar-350m) | 318M | 0.5547 |
| [sansar-125m](https://huggingface.co/MuseMesh/sansar-125m) | 97M | 0.6039 |
| [sansar-60m](https://huggingface.co/MuseMesh/sansar-60m) | 63M | 0.6434 |
| [sansar-20m](https://huggingface.co/MuseMesh/sansar-20m) | 27M | 0.7177 |
Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history.
Scored the same way, general base models need more bits per byte: Gemma 3 4B 0.6965, Qwen3-4B 0.7071, Llama 3.2 3B 0.7122, Sarvam-1 0.7465. sansar-700m beats all four on every held-out set with 704M parameters (Krutrim-2 12B, scored earlier in a pass that is not strictly comparable, is not in this list). A live demo is at **[saansar.com/demo](https://saansar.com/demo)**.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0]))
```
## English and Math models
Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. [Collection](https://huggingface.co/collections/MuseMesh/mume-english-and-math-language-models-6ac4f187f33abd6c9a9cdde6).
| Model | Parameters | Trained on | Result |
|---|---|---|---|
| [mume-english-125m](https://huggingface.co/MuseMesh/mume-english-125m) | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) |
| [mume-math-125m](https://huggingface.co/MuseMesh/mume-math-125m) | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch `sft` is a GSM8K/MATH fine-tune |
- **[mume-tokenizer-32k](https://huggingface.co/MuseMesh/mume-tokenizer-32k)**: the shared 32k unigram tokenizer.
- **[mume-eval-suites](https://huggingface.co/datasets/MuseMesh/mume-eval-suites)**: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests.
## Data and tokenizer
- **[Sansar Sanskrit Corpus](https://huggingface.co/datasets/MuseMesh/sansar-sanskrit-corpus)**: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed.
- **[sansar-sanskrit-tokenizer](https://huggingface.co/MuseMesh/sansar-sanskrit-tokenizer)**: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries.
## Licences
The corpus records keep their source licences. Model weights and the tokenizer are released under **CC BY-NC 4.0** for non-commercial research use, and the modelling code under **Apache-2.0**. For other uses, contact us.
**Links:** [muse-mesh.com](https://muse-mesh.com) · [GitHub](https://github.com/muse-mesh) · kushal@muse-mesh.com