|
Download README.md from MuseMesh/README: direct link, hf CLI and curl.
- Browser
- Download file 4.35 kB
-
https://huggingface.co/spaces/MuseMesh/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/MuseMesh/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/MuseMesh/README/resolve/main/README.md
4.35 kB
| title: README | |
| emoji: 🪷 | |
| colorFrom: yellow | |
| colorTo: red | |
| sdk: static | |
| pinned: false | |
| # Muse Mesh | |
| We build language technology for **Sanskrit** and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at the **[Sansar Lab](https://saansar.com)**. | |
| ## Sansar: Sanskrit-only language models | |
| **Sansar** is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model. | |
| | Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ | | |
| |---|---|---| | |
| | [sansar-700m](https://huggingface.co/MuseMesh/sansar-700m) | 704M | **0.5307** | | |
| | [sansar-350m](https://huggingface.co/MuseMesh/sansar-350m) | 318M | 0.5547 | | |
| | [sansar-125m](https://huggingface.co/MuseMesh/sansar-125m) | 97M | 0.6039 | | |
| | [sansar-60m](https://huggingface.co/MuseMesh/sansar-60m) | 63M | 0.6434 | | |
| | [sansar-20m](https://huggingface.co/MuseMesh/sansar-20m) | 27M | 0.7177 | | |
| Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history. | |
| Scored the same way, general base models need more bits per byte: Gemma 3 4B 0.6965, Qwen3-4B 0.7071, Llama 3.2 3B 0.7122, Sarvam-1 0.7465. sansar-700m beats all four on every held-out set with 704M parameters (Krutrim-2 12B, scored earlier in a pass that is not strictly comparable, is not in this list). A live demo is at **[saansar.com/demo](https://saansar.com/demo)**. | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True) | |
| ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt") | |
| print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0])) | |
| ``` | |
| ## English and Math models | |
| Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. [Collection](https://huggingface.co/collections/MuseMesh/mume-english-and-math-language-models-6ac4f187f33abd6c9a9cdde6). | |
| | Model | Parameters | Trained on | Result | | |
| |---|---|---|---| | |
| | [mume-english-125m](https://huggingface.co/MuseMesh/mume-english-125m) | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) | | |
| | [mume-math-125m](https://huggingface.co/MuseMesh/mume-math-125m) | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch `sft` is a GSM8K/MATH fine-tune | | |
| - **[mume-tokenizer-32k](https://huggingface.co/MuseMesh/mume-tokenizer-32k)**: the shared 32k unigram tokenizer. | |
| - **[mume-eval-suites](https://huggingface.co/datasets/MuseMesh/mume-eval-suites)**: the frozen test sets behind every number (FineWeb val, WikiText-103, enwik8, text8, GSM8K, MATH, OpenWebMath val) with contamination-clean variants, plus training manifests. | |
| ## Data and tokenizer | |
| - **[Sansar Sanskrit Corpus](https://huggingface.co/datasets/MuseMesh/sansar-sanskrit-corpus)**: 252M words of openly licensed Sanskrit from 21 sources, in Devanagari. It is split into one config per licence (CC BY, CC BY-SA, CC0, ODC-By, Apache-2.0, MIT), and every record carries its source, licence and attribution. Text overlapping our evaluation sets is removed. | |
| - **[sansar-sanskrit-tokenizer](https://huggingface.co/MuseMesh/sansar-sanskrit-tokenizer)**: an 8k unigram SentencePiece model over SLP1 transliteration with a Devanagari wrapper. 2.6 to 3.9 tokens per Sanskrit word with only 8k entries. | |
| ## Licences | |
| The corpus records keep their source licences. Model weights and the tokenizer are released under **CC BY-NC 4.0** for non-commercial research use, and the modelling code under **Apache-2.0**. For other uses, contact us. | |
| **Links:** [muse-mesh.com](https://muse-mesh.com) · [GitHub](https://github.com/muse-mesh) · kushal@muse-mesh.com | |