Nepali-Swar-Experimental
A 16.1 M parameter semantic language model for Nepali and English speech. Swar (स्वर) is Nepali for voice, sound, or musical note.
Experimental. A research checkpoint, not a production system.
What this is, and what it is not
This is one component of a TTS system, not a complete one. It converts text into CosyVoice3 speech tokens. It does not produce audio by itself.
text -> [ Nepali-Swar, 16.1 M ] -> speech tokens -> [ CosyVoice3 flow + HiFT ] -> audio
this repo frozen, not ours
The decoder is the frozen flow-matching model and HiFT vocoder from
Fun-CosyVoice3-0.5B. This model replaces only the 0.5 B Qwen LM in that stack —
which is the point: 16.1 M parameters in place of 500 M, a 31× reduction in the
component that carries the language.
Results
| dev loss | dev accuracy | |
|---|---|---|
| chance (6,569 classes) | 8.790 | 0.015% |
| Nepali-Swar (16.1 M) | 4.1980 | 18.23% |
| on call-centre register | 2.9938 | 31.53% |
| Fun-CosyVoice3-0.5B reference | 3.765 | 17.6% |
Accuracy is roughly 1,200× chance. That is less impressive than it sounds: next-speech-token prediction has many valid answers for the same text — small differences in pitch or timing are different token ids but equally correct speech — so ~18% sits near the intrinsic ceiling rather than measuring quality. Loss is the number with real headroom.
The shipped checkpoint is tuned for call-centre register, which is why its score on that slice (2.9938 / 31.53%) is far stronger than its general score.
Files
| file | |
|---|---|
swar.pt |
the model, 64 MB. Carries lm_config, char_vocab and control_buckets inside it. |
char_vocab.json |
154-character vocabulary — part of the model |
control_buckets.json |
rate/pause bucket edges — part of the model |
sample.wav |
5.9 s sample decoded through CosyVoice3's vocoder |
⚠️ char_vocab.json and control_buckets.json are not metadata. Character ids are
positions in that file, and control ids are defined by those bucket edges. Running
these weights against different ones is silent corruption — no error, just wrong
sounds. They are embedded in the .pt for that reason.
Usage
import torch
from nvoice.lm import build
from nvoice.vocab import CharVocab, SOS, TASK, ctrl_prefix
ck = torch.load("swar.pt", map_location="cpu")
model = build("M16", **{k: v for k, v in ck["lm_config"].items() if k != "head_dim"})
model.load_state_dict(ck["model"]); model.to("cuda").eval()
vocab = CharVocab(ck["char_vocab"]["chars"]) # travels inside the checkpoint
ids = vocab.encode("नमस्कार, हजुरलाई कसरी सहयोग गर्न सक्छु?")
prefix = [SOS] + ctrl_prefix("callcenter", None, None, None) + ids + [TASK]
tokens = model.generate(prefix, max_new=760, min_new=12, device="cuda")
# tokens -> audio via CosyVoice3's flow + HiFT
Control tokens
Four optional dials precede the text: style (neutral / callcenter), rate,
pause, pitch — five buckets each, or masked. pitch is reserved and
non-functional. Rate control is weak.
Limitations
- Not a standalone TTS. Requires CosyVoice3's flow + vocoder (~1.4 GB).
- Experimental quality. No MOS, no native-listener evaluation, no CER gate.
- Timbre comes from the reference clip, not the model — this is a voice-cloning stack, and the model carries language, not identity.
- Long paragraphs are best synthesised one sentence at a time; an autoregressive model of this size can drift or stall on very long inputs.
License
Apache-2.0. The decoder it depends on (Fun-CosyVoice3-0.5B) carries its own terms.