Nepali-Swar-Experimental

A 16.1 M parameter semantic language model for Nepali and English speech. Swar (स्वर) is Nepali for voice, sound, or musical note.

Experimental. A research checkpoint, not a production system.

What this is, and what it is not

This is one component of a TTS system, not a complete one. It converts text into CosyVoice3 speech tokens. It does not produce audio by itself.

text  ->  [ Nepali-Swar, 16.1 M ]  ->  speech tokens  ->  [ CosyVoice3 flow + HiFT ]  ->  audio
                  this repo                                       frozen, not ours

The decoder is the frozen flow-matching model and HiFT vocoder from Fun-CosyVoice3-0.5B. This model replaces only the 0.5 B Qwen LM in that stack — which is the point: 16.1 M parameters in place of 500 M, a 31× reduction in the component that carries the language.

Results

dev loss dev accuracy
chance (6,569 classes) 8.790 0.015%
Nepali-Swar (16.1 M) 4.1980 18.23%
on call-centre register 2.9938 31.53%
Fun-CosyVoice3-0.5B reference 3.765 17.6%

Accuracy is roughly 1,200× chance. That is less impressive than it sounds: next-speech-token prediction has many valid answers for the same text — small differences in pitch or timing are different token ids but equally correct speech — so ~18% sits near the intrinsic ceiling rather than measuring quality. Loss is the number with real headroom.

The shipped checkpoint is tuned for call-centre register, which is why its score on that slice (2.9938 / 31.53%) is far stronger than its general score.

Files

file
swar.pt the model, 64 MB. Carries lm_config, char_vocab and control_buckets inside it.
char_vocab.json 154-character vocabulary — part of the model
control_buckets.json rate/pause bucket edges — part of the model
sample.wav 5.9 s sample decoded through CosyVoice3's vocoder

⚠️ char_vocab.json and control_buckets.json are not metadata. Character ids are positions in that file, and control ids are defined by those bucket edges. Running these weights against different ones is silent corruption — no error, just wrong sounds. They are embedded in the .pt for that reason.

Usage

import torch
from nvoice.lm import build
from nvoice.vocab import CharVocab, SOS, TASK, ctrl_prefix

ck = torch.load("swar.pt", map_location="cpu")
model = build("M16", **{k: v for k, v in ck["lm_config"].items() if k != "head_dim"})
model.load_state_dict(ck["model"]); model.to("cuda").eval()
vocab = CharVocab(ck["char_vocab"]["chars"])          # travels inside the checkpoint

ids = vocab.encode("नमस्कार, हजुरलाई कसरी सहयोग गर्न सक्छु?")
prefix = [SOS] + ctrl_prefix("callcenter", None, None, None) + ids + [TASK]
tokens = model.generate(prefix, max_new=760, min_new=12, device="cuda")
# tokens -> audio via CosyVoice3's flow + HiFT

Control tokens

Four optional dials precede the text: style (neutral / callcenter), rate, pause, pitch — five buckets each, or masked. pitch is reserved and non-functional. Rate control is weak.

Limitations

  • Not a standalone TTS. Requires CosyVoice3's flow + vocoder (~1.4 GB).
  • Experimental quality. No MOS, no native-listener evaluation, no CER gate.
  • Timbre comes from the reference clip, not the model — this is a voice-cloning stack, and the model carries language, not identity.
  • Long paragraphs are best synthesised one sentence at a time; an autoregressive model of this size can drift or stall on very long inputs.

License

Apache-2.0. The decoder it depends on (Fun-CosyVoice3-0.5B) carries its own terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support