You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Access is granted manually by the repository owner.
Log in or Sign Up to review the conditions and access this model content.
VoxCPM2 Persian, round 4
VoxCPM2 fine-tuned for Persian with style captions, continued from
best_checkpoint_of_post_training.
| folder | step | val loss | |
|---|---|---|---|
best-ckp/ |
4012 (final) | 0.8738 | recommended; used for the demo samples |
step-4000/ |
4000 | 0.8742 | |
step-3000/ |
3000 | 0.8749 |
Each folder is a complete inference checkpoint (weights, AudioVAE, config, tokenizer). Optimizer state is not included.
Usage
Prefix the text with one caption, no space, then the text with optional inline tags:
(neutral)متن ... (warm)... (formal)... (agent)سلام، وقتتون بخیر... [uhm] ...
Captions trained: neutral, warm, happy, excited, sad, fear, anger,
formal (announcements, official speech), agent (call-center agent: informal
but respectful, emotionally neutral).
Tags: [breath] [short pause] [uhm] [clears throat] [sighs] [chuckles]
[yawns] [slow] [fast] [whispering].
Plain or diacritized (harakat) Persian both work; 20% of the training text was diacritized.
from voxcpm import VoxCPM
m = VoxCPM(voxcpm_model_path="best-ckp", zipenhancer_model_path=None, enable_denoiser=False)
wav = m.generate(text="(agent)سلام، وقتتون بخیر، کریمی هستم. [uhm] بله، سفارشتون رو دیدم.",
reference_wav_path="speaker.wav", cfg_value=2.0, inference_timesteps=10)
reference_wav_path clones a voice (6-14 s works best). The style comes from the
caption, not from the reference: every training reference was a neutral clip.
Training
- Init:
best_checkpoint_of_post_training - Data (251.5 h): 218.4 h Gemini 3.8 Flash TTS Persian long-form (9 captions,
7 voices, one caption per clip, 20% diacritized, same-voice neutral references)
- 33.1 h NonVerbalSpeech-38K (English + Chinese) replay against forgetting
- 4 epochs, 4,012 steps, LR 1e-5, ~1,200 s of audio per optimizer step (bucketed batches of 400 s x grad-accum 3), single B300
- Config, launcher and full log in
training/
Validation loss
| step | val loss |
|---|---|
| 0 | 0.9032 |
| 250 | 0.8865 |
| 500 | 0.8823 |
| 750 | 0.8817 |
| 1000 | 0.8784 |
| 1250 | 0.8789 |
| 1500 | 0.8795 |
| 1750 | 0.8757 |
| 2000 | 0.8761 |
| 2250 | 0.8753 |
| 2500 | 0.8737 |
| 2750 | 0.8721 |
| 3000 | 0.8749 |
| 3250 | 0.8722 |
| 3500 | 0.8767 |
| 3750 | 0.8735 |
| 4000 | 0.8742 |
| 4011 | 0.8738 |