You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access is granted manually by the repository owner.

Log in or Sign Up to review the conditions and access this model content.

VoxCPM2 Persian, round 4

VoxCPM2 fine-tuned for Persian with style captions, continued from best_checkpoint_of_post_training.

folder step val loss
best-ckp/ 4012 (final) 0.8738 recommended; used for the demo samples
step-4000/ 4000 0.8742
step-3000/ 3000 0.8749

Each folder is a complete inference checkpoint (weights, AudioVAE, config, tokenizer). Optimizer state is not included.

Usage

Prefix the text with one caption, no space, then the text with optional inline tags:

(neutral)متن ...      (warm)...      (formal)...      (agent)سلام، وقتتون بخیر... [uhm] ...

Captions trained: neutral, warm, happy, excited, sad, fear, anger, formal (announcements, official speech), agent (call-center agent: informal but respectful, emotionally neutral).

Tags: [breath] [short pause] [uhm] [clears throat] [sighs] [chuckles] [yawns] [slow] [fast] [whispering].

Plain or diacritized (harakat) Persian both work; 20% of the training text was diacritized.

from voxcpm import VoxCPM
m = VoxCPM(voxcpm_model_path="best-ckp", zipenhancer_model_path=None, enable_denoiser=False)
wav = m.generate(text="(agent)سلام، وقتتون بخیر، کریمی هستم. [uhm] بله، سفارشتون رو دیدم.",
                 reference_wav_path="speaker.wav", cfg_value=2.0, inference_timesteps=10)

reference_wav_path clones a voice (6-14 s works best). The style comes from the caption, not from the reference: every training reference was a neutral clip.

Training

  • Init: best_checkpoint_of_post_training
  • Data (251.5 h): 218.4 h Gemini 3.8 Flash TTS Persian long-form (9 captions, 7 voices, one caption per clip, 20% diacritized, same-voice neutral references)
    • 33.1 h NonVerbalSpeech-38K (English + Chinese) replay against forgetting
  • 4 epochs, 4,012 steps, LR 1e-5, ~1,200 s of audio per optimizer step (bucketed batches of 400 s x grad-accum 3), single B300
  • Config, launcher and full log in training/

Validation loss

step val loss
0 0.9032
250 0.8865
500 0.8823
750 0.8817
1000 0.8784
1250 0.8789
1500 0.8795
1750 0.8757
2000 0.8761
2250 0.8753
2500 0.8737
2750 0.8721
3000 0.8749
3250 0.8722
3500 0.8767
3750 0.8735
4000 0.8742
4011 0.8738
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support