VoxCPM2 Persian β round 3
Continued from markmuller/TTS_POST_trainnig_1_emo / best_checkpoint_of_post_training,
2 epochs (2,917 steps) on 98,061 rows / 270.5 h:
| slice | rows | hours | text |
|---|---|---|---|
| gold (TTS_DATA_GOLD) | 77,025 | 141.4 | (emotion)[tag] captions, plain text |
| fidibo50 | 8,361 | 49.4 | (fidibo) + diacritized text, same-speaker refs |
| neutral50 | 2,236 | 46.6 | Gemini long-form (43β150 s), inline per-sentence captions, same-voice refs |
| nonverbal38k | 10,439 | 33.1 | EN + ZH replay with nonverbal tags, zero-shot (CC-BY-NC-4.0 source) |
Length-bucketed batches (400 s padded audio per micro-batch, x3 accumulation), LR 1e-5, one B300.
Checkpoints
| folder | step | why |
|---|---|---|
step_0002917 |
2917 (end of epoch 2) | best flow-matching loss on every slice |
step_0001000 |
1000 | best stop-head loss β try it if endings are cut short or run on |
Each folder is a full trainer checkpoint (weights + optimizer + scheduler), so it can be loaded for inference or resumed.
Per-slice validation (full val sets, identical batches and noise per checkpoint)
diffusion loss / stop loss, lower is better:
| ckpt | gold | neutral50 | fidibo50 | nonverbal38k |
|---|---|---|---|---|
| init | 0.8983 / 0.0137 | 0.9345 / 0.0023 | 0.8031 / 0.0038 | 1.0056 / 0.0255 |
| step 1000 | 0.8902 / 0.0125 | 0.9271 / 0.0017 | 0.8048 / 0.0028 | 0.9916 / 0.0209 |
| step 2917 | 0.8837 / 0.0137 | 0.9235 / 0.0027 | 0.8035 / 0.0025 | 0.9901 / 0.0242 |
run/ holds the generated training config, the training log and the full evaluation JSON.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support