YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
CosyVoice2 Hebrew Fine-tune
Private repository containing CosyVoice2 fine-tuned for Hebrew speech synthesis.
Repository Structure
βββ cosyvoice/ # CosyVoice2 library code
βββ pretrained_models/
β βββ CosyVoice2-0.5B/ # Base pretrained model
β βββ llm.pt
β βββ flow.pt
β βββ hift.pt
β βββ CosyVoice-BlankEN/model.safetensors
βββ examples/hebrew/cosyvoice2/
β βββ conf/cosyvoice2.yaml # Training config
β βββ run.sh # Full training pipeline
β βββ path.sh # Environment setup
β βββ model_hebrew/ # Hebrew model files (for inference)
β β βββ llm.pt, flow.pt, hift.pt, spk2info.pt
β β βββ *.onnx (tokenizer, speaker embedding)
β βββ exp/cosyvoice2_hebrew/llm/torch_ddp/
β βββ epoch_3_whole.pt # Best by loss (2.698, acc 0.280)
β βββ epoch_4_whole.pt # Best by accuracy (2.706, acc 0.280)
βββ third_party/Matcha-TTS/ # Required dependency
βββ tools/ # Data preparation tools
Training Summary
- Base model: CosyVoice2-0.5B
- Task: Hebrew SFT (Supervised Fine-Tuning) on LLM component
- Optimizer: Adam, lr=1e-5, constant LR with 1000 warmup steps
- Batch: dynamic batching, max_frames_in_batch=2000, accum_grad=4
- Mixed precision: bf16
- Training ran: 9 epochs (~50k steps), best around epoch 3-4
Loss Curve (key checkpoints)
| Checkpoint | Loss | Accuracy | Step |
|---|---|---|---|
| epoch_0_whole | 2.795 | 0.265 | 5,957 |
| epoch_1_whole | 2.735 | 0.274 | 11,911 |
| epoch_2_whole | 2.707 | 0.278 | 17,869 |
| epoch_3_whole | 2.698 | 0.280 | 23,823 |
| epoch_4_whole | 2.706 | 0.280 | 29,780 |
| epoch_5_whole | 2.737 | 0.279 | 35,736 |
| epoch_7_whole | 2.899 | 0.271 | 47,641 |
| epoch_8_step_50000 | 3.066 | 0.263 | 49,999 |
Loss starts increasing after epoch 4-5 (overfitting).
How to Resume Fine-Tuning on a New System
1. Clone this repo
git lfs install
git clone https://huggingface.co/dm15/cosyvoice2-hebrew
cd cosyvoice2-hebrew
2. Install dependencies
pip install -r requirements.txt
# Make sure you have: torch, torchaudio, hyperpyyaml, onnxruntime-gpu,
# conformer, diffusers, etc.
cd third_party/Matcha-TTS && pip install -e . && cd ../..
3. Prepare your training data
You need train.data.list and dev.data.list pointing to parquet files.
If you have the original Hebrew data, run stages 0-4 of run.sh:
cd examples/hebrew/cosyvoice2
# Edit run.sh: set metadata= and wav_dir= to your data paths
bash run.sh # with stage=0 stop_stage=4
If you already have parquet data from a previous run, just make sure
data/train.data.list and data/dev.data.list exist with correct paths.
4. Resume training from best checkpoint
The key is the --checkpoint flag β point it to the checkpoint you want to resume from.
The training script reads step and epoch from the checkpoint and continues from there.
Resume from epoch_4_whole.pt (recommended starting point):
cd examples/hebrew/cosyvoice2
. ./path.sh
export CUDA_VISIBLE_DEVICES="0" # adjust to your GPU
torchrun --nnodes=1 --nproc_per_node=1 \
--rdzv_id=2024 --rdzv_backend="c10d" --rdzv_endpoint="localhost:1235" \
../../../cosyvoice/bin/train.py \
--train_engine torch_ddp \
--config conf/cosyvoice2.yaml \
--train_data data/train.data.list \
--cv_data data/dev.data.list \
--qwen_pretrain_path ../../../pretrained_models/CosyVoice2-0.5B/CosyVoice-BlankEN \
--onnx_path ../../../pretrained_models/CosyVoice2-0.5B \
--model llm \
--checkpoint exp/cosyvoice2_hebrew/llm/torch_ddp/epoch_4_whole.pt \
--model_dir $(pwd)/exp/cosyvoice2_hebrew/llm/torch_ddp \
--tensorboard_dir $(pwd)/tensorboard/cosyvoice2_hebrew/llm/torch_ddp \
--ddp.dist_backend nccl \
--num_workers 4 \
--prefetch 100 \
--pin_memory \
--use_amp
5. Resume tips
- Checkpoint stores: model weights +
stepcount +epochnumber. Training resumes fromepoch + 1. - Optimizer state is NOT saved (save_states=model_only), so momentum/Adam state resets on resume. This is normal for this codebase.
- To lower the learning rate for continued training, edit
conf/cosyvoice2.yamlβtrain_conf.optim_conf.lr - To reduce overfitting (loss was rising after epoch 4): try reducing
max_epoch, increasing data, or lowering lr. - Multi-GPU: change
--nproc_per_node=NandCUDA_VISIBLE_DEVICESaccordingly.
6. Inference with fine-tuned model
cd examples/hebrew/cosyvoice2
python3 infer_cpu.py # uses model_hebrew/ directory
Or to use a specific checkpoint as the LLM, copy it to model_hebrew/llm.pt:
cp exp/cosyvoice2_hebrew/llm/torch_ddp/epoch_3_whole.pt model_hebrew/llm.pt
Original Training Command (for reference)
The full pipeline is in examples/hebrew/cosyvoice2/run.sh. It includes:
- Stage 0: Data preparation from metadata.csv
- Stage 1: Speaker embedding extraction (CampPlus)
- Stage 2: Speech token extraction
- Stage 3: Parquet file creation
- Stage 4: Data list creation
- Stage 5: Fine-tuning (LLM + Flow)
- Stage 6: Checkpoint averaging