File size: 2,463 Bytes
8c2913a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 | # Russian CosyVoice 3 dubbing pack
This directory adds a reproducible Russian voice-conversion route and 20 curated
Russian GLaDOS samples. It does not add a native Russian checkpoint to the
original GPT-SoVITS or Style-Bert-VITS2 models.
## Pipeline
1. The English Portal clip supplies the GLaDOS voice, timbre, formants, and delivery.
2. RuAccent and Silero v4 generate a clean Russian content track.
3. Rubber Band matches the reference duration before VC, or after VC for lines listed in `pipeline_controls.json`.
4. CosyVoice 3 `inference_vc` transfers the English reference voice onto the Russian content.
5. A verified, formant-preserving Rubber Band pass aligns median F0 to the English reference.
`manifest.jsonl` maps every English reference, Russian translation, output sample,
prosody profile, per-line control, and QA result. The references come from
[`ray0rf1re/GLaDOS-audio-v2`](https://huggingface.co/datasets/ray0rf1re/GLaDOS-audio-v2).
## Reproduce
Clone [FunAudioLLM/CosyVoice](https://github.com/FunAudioLLM/CosyVoice) with its
submodules and install its dependencies. Install the packages in
`requirements.txt` and the `rubberband` command-line tool. Download
[`FunAudioLLM/Fun-CosyVoice3-0.5B-2512`](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512),
then run:
```bash
python download_refs.py
python generate_silero_sources.py --ruaccent-workdir .cache/ruaccent
python generate_cosyvoice_vc_batch.py \
--cosyvoice-root /path/to/CosyVoice \
--model-dir /path/to/Fun-CosyVoice3-0.5B-2512
```
Use `--overwrite` to rebuild existing intermediates. The default directories are
relative to this folder, and `--ids 0234 0258` can limit a run.
## Validation
All 20 published WAV files pass these checks:
- 24 kHz, mono, PCM16;
- absolute duration difference from the English reference at most 0.08 seconds;
- peak amplitude at most 0.951;
- median-F0 relative error at most 2.5%;
- GigaAM multilingual CTC word error rate at most 10% per clip.
The checked run produced 17 word-exact clips, 1.36% mean WER, and 9.09% maximum
WER. Re-run it with:
```bash
python qa_cosyvoice_batch.py --asr --outputs samples
```
See `qa_report.json` for per-clip measurements and transcripts. The sample pack
was generated with [Silero](https://github.com/snakers4/silero-models),
[RuAccent](https://github.com/Den4ikAI/ruaccent), CosyVoice 3, and Rubber Band.
The original repository license and the source/model terms continue to apply.
|