Download Models/Russian_CosyVoice3/README.md from Random118/GLaDOS_TTS: direct link, hf CLI and curl.
- Browser
- Download file 2.46 kB
-
https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/README.md
- Command line
-
hf download hf://Random118/GLaDOS_TTS/Models/Russian_CosyVoice3/README.md
-
curl -L -o README.md https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/README.md
Russian CosyVoice 3 dubbing pack
This directory adds a reproducible Russian voice-conversion route and 20 curated Russian GLaDOS samples. It does not add a native Russian checkpoint to the original GPT-SoVITS or Style-Bert-VITS2 models.
Pipeline
- The English Portal clip supplies the GLaDOS voice, timbre, formants, and delivery.
- RuAccent and Silero v4 generate a clean Russian content track.
- Rubber Band matches the reference duration before VC, or after VC for lines listed in
pipeline_controls.json. - CosyVoice 3
inference_vctransfers the English reference voice onto the Russian content. - A verified, formant-preserving Rubber Band pass aligns median F0 to the English reference.
manifest.jsonl maps every English reference, Russian translation, output sample,
prosody profile, per-line control, and QA result. The references come from
ray0rf1re/GLaDOS-audio-v2.
Reproduce
Clone FunAudioLLM/CosyVoice with its
submodules and install its dependencies. Install the packages in
requirements.txt and the rubberband command-line tool. Download
FunAudioLLM/Fun-CosyVoice3-0.5B-2512,
then run:
python download_refs.py
python generate_silero_sources.py --ruaccent-workdir .cache/ruaccent
python generate_cosyvoice_vc_batch.py \
--cosyvoice-root /path/to/CosyVoice \
--model-dir /path/to/Fun-CosyVoice3-0.5B-2512
Use --overwrite to rebuild existing intermediates. The default directories are
relative to this folder, and --ids 0234 0258 can limit a run.
Validation
All 20 published WAV files pass these checks:
- 24 kHz, mono, PCM16;
- absolute duration difference from the English reference at most 0.08 seconds;
- peak amplitude at most 0.951;
- median-F0 relative error at most 2.5%;
- GigaAM multilingual CTC word error rate at most 10% per clip.
The checked run produced 17 word-exact clips, 1.36% mean WER, and 9.09% maximum WER. Re-run it with:
python qa_cosyvoice_batch.py --asr --outputs samples
See qa_report.json for per-clip measurements and transcripts. The sample pack
was generated with Silero,
RuAccent, CosyVoice 3, and Rubber Band.
The original repository license and the source/model terms continue to apply.