File size: 2,463 Bytes
8c2913a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
# Russian CosyVoice 3 dubbing pack

This directory adds a reproducible Russian voice-conversion route and 20 curated
Russian GLaDOS samples. It does not add a native Russian checkpoint to the
original GPT-SoVITS or Style-Bert-VITS2 models.

## Pipeline

1. The English Portal clip supplies the GLaDOS voice, timbre, formants, and delivery.
2. RuAccent and Silero v4 generate a clean Russian content track.
3. Rubber Band matches the reference duration before VC, or after VC for lines listed in `pipeline_controls.json`.
4. CosyVoice 3 `inference_vc` transfers the English reference voice onto the Russian content.
5. A verified, formant-preserving Rubber Band pass aligns median F0 to the English reference.

`manifest.jsonl` maps every English reference, Russian translation, output sample,
prosody profile, per-line control, and QA result. The references come from
[`ray0rf1re/GLaDOS-audio-v2`](https://huggingface.co/datasets/ray0rf1re/GLaDOS-audio-v2).

## Reproduce

Clone [FunAudioLLM/CosyVoice](https://github.com/FunAudioLLM/CosyVoice) with its
submodules and install its dependencies. Install the packages in
`requirements.txt` and the `rubberband` command-line tool. Download
[`FunAudioLLM/Fun-CosyVoice3-0.5B-2512`](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512),
then run:

```bash
python download_refs.py
python generate_silero_sources.py --ruaccent-workdir .cache/ruaccent
python generate_cosyvoice_vc_batch.py \
  --cosyvoice-root /path/to/CosyVoice \
  --model-dir /path/to/Fun-CosyVoice3-0.5B-2512
```

Use `--overwrite` to rebuild existing intermediates. The default directories are
relative to this folder, and `--ids 0234 0258` can limit a run.

## Validation

All 20 published WAV files pass these checks:

- 24 kHz, mono, PCM16;
- absolute duration difference from the English reference at most 0.08 seconds;
- peak amplitude at most 0.951;
- median-F0 relative error at most 2.5%;
- GigaAM multilingual CTC word error rate at most 10% per clip.

The checked run produced 17 word-exact clips, 1.36% mean WER, and 9.09% maximum
WER. Re-run it with:

```bash
python qa_cosyvoice_batch.py --asr --outputs samples
```

See `qa_report.json` for per-clip measurements and transcripts. The sample pack
was generated with [Silero](https://github.com/snakers4/silero-models),
[RuAccent](https://github.com/Den4ikAI/ruaccent), CosyVoice 3, and Rubber Band.
The original repository license and the source/model terms continue to apply.