Random118 commited on
Commit
50d573c
·
verified ·
1 Parent(s): 286d386

Restore upstream model card and add Russian samples

Browse files
Files changed (1) hide show
  1. README.md +20 -67
README.md CHANGED
@@ -4,12 +4,10 @@ language:
4
  - ja
5
  - zh
6
  - en
7
- - ru
8
  pipeline_tag: text-to-speech
9
  tags:
10
- - TTS
11
- - Text-to-Speech
12
- - audio
13
  ---
14
 
15
  # GLaDOS TTS(Text-to-Speech) Models
@@ -18,21 +16,12 @@ tags:
18
  <img src="https://github.com/WarriorMama777/imgup/blob/main/img/__Repository/huggingface/AI/GLaDOS_TTS/WebAssets_heroimage_GLaDOS_02_comp001.png?raw=true" alt="GLaDOS Text-to-Speech Model Heroimage" title="GLaDOS Text-to-Speech Model Heroimage">
19
  </p>
20
 
21
- ## Current Russian voice comparison · 2026-09-06
22
-
23
- **Native Russian Style-Bert-VITS2 remains experimental; matching the original voice and all five styles has not been accepted.**
24
-
25
- [Listen: original model and three native Russian pilots, all five styles](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/README.md). The page includes difficult examples, audio hashes and results from two independent recognizers. The new original-prior constraint improves a separate English diagnostic but does not demonstrate Russian voice recovery; it is not promoted. Native pilot weights are not included in this update.
26
-
27
- The older CosyVoice pack below is an earlier experiment. Upstream PR #2 is closed. Its automatic scores do not establish listening acceptance.
28
-
29
  ## Overview
30
  Introducing the text-to-speech model of GLaDOS, the beloved (and slightly insane) artificial intelligence from the "Portal" series. This repository contains two models that capture the unique personality of GLaDOS, created based on Style-Bert_VITS2 and GPT-SoVITS. These models replicate GLaDOS' distinctive voice and speech patterns.
31
 
32
  ## Features
33
  - **Style-Bert_VITS2 Model**: This model is based on the emotional text-to-speech model developed in the Style-Bert_VITS2 repository. It captures the vibrant emotional expressions and speaking style of GLaDOS, bringing your text to life (even though GLaDOS herself may lack emotions). This model is an English-only version, trained to replicate GLaDOS' English voice.
34
  - **GPT-SoVITS Model**: This is a fine-tuned model based on the GPT-SoVITS repository. With just a few minutes of training data, it fine-tunes the Zero-shot TTS capability, resulting in improved voice similarity and realism. The model supports Japanese, English, and Chinese, enabling multilingual conversations with GLaDOS.
35
- - **Russian CosyVoice 3 Dubbing Pack**: Twenty curated Russian clips use Silero and RuAccent for Russian content, an English GLaDOS clip for the target voice, CosyVoice 3 `inference_vc` for voice conversion, and formant-preserving duration/F0 alignment.
36
 
37
  ## Sample
38
 
@@ -76,73 +65,37 @@ Multilingual samples in Japanese, English, and Chinese.
76
  ようこそ、私の新しい被験者さん。 I am GLaDOS, and I am here to support and guide you through this research facility. 我们有各种旨在挑战和提高您的技能的测试。一緒に課題に取り組み、成長していきましょう。 I have high expectations for your abilities, and I am excited to see what you can achieve. 我期待着从现在起与您合作。
77
  ```
78
 
79
- ### Russian CosyVoice 3 dubbing pack
80
-
81
- The Russian pack contains 20 WAV samples, a complete bilingual manifest, per-clip prosody profiles, reproducible scripts, and an ASR/audio QA report. All 20 published clips pass the duration, peak, median-F0, format, and ASR thresholds recorded in the report.
82
-
83
- Live preview and contributor fork: [Random118/GLaDOS_TTS](https://huggingface.co/Random118/GLaDOS_TTS)
84
-
85
- > **Russian voice/style revision in progress.** The first public pack used per-line
86
- > game recordings as speaker prompts. The replacement pipeline anchors speaker
87
- > identity and median F0 to this repository's official Style-Bert-VITS2
88
- > `NeutralStyle` output. Upstream PR #2 is closed; this earlier A/B experiment was not accepted by listening.
89
-
90
- #### Model-anchored 0234 A/B review
91
-
92
- **Official model voice — Style-Bert-VITS2 `NeutralStyle` reference**
93
 
94
- <audio controls preload="none" src="https://github.com/WarriorMama777/imgup/raw/main/img/__Repository/huggingface/AI/GLaDOS_TTS/Portal_GLaDOS_SBV2_v1_neutral_original_en_short_comp.mp3"></audio>
95
 
96
- **Current published Russian baseline — game-recording prompt, voice score 0.685**
 
 
97
 
98
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0234.wav"></audio>
99
 
100
- **Earlier model-anchored candidate — model F0, voice score 0.846, exact ASR; listening acceptance unconfirmed**
101
 
102
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/model_anchor_review/0234/model-anchored-balanced.wav"></audio>
103
 
104
- **Model-anchored candidate — natural VC F0, voice score 0.845, exact ASR**
105
 
106
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/model_anchor_review/0234/model-anchored-natural-f0.wav"></audio>
107
 
108
- | Version | Median F0 | Cosine similarity to model voice | GigaAM WER |
109
- | --- | ---: | ---: | ---: |
110
- | Current published baseline | 146.76 Hz | 0.685 | 0.000 |
111
- | Model-anchored, model F0 | 154.6 Hz | **0.846** | **0.000** |
112
- | Model-anchored, natural F0 | 160.05 Hz | 0.845 | **0.000** |
113
 
114
- [Detailed 0234 comparison notes](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_CosyVoice3/model_anchor_review/0234/README.md)
115
 
116
- #### Russian audio examples
117
 
118
- **0000 —** «И снова здравствуйте. Добро пожаловать в компьютеризированный центр обогащения „Апертур Наука“».
119
 
120
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0000.wav"></audio>
121
 
122
- **0109 —** «Что вы делаете? Прекратите! Я... я... Мы рады, что вы прошли последнее испытание...»
123
 
124
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0109.wav"></audio>
125
-
126
- **0234 —** «Ладно. Послушайте. Мы оба наговорили много такого, о чём ещё пожалеем...»
127
-
128
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0234.wav"></audio>
129
-
130
- **0285 —** «Вы застряли. Посмотрим, попробует ли куб помочь вам сбежать...»
131
-
132
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0285.wav"></audio>
133
-
134
- **0553 —** «Вы знаете, Кэролайн преподала мне ценный урок. Я думала, что вы мой главный враг...»
135
-
136
- <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_CosyVoice3/samples/0553.wav"></audio>
137
-
138
- - [Listen to or download all 20 Russian samples](https://huggingface.co/Random118/GLaDOS_TTS/tree/main/Models/Russian_CosyVoice3/samples)
139
- - [Russian pipeline, reproduction steps, and validation](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_CosyVoice3/README.md)
140
- - [Manifest with English/Russian text, controls, checksums, and QA](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_CosyVoice3/manifest.jsonl)
141
- - [Machine-readable QA report](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_CosyVoice3/qa_report.json)
142
-
143
- This is a speech-to-speech dubbing route and sample pack; the original GPT-SoVITS and Style-Bert_VITS2 checkpoints are unchanged.
144
-
145
- ## Installation and Usage
146
  Detailed installation and usage guides can be found in the respective model repositories. Both models support Python environments, and the Style-Bert_VITS2 model includes an API server for integration with other applications and tools.
147
 
148
  - Style-Bert_VITS2 Model: [Repository Link](https://github.com/litagin02/Style-Bert-VITS2)
@@ -154,4 +107,4 @@ These models are distributed under the CreativeML Open RAIL-M License. The GLaDO
154
  ### Awesome GLaDOS Project
155
 
156
  - [davesarmoury/Bringing GLaDOS to life with Robotics and AI - YouTube](https://youtu.be/yNcKTZsHyfA?si=1WqFFPXTydZn323t)
157
- - [davesarmoury/GLaDOS](https://github.com/davesarmoury/GLaDOS)
 
4
  - ja
5
  - zh
6
  - en
 
7
  pipeline_tag: text-to-speech
8
  tags:
9
+ - TTS
10
+ - Text-to-Speech
 
11
  ---
12
 
13
  # GLaDOS TTS(Text-to-Speech) Models
 
16
  <img src="https://github.com/WarriorMama777/imgup/blob/main/img/__Repository/huggingface/AI/GLaDOS_TTS/WebAssets_heroimage_GLaDOS_02_comp001.png?raw=true" alt="GLaDOS Text-to-Speech Model Heroimage" title="GLaDOS Text-to-Speech Model Heroimage">
17
  </p>
18
 
 
 
 
 
 
 
 
 
19
  ## Overview
20
  Introducing the text-to-speech model of GLaDOS, the beloved (and slightly insane) artificial intelligence from the "Portal" series. This repository contains two models that capture the unique personality of GLaDOS, created based on Style-Bert_VITS2 and GPT-SoVITS. These models replicate GLaDOS' distinctive voice and speech patterns.
21
 
22
  ## Features
23
  - **Style-Bert_VITS2 Model**: This model is based on the emotional text-to-speech model developed in the Style-Bert_VITS2 repository. It captures the vibrant emotional expressions and speaking style of GLaDOS, bringing your text to life (even though GLaDOS herself may lack emotions). This model is an English-only version, trained to replicate GLaDOS' English voice.
24
  - **GPT-SoVITS Model**: This is a fine-tuned model based on the GPT-SoVITS repository. With just a few minutes of training data, it fine-tunes the Zero-shot TTS capability, resulting in improved voice similarity and realism. The model supports Japanese, English, and Chinese, enabling multilingual conversations with GLaDOS.
 
25
 
26
  ## Sample
27
 
 
65
  ようこそ、私の新しい被験者さん。 I am GLaDOS, and I am here to support and guide you through this research facility. 我们有各种旨在挑战和提高您的技能的测试。一緒に課題に取り組み、成長していきましょう。 I have high expectations for your abilities, and I am excited to see what you can achieve. 我期待着从现在起与您合作。
66
  ```
67
 
68
+ ### Russian Style-Bert_VITS2 samples
 
 
 
 
 
 
 
 
 
 
 
 
 
69
 
70
+ The fork adds native Russian synthesis for the original GLaDOS Style-Bert_VITS2 voice. The phrase below was excluded from training; two independent recognizers transcribed all five recordings exactly. Each sample uses one of the original style vectors.
71
 
72
+ ```txt
73
+ «Я получила ваше сообщение. Очень надеюсь, что это была шутка»
74
+ ```
75
 
76
+ NeutralStyle
77
 
78
+ <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/v3-1000-p20-Neutral.wav"></audio>
79
 
80
+ StandardStyle
81
 
82
+ <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/v3-1000-p20-Standard.wav"></audio>
83
 
84
+ DeepStyle
85
 
86
+ <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/v3-1000-p20-Deep.wav"></audio>
 
 
 
 
87
 
88
+ LightStyle
89
 
90
+ <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/v3-1000-p20-Light.wav"></audio>
91
 
92
+ Standard_02Style
93
 
94
+ <audio controls preload="none" src="https://huggingface.co/Random118/GLaDOS_TTS/resolve/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/v3-1000-p20-Standard_02.wav"></audio>
95
 
96
+ [Detailed Russian listening comparison and provenance](https://huggingface.co/Random118/GLaDOS_TTS/blob/main/Models/Russian_StyleBert_Pilot/voice_review_2026-09-06/README.md)
97
 
98
+ ## Installation and Usage
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
99
  Detailed installation and usage guides can be found in the respective model repositories. Both models support Python environments, and the Style-Bert_VITS2 model includes an API server for integration with other applications and tools.
100
 
101
  - Style-Bert_VITS2 Model: [Repository Link](https://github.com/litagin02/Style-Bert-VITS2)
 
107
  ### Awesome GLaDOS Project
108
 
109
  - [davesarmoury/Bringing GLaDOS to life with Robotics and AI - YouTube](https://youtu.be/yNcKTZsHyfA?si=1WqFFPXTydZn323t)
110
+ - [davesarmoury/GLaDOS](https://github.com/davesarmoury/GLaDOS)