0xra commited on
Commit
1eeaa77
Β·
verified Β·
1 Parent(s): 9c25162

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +326 -0
README.md ADDED
@@ -0,0 +1,326 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ pipeline_tag: audio-to-audio
3
+ tags:
4
+ - rvc
5
+ - rvc-v2
6
+ - applio
7
+ - voice-conversion
8
+ - retrieval-based-voice-conversion
9
+ - contentvec
10
+ - rmvpe
11
+ - hifigan
12
+ - 48khz
13
+ - cyberpunk-2077
14
+ - rogue
15
+ ---
16
+
17
+ # Rogue β€” RVC v2 Voice Conversion Model for Applio
18
+
19
+ RVC v2 voice-conversion model for **Rogue**, trained with **Applio** from extracted in-game voice audio.
20
+
21
+ > **Important:** this is a **voice conversion** model, not a standalone text-to-speech model. It converts an existing speech or singing performance into the target timbre while retaining much of the source timing, phrasing, and pitch contour.
22
+
23
+ ## Model files
24
+
25
+ The supplied Applio/RVC package contains:
26
+
27
+ | File | Purpose | Size |
28
+ |---|---|---:|
29
+ | `rogue.pth` | RVC model weights | 57,532,660 bytes / 54.87 MiB |
30
+ | `rogue.index` | Retrieval index used to reinforce the target timbre | 333,505,619 bytes / 318.06 MiB |
31
+
32
+ The uploaded package `rogue.zip` contains both files.
33
+
34
+ ### SHA-256
35
+
36
+ ```text
37
+ rogue.pth
38
+ b13ed63734886b2894cd667eafa26e4e966c6aea51e4422b6ee7457a71cdc210
39
+
40
+ rogue.index
41
+ dbe27e2d354f4e690b1459a2872a7f65a4d9c9f9efadc822d361abd06264bb9e
42
+
43
+ rogue.zip
44
+ d556db0d73b0e56ee48012b8e180d2e4c19a5b3b841ae7561c409b6a26435a21
45
+ ```
46
+
47
+ ## Model details
48
+
49
+ | Property | Value |
50
+ |---|---|
51
+ | Model type | Retrieval-Based Voice Conversion (RVC) |
52
+ | RVC version | **v2** |
53
+ | Training / inference app | **Applio** |
54
+ | Packaged model filename | `rogue.pth` |
55
+ | Packaged index filename | `rogue.index` |
56
+ | Internal training model name | `mara` |
57
+ | Target / donor voice | `rogue` |
58
+ | Sample rate | **48,000 Hz** |
59
+ | F0 / pitch guidance | **Enabled** |
60
+ | Pitch extractor used for feature extraction | **RMVPE** |
61
+ | Embedder | **ContentVec** |
62
+ | Vocoder | **HiFi-GAN** |
63
+ | Number of speakers | **1** |
64
+ | Published checkpoint epoch | **200** |
65
+ | Published checkpoint step | **26,200** |
66
+ | Embedded dataset length | `00:35:00` |
67
+ | Model creation timestamp | `2026-09-09T10:45:38.675607` |
68
+ | Embedded model hash | `802d9890b58d28f51309e9db116ff921280d948615abfcbb0e9069c2fb530720` |
69
+
70
+ ### Naming note
71
+
72
+ The original Applio training sheet uses **`mara`** as the training model name and **`rogue`** as the donor/target voice. The distributed files were packaged as **`rogue.pth`** and **`rogue.index`**. The internal metadata inside `rogue.pth` still reports `model_name: mara`.
73
+
74
+ ## Training dataset
75
+
76
+ The supplied dataset archive was verified as:
77
+
78
+ | Property | Value |
79
+ |---|---:|
80
+ | WAV files | **626** |
81
+ | Total duration | **35:00.422** |
82
+ | Sample rate | **48,000 Hz** for all 626 files |
83
+ | Channels | **Mono** for all 626 files |
84
+ | PCM sample width | **16-bit** for all 626 files |
85
+ | Shortest clip | **1.210 s** |
86
+ | Median clip | **2.900 s** |
87
+ | Mean clip | **3.355 s** |
88
+ | Longest clip | **11.732 s** |
89
+
90
+ The archive stores the clips under:
91
+
92
+ ```text
93
+ mara/
94
+ rogue_*.wav
95
+ ```
96
+
97
+ The provided dataset contains audio files rather than a transcription dataset.
98
+
99
+ ### Dataset preprocessing rationale
100
+
101
+ The source audio was extracted from a game rather than recorded from a microphone for this training run. The training sheet therefore keeps both **Process Effects** and **Noise Reduction** disabled. The source was already produced/mastered audio, and an additional cleanup pass could remove breath, sibilance, and other speaker-specific details useful to voice conversion.
102
+
103
+ ## Applio training configuration
104
+
105
+ ### 1. Preprocess
106
+
107
+ | Setting | Value |
108
+ |---|---|
109
+ | Model Name | `mara` |
110
+ | Sample Rate | `48000` |
111
+ | Cut Preprocess | `Automatic` |
112
+ | Chunk Length | `3.0` s |
113
+ | Overlap Length | `0.3` s |
114
+ | Process Effects | **OFF** |
115
+ | Noise Reduction | **OFF** |
116
+ | CPU Cores | Default |
117
+
118
+ ### 2. Extract
119
+
120
+ | Setting | Value |
121
+ |---|---|
122
+ | Model Name | `mara` |
123
+ | Sample Rate | `48000` |
124
+ | Pitch Extractor / F0 | **RMVPE** |
125
+ | Embedder Model | **ContentVec** |
126
+ | GPU | `0` |
127
+
128
+ ### 3. Train
129
+
130
+ | Setting | Value |
131
+ |---|---|
132
+ | Model Name | `mara` |
133
+ | Sample Rate | `48000` |
134
+ | Vocoder | **HiFi-GAN** |
135
+ | Batch Size | **8** |
136
+ | Configured Total Epochs | **300** |
137
+ | Save Every Epoch | **10** |
138
+ | Save Only Latest | **OFF** |
139
+ | Save Every Weights | **ON** |
140
+ | Pretrained | **ON** |
141
+ | Cache Dataset in GPU | **OFF** |
142
+ | GPU | `0` |
143
+
144
+ ### Published checkpoint vs. configured training length
145
+
146
+ The training configuration was set to **300 total epochs** and to save weights every 10 epochs. However, the `rogue.pth` file published in this package identifies itself as **epoch 200, step 26,200**.
147
+
148
+ For that reason, this model card describes the distributed model as the **epoch-200 checkpoint** rather than calling it a 300-epoch model.
149
+
150
+ ## Using the model in Applio
151
+
152
+ Applio expects an RVC model to use the `.pth` weights file and, when available, the matching `.index` retrieval file.
153
+
154
+ ### Option 1 β€” Applio model downloader
155
+
156
+ If `rogue.zip` is uploaded to this Hugging Face repository, its direct file URL can be used in Applio's **Download Model** panel.
157
+
158
+ The URL format is:
159
+
160
+ ```text
161
+ https://huggingface.co/<USERNAME>/<REPOSITORY>/resolve/main/rogue.zip
162
+ ```
163
+
164
+ Replace `<USERNAME>` and `<REPOSITORY>` with the actual Hugging Face repository path.
165
+
166
+ ### Option 2 β€” Manual installation
167
+
168
+ Extract `rogue.zip` and place both files in one model folder, for example:
169
+
170
+ ```text
171
+ Applio/
172
+ └── logs/
173
+ └── Rogue/
174
+ β”œβ”€β”€ rogue.pth
175
+ └── rogue.index
176
+ ```
177
+
178
+ Then refresh the model list in Applio and select the matching `.pth` and `.index`.
179
+
180
+ Official Applio model-installation documentation:
181
+
182
+ https://docs.applio.org/getting-started/installing-inference-models/
183
+
184
+ ## Example Applio CLI inference
185
+
186
+ Current Applio versions expose inference through `core.py`.
187
+
188
+ Example starting point:
189
+
190
+ ```bash
191
+ python core.py infer \
192
+ --input_path input.wav \
193
+ --output_path rogue_output.wav \
194
+ --pth_path logs/Rogue/rogue.pth \
195
+ --index_path logs/Rogue/rogue.index \
196
+ --pitch 0 \
197
+ --index_rate 0.75 \
198
+ --volume_envelope 1 \
199
+ --protect 0.5 \
200
+ --f0_method rmvpe \
201
+ --embedder_model contentvec
202
+ ```
203
+
204
+ The inference values above are **starting values, not training parameters**. Adjust them for the source recording.
205
+
206
+ Useful notes:
207
+
208
+ - Keep the embedder as **ContentVec**, matching training.
209
+ - **RMVPE** is a sensible default for this model because it was also used during feature extraction.
210
+ - `pitch 0` preserves the source pitch by default.
211
+ - The retrieval index can improve target-timbre similarity, but an excessively high index influence may also introduce artifacts.
212
+ - Source audio quality, pitch range, delivery, noise, and pronunciation can materially affect the result.
213
+
214
+ Official Applio CLI documentation:
215
+
216
+ https://docs.applio.org/reference/cli/
217
+
218
+ ## Reproducing the training setup
219
+
220
+ The supplied training sheet also records a headless helper command:
221
+
222
+ ```bash
223
+ tools/rvc_train.sh mara
224
+ ```
225
+
226
+ The recorded expected Applio outputs were:
227
+
228
+ ```text
229
+ /mnt/sata/Applio/logs/mara/mara.pth
230
+ /mnt/sata/Applio/logs/mara/*.index
231
+ ```
232
+
233
+ These are local paths from the original training environment and are not required for inference after the files have been packaged.
234
+
235
+ ## Evaluation
236
+
237
+ No formal objective evaluation results, AB listening test, benchmark scores, or reference inference samples were supplied with the uploaded files.
238
+
239
+ The package therefore does **not** claim that epoch 200 is objectively superior to another saved checkpoint. Users should evaluate the model by listening on representative speech/singing inputs and tuning inference parameters for their use case.
240
+
241
+ ## Limitations
242
+
243
+ - Trained from approximately **35 minutes** of a single target voice.
244
+ - Training material is game dialogue, so the model may inherit characteristics of that recording, performance style, mastering, or available pitch range.
245
+ - Voice conversion quality depends strongly on the input speaker and recording.
246
+ - Pronunciation, emotion, singing range, very high/low pitch, background noise, and aggressive processing can reduce similarity or create artifacts.
247
+ - No multilingual or cross-language evaluation was provided.
248
+ - No formal quality benchmark was provided.
249
+ - This model performs voice conversion; it does not generate speech from text on its own.
250
+
251
+ ## Responsible use
252
+
253
+ Use the model transparently and responsibly.
254
+
255
+ - Do not present generated audio as an authentic recording of the original performer or as an official game asset.
256
+ - Do not use generated audio to mislead, defraud, harass, or impersonate someone deceptively.
257
+ - Clearly label synthetic or voice-converted audio when context could otherwise make its origin unclear.
258
+ - Respect applicable rights, platform rules, and local law.
259
+
260
+ ## Rights and redistribution
261
+
262
+ The training material was extracted from game audio. This model card does **not** grant rights to the original game recordings, characters, performances, trademarks, or other underlying assets.
263
+
264
+ No repository license is declared in the YAML metadata above because the supplied files did not include a license specifying what redistribution or reuse rights can be granted.
265
+
266
+ If the raw dataset is going to be made public, verify that you have the right to redistribute the source audio. If you do not, a safer public repository layout is to publish the trained model package and documentation without publishing the extracted source WAV files.
267
+
268
+ This is an unofficial community model and is not presented as affiliated with or endorsed by the game's developers, publishers, performers, or other rights holders.
269
+
270
+ ## Integrity information
271
+
272
+ Additional uploaded-source checksums:
273
+
274
+ ```text
275
+ rogue_datasets.zip
276
+ e5531e041a4d3e33baca6879509c58e03d6b7524239c918d800126eeba321fff
277
+
278
+ SETTINGS.txt
279
+ 5e6ba830f607c276c3cd1e1acbf0184140e0916c6f8492959df26c282db60fae
280
+ ```
281
+
282
+ ## Recommended Hugging Face repository layout
283
+
284
+ For a compact repository that supports Applio's ZIP downloader:
285
+
286
+ ```text
287
+ README.md
288
+ rogue.zip
289
+ SETTINGS.txt
290
+ ```
291
+
292
+ If you also want people to access the two RVC files directly, you can additionally upload:
293
+
294
+ ```text
295
+ rogue.pth
296
+ rogue.index
297
+ ```
298
+
299
+ Be aware that keeping both the ZIP and the extracted model files duplicates the model storage.
300
+
301
+ ## Frameworks and references
302
+
303
+ - Applio: https://github.com/IAHispano/Applio
304
+ - Applio documentation: https://docs.applio.org/
305
+ - Hugging Face model cards: https://huggingface.co/docs/hub/model-cards
306
+
307
+ ---
308
+
309
+ ### Quick technical summary
310
+
311
+ ```text
312
+ RVC version: v2
313
+ Sample rate: 48000 Hz
314
+ Pitch guidance: yes
315
+ Pitch extractor: RMVPE
316
+ Embedder: ContentVec
317
+ Vocoder: HiFi-GAN
318
+ Dataset: 626 mono 16-bit WAV files
319
+ Dataset duration: 35:00.422
320
+ Batch size: 8
321
+ Configured epochs: 300
322
+ Published epoch: 200
323
+ Published step: 26200
324
+ Weights: rogue.pth
325
+ Index: rogue.index
326
+ ```