File size: 10,574 Bytes
1eeaa77 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 | ---
pipeline_tag: audio-to-audio
tags:
- rvc
- rvc-v2
- applio
- voice-conversion
- retrieval-based-voice-conversion
- contentvec
- rmvpe
- hifigan
- 48khz
- cyberpunk-2077
- rogue
---
# Rogue β RVC v2 Voice Conversion Model for Applio
RVC v2 voice-conversion model for **Rogue**, trained with **Applio** from extracted in-game voice audio.
> **Important:** this is a **voice conversion** model, not a standalone text-to-speech model. It converts an existing speech or singing performance into the target timbre while retaining much of the source timing, phrasing, and pitch contour.
## Model files
The supplied Applio/RVC package contains:
| File | Purpose | Size |
|---|---|---:|
| `rogue.pth` | RVC model weights | 57,532,660 bytes / 54.87 MiB |
| `rogue.index` | Retrieval index used to reinforce the target timbre | 333,505,619 bytes / 318.06 MiB |
The uploaded package `rogue.zip` contains both files.
### SHA-256
```text
rogue.pth
b13ed63734886b2894cd667eafa26e4e966c6aea51e4422b6ee7457a71cdc210
rogue.index
dbe27e2d354f4e690b1459a2872a7f65a4d9c9f9efadc822d361abd06264bb9e
rogue.zip
d556db0d73b0e56ee48012b8e180d2e4c19a5b3b841ae7561c409b6a26435a21
```
## Model details
| Property | Value |
|---|---|
| Model type | Retrieval-Based Voice Conversion (RVC) |
| RVC version | **v2** |
| Training / inference app | **Applio** |
| Packaged model filename | `rogue.pth` |
| Packaged index filename | `rogue.index` |
| Internal training model name | `mara` |
| Target / donor voice | `rogue` |
| Sample rate | **48,000 Hz** |
| F0 / pitch guidance | **Enabled** |
| Pitch extractor used for feature extraction | **RMVPE** |
| Embedder | **ContentVec** |
| Vocoder | **HiFi-GAN** |
| Number of speakers | **1** |
| Published checkpoint epoch | **200** |
| Published checkpoint step | **26,200** |
| Embedded dataset length | `00:35:00` |
| Model creation timestamp | `2026-09-09T10:45:38.675607` |
| Embedded model hash | `802d9890b58d28f51309e9db116ff921280d948615abfcbb0e9069c2fb530720` |
### Naming note
The original Applio training sheet uses **`mara`** as the training model name and **`rogue`** as the donor/target voice. The distributed files were packaged as **`rogue.pth`** and **`rogue.index`**. The internal metadata inside `rogue.pth` still reports `model_name: mara`.
## Training dataset
The supplied dataset archive was verified as:
| Property | Value |
|---|---:|
| WAV files | **626** |
| Total duration | **35:00.422** |
| Sample rate | **48,000 Hz** for all 626 files |
| Channels | **Mono** for all 626 files |
| PCM sample width | **16-bit** for all 626 files |
| Shortest clip | **1.210 s** |
| Median clip | **2.900 s** |
| Mean clip | **3.355 s** |
| Longest clip | **11.732 s** |
The archive stores the clips under:
```text
mara/
rogue_*.wav
```
The provided dataset contains audio files rather than a transcription dataset.
### Dataset preprocessing rationale
The source audio was extracted from a game rather than recorded from a microphone for this training run. The training sheet therefore keeps both **Process Effects** and **Noise Reduction** disabled. The source was already produced/mastered audio, and an additional cleanup pass could remove breath, sibilance, and other speaker-specific details useful to voice conversion.
## Applio training configuration
### 1. Preprocess
| Setting | Value |
|---|---|
| Model Name | `mara` |
| Sample Rate | `48000` |
| Cut Preprocess | `Automatic` |
| Chunk Length | `3.0` s |
| Overlap Length | `0.3` s |
| Process Effects | **OFF** |
| Noise Reduction | **OFF** |
| CPU Cores | Default |
### 2. Extract
| Setting | Value |
|---|---|
| Model Name | `mara` |
| Sample Rate | `48000` |
| Pitch Extractor / F0 | **RMVPE** |
| Embedder Model | **ContentVec** |
| GPU | `0` |
### 3. Train
| Setting | Value |
|---|---|
| Model Name | `mara` |
| Sample Rate | `48000` |
| Vocoder | **HiFi-GAN** |
| Batch Size | **8** |
| Configured Total Epochs | **300** |
| Save Every Epoch | **10** |
| Save Only Latest | **OFF** |
| Save Every Weights | **ON** |
| Pretrained | **ON** |
| Cache Dataset in GPU | **OFF** |
| GPU | `0` |
### Published checkpoint vs. configured training length
The training configuration was set to **300 total epochs** and to save weights every 10 epochs. However, the `rogue.pth` file published in this package identifies itself as **epoch 200, step 26,200**.
For that reason, this model card describes the distributed model as the **epoch-200 checkpoint** rather than calling it a 300-epoch model.
## Using the model in Applio
Applio expects an RVC model to use the `.pth` weights file and, when available, the matching `.index` retrieval file.
### Option 1 β Applio model downloader
If `rogue.zip` is uploaded to this Hugging Face repository, its direct file URL can be used in Applio's **Download Model** panel.
The URL format is:
```text
https://huggingface.co/<USERNAME>/<REPOSITORY>/resolve/main/rogue.zip
```
Replace `<USERNAME>` and `<REPOSITORY>` with the actual Hugging Face repository path.
### Option 2 β Manual installation
Extract `rogue.zip` and place both files in one model folder, for example:
```text
Applio/
βββ logs/
βββ Rogue/
βββ rogue.pth
βββ rogue.index
```
Then refresh the model list in Applio and select the matching `.pth` and `.index`.
Official Applio model-installation documentation:
https://docs.applio.org/getting-started/installing-inference-models/
## Example Applio CLI inference
Current Applio versions expose inference through `core.py`.
Example starting point:
```bash
python core.py infer \
--input_path input.wav \
--output_path rogue_output.wav \
--pth_path logs/Rogue/rogue.pth \
--index_path logs/Rogue/rogue.index \
--pitch 0 \
--index_rate 0.75 \
--volume_envelope 1 \
--protect 0.5 \
--f0_method rmvpe \
--embedder_model contentvec
```
The inference values above are **starting values, not training parameters**. Adjust them for the source recording.
Useful notes:
- Keep the embedder as **ContentVec**, matching training.
- **RMVPE** is a sensible default for this model because it was also used during feature extraction.
- `pitch 0` preserves the source pitch by default.
- The retrieval index can improve target-timbre similarity, but an excessively high index influence may also introduce artifacts.
- Source audio quality, pitch range, delivery, noise, and pronunciation can materially affect the result.
Official Applio CLI documentation:
https://docs.applio.org/reference/cli/
## Reproducing the training setup
The supplied training sheet also records a headless helper command:
```bash
tools/rvc_train.sh mara
```
The recorded expected Applio outputs were:
```text
/mnt/sata/Applio/logs/mara/mara.pth
/mnt/sata/Applio/logs/mara/*.index
```
These are local paths from the original training environment and are not required for inference after the files have been packaged.
## Evaluation
No formal objective evaluation results, AB listening test, benchmark scores, or reference inference samples were supplied with the uploaded files.
The package therefore does **not** claim that epoch 200 is objectively superior to another saved checkpoint. Users should evaluate the model by listening on representative speech/singing inputs and tuning inference parameters for their use case.
## Limitations
- Trained from approximately **35 minutes** of a single target voice.
- Training material is game dialogue, so the model may inherit characteristics of that recording, performance style, mastering, or available pitch range.
- Voice conversion quality depends strongly on the input speaker and recording.
- Pronunciation, emotion, singing range, very high/low pitch, background noise, and aggressive processing can reduce similarity or create artifacts.
- No multilingual or cross-language evaluation was provided.
- No formal quality benchmark was provided.
- This model performs voice conversion; it does not generate speech from text on its own.
## Responsible use
Use the model transparently and responsibly.
- Do not present generated audio as an authentic recording of the original performer or as an official game asset.
- Do not use generated audio to mislead, defraud, harass, or impersonate someone deceptively.
- Clearly label synthetic or voice-converted audio when context could otherwise make its origin unclear.
- Respect applicable rights, platform rules, and local law.
## Rights and redistribution
The training material was extracted from game audio. This model card does **not** grant rights to the original game recordings, characters, performances, trademarks, or other underlying assets.
No repository license is declared in the YAML metadata above because the supplied files did not include a license specifying what redistribution or reuse rights can be granted.
If the raw dataset is going to be made public, verify that you have the right to redistribute the source audio. If you do not, a safer public repository layout is to publish the trained model package and documentation without publishing the extracted source WAV files.
This is an unofficial community model and is not presented as affiliated with or endorsed by the game's developers, publishers, performers, or other rights holders.
## Integrity information
Additional uploaded-source checksums:
```text
rogue_datasets.zip
e5531e041a4d3e33baca6879509c58e03d6b7524239c918d800126eeba321fff
SETTINGS.txt
5e6ba830f607c276c3cd1e1acbf0184140e0916c6f8492959df26c282db60fae
```
## Recommended Hugging Face repository layout
For a compact repository that supports Applio's ZIP downloader:
```text
README.md
rogue.zip
SETTINGS.txt
```
If you also want people to access the two RVC files directly, you can additionally upload:
```text
rogue.pth
rogue.index
```
Be aware that keeping both the ZIP and the extracted model files duplicates the model storage.
## Frameworks and references
- Applio: https://github.com/IAHispano/Applio
- Applio documentation: https://docs.applio.org/
- Hugging Face model cards: https://huggingface.co/docs/hub/model-cards
---
### Quick technical summary
```text
RVC version: v2
Sample rate: 48000 Hz
Pitch guidance: yes
Pitch extractor: RMVPE
Embedder: ContentVec
Vocoder: HiFi-GAN
Dataset: 626 mono 16-bit WAV files
Dataset duration: 35:00.422
Batch size: 8
Configured epochs: 300
Published epoch: 200
Published step: 26200
Weights: rogue.pth
Index: rogue.index
```
|