| --- |
| pipeline_tag: audio-to-audio |
| tags: |
| - rvc |
| - rvc-v2 |
| - applio |
| - voice-conversion |
| - retrieval-based-voice-conversion |
| - contentvec |
| - rmvpe |
| - hifigan |
| - 48khz |
| - cyberpunk-2077 |
| - rogue |
| --- |
| |
| # Rogue — RVC v2 Voice Conversion Model for Applio |
|
|
| RVC v2 voice-conversion model for **Rogue**, trained with **Applio** from extracted in-game voice audio. |
|
|
| > **Important:** this is a **voice conversion** model, not a standalone text-to-speech model. It converts an existing speech or singing performance into the target timbre while retaining much of the source timing, phrasing, and pitch contour. |
|
|
| ## Model files |
|
|
| The supplied Applio/RVC package contains: |
|
|
| | File | Purpose | Size | |
| |---|---|---:| |
| | `rogue.pth` | RVC model weights | 57,532,660 bytes / 54.87 MiB | |
| | `rogue.index` | Retrieval index used to reinforce the target timbre | 333,505,619 bytes / 318.06 MiB | |
|
|
| The uploaded package `rogue.zip` contains both files. |
|
|
| ### SHA-256 |
|
|
| ```text |
| rogue.pth |
| b13ed63734886b2894cd667eafa26e4e966c6aea51e4422b6ee7457a71cdc210 |
| |
| rogue.index |
| dbe27e2d354f4e690b1459a2872a7f65a4d9c9f9efadc822d361abd06264bb9e |
| |
| rogue.zip |
| d556db0d73b0e56ee48012b8e180d2e4c19a5b3b841ae7561c409b6a26435a21 |
| ``` |
|
|
| ## Model details |
|
|
| | Property | Value | |
| |---|---| |
| | Model type | Retrieval-Based Voice Conversion (RVC) | |
| | RVC version | **v2** | |
| | Training / inference app | **Applio** | |
| | Packaged model filename | `rogue.pth` | |
| | Packaged index filename | `rogue.index` | |
| | Internal training model name | `mara` | |
| | Target / donor voice | `rogue` | |
| | Sample rate | **48,000 Hz** | |
| | F0 / pitch guidance | **Enabled** | |
| | Pitch extractor used for feature extraction | **RMVPE** | |
| | Embedder | **ContentVec** | |
| | Vocoder | **HiFi-GAN** | |
| | Number of speakers | **1** | |
| | Published checkpoint epoch | **200** | |
| | Published checkpoint step | **26,200** | |
| | Embedded dataset length | `00:35:00` | |
| | Model creation timestamp | `2026-09-09T10:45:38.675607` | |
| | Embedded model hash | `802d9890b58d28f51309e9db116ff921280d948615abfcbb0e9069c2fb530720` | |
|
|
| ### Naming note |
|
|
| The original Applio training sheet uses **`mara`** as the training model name and **`rogue`** as the donor/target voice. The distributed files were packaged as **`rogue.pth`** and **`rogue.index`**. The internal metadata inside `rogue.pth` still reports `model_name: mara`. |
|
|
| ## Training dataset |
|
|
| The supplied dataset archive was verified as: |
|
|
| | Property | Value | |
| |---|---:| |
| | WAV files | **626** | |
| | Total duration | **35:00.422** | |
| | Sample rate | **48,000 Hz** for all 626 files | |
| | Channels | **Mono** for all 626 files | |
| | PCM sample width | **16-bit** for all 626 files | |
| | Shortest clip | **1.210 s** | |
| | Median clip | **2.900 s** | |
| | Mean clip | **3.355 s** | |
| | Longest clip | **11.732 s** | |
|
|
| The archive stores the clips under: |
|
|
| ```text |
| mara/ |
| rogue_*.wav |
| ``` |
|
|
| The provided dataset contains audio files rather than a transcription dataset. |
|
|
| ### Dataset preprocessing rationale |
|
|
| The source audio was extracted from a game rather than recorded from a microphone for this training run. The training sheet therefore keeps both **Process Effects** and **Noise Reduction** disabled. The source was already produced/mastered audio, and an additional cleanup pass could remove breath, sibilance, and other speaker-specific details useful to voice conversion. |
|
|
| ## Applio training configuration |
|
|
| ### 1. Preprocess |
|
|
| | Setting | Value | |
| |---|---| |
| | Model Name | `mara` | |
| | Sample Rate | `48000` | |
| | Cut Preprocess | `Automatic` | |
| | Chunk Length | `3.0` s | |
| | Overlap Length | `0.3` s | |
| | Process Effects | **OFF** | |
| | Noise Reduction | **OFF** | |
| | CPU Cores | Default | |
|
|
| ### 2. Extract |
|
|
| | Setting | Value | |
| |---|---| |
| | Model Name | `mara` | |
| | Sample Rate | `48000` | |
| | Pitch Extractor / F0 | **RMVPE** | |
| | Embedder Model | **ContentVec** | |
| | GPU | `0` | |
|
|
| ### 3. Train |
|
|
| | Setting | Value | |
| |---|---| |
| | Model Name | `mara` | |
| | Sample Rate | `48000` | |
| | Vocoder | **HiFi-GAN** | |
| | Batch Size | **8** | |
| | Configured Total Epochs | **300** | |
| | Save Every Epoch | **10** | |
| | Save Only Latest | **OFF** | |
| | Save Every Weights | **ON** | |
| | Pretrained | **ON** | |
| | Cache Dataset in GPU | **OFF** | |
| | GPU | `0` | |
|
|
| ### Published checkpoint vs. configured training length |
|
|
| The training configuration was set to **300 total epochs** and to save weights every 10 epochs. However, the `rogue.pth` file published in this package identifies itself as **epoch 200, step 26,200**. |
|
|
| For that reason, this model card describes the distributed model as the **epoch-200 checkpoint** rather than calling it a 300-epoch model. |
|
|
| ## Using the model in Applio |
|
|
| Applio expects an RVC model to use the `.pth` weights file and, when available, the matching `.index` retrieval file. |
|
|
| ### Option 1 — Applio model downloader |
|
|
| If `rogue.zip` is uploaded to this Hugging Face repository, its direct file URL can be used in Applio's **Download Model** panel. |
|
|
| The URL format is: |
|
|
| ```text |
| https://huggingface.co/<USERNAME>/<REPOSITORY>/resolve/main/rogue.zip |
| ``` |
|
|
| Replace `<USERNAME>` and `<REPOSITORY>` with the actual Hugging Face repository path. |
|
|
| ### Option 2 — Manual installation |
|
|
| Extract `rogue.zip` and place both files in one model folder, for example: |
|
|
| ```text |
| Applio/ |
| └── logs/ |
| └── Rogue/ |
| ├── rogue.pth |
| └── rogue.index |
| ``` |
|
|
| Then refresh the model list in Applio and select the matching `.pth` and `.index`. |
|
|
| Official Applio model-installation documentation: |
|
|
| https://docs.applio.org/getting-started/installing-inference-models/ |
|
|
| ## Example Applio CLI inference |
|
|
| Current Applio versions expose inference through `core.py`. |
|
|
| Example starting point: |
|
|
| ```bash |
| python core.py infer \ |
| --input_path input.wav \ |
| --output_path rogue_output.wav \ |
| --pth_path logs/Rogue/rogue.pth \ |
| --index_path logs/Rogue/rogue.index \ |
| --pitch 0 \ |
| --index_rate 0.75 \ |
| --volume_envelope 1 \ |
| --protect 0.5 \ |
| --f0_method rmvpe \ |
| --embedder_model contentvec |
| ``` |
|
|
| The inference values above are **starting values, not training parameters**. Adjust them for the source recording. |
|
|
| Useful notes: |
|
|
| - Keep the embedder as **ContentVec**, matching training. |
| - **RMVPE** is a sensible default for this model because it was also used during feature extraction. |
| - `pitch 0` preserves the source pitch by default. |
| - The retrieval index can improve target-timbre similarity, but an excessively high index influence may also introduce artifacts. |
| - Source audio quality, pitch range, delivery, noise, and pronunciation can materially affect the result. |
|
|
| Official Applio CLI documentation: |
|
|
| https://docs.applio.org/reference/cli/ |
|
|
| ## Reproducing the training setup |
|
|
| The supplied training sheet also records a headless helper command: |
|
|
| ```bash |
| tools/rvc_train.sh mara |
| ``` |
|
|
| The recorded expected Applio outputs were: |
|
|
| ```text |
| /mnt/sata/Applio/logs/mara/mara.pth |
| /mnt/sata/Applio/logs/mara/*.index |
| ``` |
|
|
| These are local paths from the original training environment and are not required for inference after the files have been packaged. |
|
|
| ## Evaluation |
|
|
| No formal objective evaluation results, AB listening test, benchmark scores, or reference inference samples were supplied with the uploaded files. |
|
|
| The package therefore does **not** claim that epoch 200 is objectively superior to another saved checkpoint. Users should evaluate the model by listening on representative speech/singing inputs and tuning inference parameters for their use case. |
|
|
| ## Limitations |
|
|
| - Trained from approximately **35 minutes** of a single target voice. |
| - Training material is game dialogue, so the model may inherit characteristics of that recording, performance style, mastering, or available pitch range. |
| - Voice conversion quality depends strongly on the input speaker and recording. |
| - Pronunciation, emotion, singing range, very high/low pitch, background noise, and aggressive processing can reduce similarity or create artifacts. |
| - No multilingual or cross-language evaluation was provided. |
| - No formal quality benchmark was provided. |
| - This model performs voice conversion; it does not generate speech from text on its own. |
|
|
| ## Responsible use |
|
|
| Use the model transparently and responsibly. |
|
|
| - Do not present generated audio as an authentic recording of the original performer or as an official game asset. |
| - Do not use generated audio to mislead, defraud, harass, or impersonate someone deceptively. |
| - Clearly label synthetic or voice-converted audio when context could otherwise make its origin unclear. |
| - Respect applicable rights, platform rules, and local law. |
|
|
| ## Rights and redistribution |
|
|
| The training material was extracted from game audio. This model card does **not** grant rights to the original game recordings, characters, performances, trademarks, or other underlying assets. |
|
|
| No repository license is declared in the YAML metadata above because the supplied files did not include a license specifying what redistribution or reuse rights can be granted. |
|
|
| If the raw dataset is going to be made public, verify that you have the right to redistribute the source audio. If you do not, a safer public repository layout is to publish the trained model package and documentation without publishing the extracted source WAV files. |
|
|
| This is an unofficial community model and is not presented as affiliated with or endorsed by the game's developers, publishers, performers, or other rights holders. |
|
|
| ## Integrity information |
|
|
| Additional uploaded-source checksums: |
|
|
| ```text |
| rogue_datasets.zip |
| e5531e041a4d3e33baca6879509c58e03d6b7524239c918d800126eeba321fff |
| |
| SETTINGS.txt |
| 5e6ba830f607c276c3cd1e1acbf0184140e0916c6f8492959df26c282db60fae |
| ``` |
|
|
| ## Recommended Hugging Face repository layout |
|
|
| For a compact repository that supports Applio's ZIP downloader: |
|
|
| ```text |
| README.md |
| rogue.zip |
| SETTINGS.txt |
| ``` |
|
|
| If you also want people to access the two RVC files directly, you can additionally upload: |
|
|
| ```text |
| rogue.pth |
| rogue.index |
| ``` |
|
|
| Be aware that keeping both the ZIP and the extracted model files duplicates the model storage. |
|
|
| ## Frameworks and references |
|
|
| - Applio: https://github.com/IAHispano/Applio |
| - Applio documentation: https://docs.applio.org/ |
| - Hugging Face model cards: https://huggingface.co/docs/hub/model-cards |
|
|
| --- |
|
|
| ### Quick technical summary |
|
|
| ```text |
| RVC version: v2 |
| Sample rate: 48000 Hz |
| Pitch guidance: yes |
| Pitch extractor: RMVPE |
| Embedder: ContentVec |
| Vocoder: HiFi-GAN |
| Dataset: 626 mono 16-bit WAV files |
| Dataset duration: 35:00.422 |
| Batch size: 8 |
| Configured epochs: 300 |
| Published epoch: 200 |
| Published step: 26200 |
| Weights: rogue.pth |
| Index: rogue.index |
| ``` |
|
|