Rogue-RVC / README.md
0xra's picture
Upload README.md
1eeaa77 verified
|
Raw
History Blame Contribute Delete
10.6 kB
metadata
pipeline_tag: audio-to-audio
tags:
  - rvc
  - rvc-v2
  - applio
  - voice-conversion
  - retrieval-based-voice-conversion
  - contentvec
  - rmvpe
  - hifigan
  - 48khz
  - cyberpunk-2077
  - rogue

Rogue β€” RVC v2 Voice Conversion Model for Applio

RVC v2 voice-conversion model for Rogue, trained with Applio from extracted in-game voice audio.

Important: this is a voice conversion model, not a standalone text-to-speech model. It converts an existing speech or singing performance into the target timbre while retaining much of the source timing, phrasing, and pitch contour.

Model files

The supplied Applio/RVC package contains:

File Purpose Size
rogue.pth RVC model weights 57,532,660 bytes / 54.87 MiB
rogue.index Retrieval index used to reinforce the target timbre 333,505,619 bytes / 318.06 MiB

The uploaded package rogue.zip contains both files.

SHA-256

rogue.pth
b13ed63734886b2894cd667eafa26e4e966c6aea51e4422b6ee7457a71cdc210

rogue.index
dbe27e2d354f4e690b1459a2872a7f65a4d9c9f9efadc822d361abd06264bb9e

rogue.zip
d556db0d73b0e56ee48012b8e180d2e4c19a5b3b841ae7561c409b6a26435a21

Model details

Property Value
Model type Retrieval-Based Voice Conversion (RVC)
RVC version v2
Training / inference app Applio
Packaged model filename rogue.pth
Packaged index filename rogue.index
Internal training model name mara
Target / donor voice rogue
Sample rate 48,000 Hz
F0 / pitch guidance Enabled
Pitch extractor used for feature extraction RMVPE
Embedder ContentVec
Vocoder HiFi-GAN
Number of speakers 1
Published checkpoint epoch 200
Published checkpoint step 26,200
Embedded dataset length 00:35:00
Model creation timestamp 2026-09-09T10:45:38.675607
Embedded model hash 802d9890b58d28f51309e9db116ff921280d948615abfcbb0e9069c2fb530720

Naming note

The original Applio training sheet uses mara as the training model name and rogue as the donor/target voice. The distributed files were packaged as rogue.pth and rogue.index. The internal metadata inside rogue.pth still reports model_name: mara.

Training dataset

The supplied dataset archive was verified as:

Property Value
WAV files 626
Total duration 35:00.422
Sample rate 48,000 Hz for all 626 files
Channels Mono for all 626 files
PCM sample width 16-bit for all 626 files
Shortest clip 1.210 s
Median clip 2.900 s
Mean clip 3.355 s
Longest clip 11.732 s

The archive stores the clips under:

mara/
  rogue_*.wav

The provided dataset contains audio files rather than a transcription dataset.

Dataset preprocessing rationale

The source audio was extracted from a game rather than recorded from a microphone for this training run. The training sheet therefore keeps both Process Effects and Noise Reduction disabled. The source was already produced/mastered audio, and an additional cleanup pass could remove breath, sibilance, and other speaker-specific details useful to voice conversion.

Applio training configuration

1. Preprocess

Setting Value
Model Name mara
Sample Rate 48000
Cut Preprocess Automatic
Chunk Length 3.0 s
Overlap Length 0.3 s
Process Effects OFF
Noise Reduction OFF
CPU Cores Default

2. Extract

Setting Value
Model Name mara
Sample Rate 48000
Pitch Extractor / F0 RMVPE
Embedder Model ContentVec
GPU 0

3. Train

Setting Value
Model Name mara
Sample Rate 48000
Vocoder HiFi-GAN
Batch Size 8
Configured Total Epochs 300
Save Every Epoch 10
Save Only Latest OFF
Save Every Weights ON
Pretrained ON
Cache Dataset in GPU OFF
GPU 0

Published checkpoint vs. configured training length

The training configuration was set to 300 total epochs and to save weights every 10 epochs. However, the rogue.pth file published in this package identifies itself as epoch 200, step 26,200.

For that reason, this model card describes the distributed model as the epoch-200 checkpoint rather than calling it a 300-epoch model.

Using the model in Applio

Applio expects an RVC model to use the .pth weights file and, when available, the matching .index retrieval file.

Option 1 β€” Applio model downloader

If rogue.zip is uploaded to this Hugging Face repository, its direct file URL can be used in Applio's Download Model panel.

The URL format is:

https://huggingface.co/<USERNAME>/<REPOSITORY>/resolve/main/rogue.zip

Replace <USERNAME> and <REPOSITORY> with the actual Hugging Face repository path.

Option 2 β€” Manual installation

Extract rogue.zip and place both files in one model folder, for example:

Applio/
└── logs/
    └── Rogue/
        β”œβ”€β”€ rogue.pth
        └── rogue.index

Then refresh the model list in Applio and select the matching .pth and .index.

Official Applio model-installation documentation:

https://docs.applio.org/getting-started/installing-inference-models/

Example Applio CLI inference

Current Applio versions expose inference through core.py.

Example starting point:

python core.py infer \
  --input_path input.wav \
  --output_path rogue_output.wav \
  --pth_path logs/Rogue/rogue.pth \
  --index_path logs/Rogue/rogue.index \
  --pitch 0 \
  --index_rate 0.75 \
  --volume_envelope 1 \
  --protect 0.5 \
  --f0_method rmvpe \
  --embedder_model contentvec

The inference values above are starting values, not training parameters. Adjust them for the source recording.

Useful notes:

  • Keep the embedder as ContentVec, matching training.
  • RMVPE is a sensible default for this model because it was also used during feature extraction.
  • pitch 0 preserves the source pitch by default.
  • The retrieval index can improve target-timbre similarity, but an excessively high index influence may also introduce artifacts.
  • Source audio quality, pitch range, delivery, noise, and pronunciation can materially affect the result.

Official Applio CLI documentation:

https://docs.applio.org/reference/cli/

Reproducing the training setup

The supplied training sheet also records a headless helper command:

tools/rvc_train.sh mara

The recorded expected Applio outputs were:

/mnt/sata/Applio/logs/mara/mara.pth
/mnt/sata/Applio/logs/mara/*.index

These are local paths from the original training environment and are not required for inference after the files have been packaged.

Evaluation

No formal objective evaluation results, AB listening test, benchmark scores, or reference inference samples were supplied with the uploaded files.

The package therefore does not claim that epoch 200 is objectively superior to another saved checkpoint. Users should evaluate the model by listening on representative speech/singing inputs and tuning inference parameters for their use case.

Limitations

  • Trained from approximately 35 minutes of a single target voice.
  • Training material is game dialogue, so the model may inherit characteristics of that recording, performance style, mastering, or available pitch range.
  • Voice conversion quality depends strongly on the input speaker and recording.
  • Pronunciation, emotion, singing range, very high/low pitch, background noise, and aggressive processing can reduce similarity or create artifacts.
  • No multilingual or cross-language evaluation was provided.
  • No formal quality benchmark was provided.
  • This model performs voice conversion; it does not generate speech from text on its own.

Responsible use

Use the model transparently and responsibly.

  • Do not present generated audio as an authentic recording of the original performer or as an official game asset.
  • Do not use generated audio to mislead, defraud, harass, or impersonate someone deceptively.
  • Clearly label synthetic or voice-converted audio when context could otherwise make its origin unclear.
  • Respect applicable rights, platform rules, and local law.

Rights and redistribution

The training material was extracted from game audio. This model card does not grant rights to the original game recordings, characters, performances, trademarks, or other underlying assets.

No repository license is declared in the YAML metadata above because the supplied files did not include a license specifying what redistribution or reuse rights can be granted.

If the raw dataset is going to be made public, verify that you have the right to redistribute the source audio. If you do not, a safer public repository layout is to publish the trained model package and documentation without publishing the extracted source WAV files.

This is an unofficial community model and is not presented as affiliated with or endorsed by the game's developers, publishers, performers, or other rights holders.

Integrity information

Additional uploaded-source checksums:

rogue_datasets.zip
e5531e041a4d3e33baca6879509c58e03d6b7524239c918d800126eeba321fff

SETTINGS.txt
5e6ba830f607c276c3cd1e1acbf0184140e0916c6f8492959df26c282db60fae

Recommended Hugging Face repository layout

For a compact repository that supports Applio's ZIP downloader:

README.md
rogue.zip
SETTINGS.txt

If you also want people to access the two RVC files directly, you can additionally upload:

rogue.pth
rogue.index

Be aware that keeping both the ZIP and the extracted model files duplicates the model storage.

Frameworks and references


Quick technical summary

RVC version:       v2
Sample rate:       48000 Hz
Pitch guidance:    yes
Pitch extractor:   RMVPE
Embedder:          ContentVec
Vocoder:           HiFi-GAN
Dataset:           626 mono 16-bit WAV files
Dataset duration:  35:00.422
Batch size:        8
Configured epochs: 300
Published epoch:   200
Published step:    26200
Weights:           rogue.pth
Index:             rogue.index