File size: 8,984 Bytes
8733ee3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 | ---
title: RVC-v2 + Beatrice-v2 Voice Conversion + training
emoji: 🎤
colorFrom: red
colorTo: yellow
sdk: gradio
python_version: "3.10"
sdk_version: 6.3.0
app_file: app.py
pinned: false
license: mit
tags:
- voice-conversion
- rvc
- beatrice
- beatrice-v2
- audio
- mcp-server
short_description: RVC-v2 Beatrice-v2 - CPU inference + training
---
# RVC + Beatrice Voice Conversion
**CPU Inference + Training** - Single-file app for HuggingFace Spaces.
## Features
- **Voice Conversion** - RVC v2 (.pth) + Beatrice v2 (.pt.gz)
- **Training** - Train both model types (GPU recommended, CPU works but slow)
- **Single File** - Everything in one `app.py`
- **CLI Support** - Command-line interface for batch processing
## Voice Conversion
1. Upload source audio (any format)
2. Select **Model Type**: RVC v2 or Beatrice v2
3. Get a model into the inputs one of three ways:
- **Upload manually** — .pth for RVC, .pt.gz for Beatrice.
- **Retrieve by tracking code** — under **"📥 بازیابی خودکار مدل"**, paste the tracking code from a finished Training job and click **"🔄 بازیابی مدل"**. The already-trained model (and index, for RVC) is pulled straight from the server — nothing to upload.
- **Load from a URL** — paste a direct link to a `.zip` (containing the `.pth`/`.pt.gz` and optionally `.index`) or to a raw model file, and click **"⬇️ دانلود از لینک"**. The server downloads and unpacks it and drops the files into the right slots automatically — this doesn't use your own upload bandwidth at all, only the server's.
4. Adjust pitch shift if needed
5. Click Convert
**Default model:** [audo/Benee-RVC](https://huggingface.co/audo/Benee-RVC) (RVC v2)
## Training
1. Select **Trainer**: RVC v2 or Beatrice v2
2. Upload training audio, either way:
- **ZIP upload (recommended for many files):** zip up all your training audio files (any mix of .wav/.mp3/.flac/.ogg/.m4a/.aac/.wma/.opus/.aiff, subfolders ok) and upload the ZIP. It's auto-extracted, the file count is auto-detected, and every file is auto-added to the training slots — no manual per-file selection.
- **One-by-one:** start with one file; after each upload, an **"➕ افزودن فایل صوتی دیگر" (add another file)** button appears right below it, letting you add up to 100 files total.
All uploaded files (from either method) are combined (in order, with a short silence gap between them) into one training set.
3. Enter model name
4. Adjust epochs (up to 5000 for RVC v2) and batch size
5. Click **Start Training**
Training runs as a **background job on the server**, not tied to your browser tab:
- As soon as you click Start Training, you instantly get a **tracking code** (e.g. `4F9A2C`) — you don't wait for training to finish.
- You can close the tab or refresh the page; training keeps going on the server.
- Go to the **📡 پیگیری آموزش / Track Training** tab any time, enter your tracking code, and click **Check Status** to see live progress/logs.
- Once training finishes, the same tab lets you download the trained model (and index file, for RVC).
- **You don't have to wait for training to finish.** A checkpoint is saved periodically during the run (e.g. the 600-epoch checkpoint of a 1000-epoch target) — download it from the tracking tab any time, or go to **🎵 Voice Conversion → "📥 بازیابی خودکار مدل"**, enter the same tracking code, and test-convert with that in-progress checkpoint right away while training keeps going in the background.
- You can also cancel a queued or running job from that tab using the same tracking code.
**Note:** CPU training works but is slow. For faster training, clone locally with CUDA GPU. Keep your tracking code — it stays valid as long as the Space itself hasn't restarted.
## Compatibility
- **RVC v1** (256-dim HuBERT) - with f0 or no-f0
- **RVC v2** (768-dim HuBERT) - with f0 or no-f0
- **Beatrice v2** - 16kHz input, 24kHz output, per-speaker VQ
- **Index retrieval** (.index files) for RVC voice matching
Model version and f0 flag are auto-detected from the checkpoint.
Find models: [HuggingFace](https://huggingface.co/models?search=rvc) | [Weights.gg](https://weights.gg)
---
## API
### Python Client - Voice Conversion
```python
from gradio_client import Client, handle_file
client = Client("Luminia/rvc-voice-conversion")
# RVC v2 inference
result = client.predict(
source_audio=handle_file("voice.wav"),
model_type="RVC v2",
model_file=handle_file("model.pth"),
index_file=None, # Optional .index file
beatrice_model_file=None, # Not used for RVC
beatrice_target_speaker=0, # Not used for RVC
beatrice_formant_shift=0.0, # Not used for RVC
pitch_shift=0, # -12 to 12 semitones
f0_method="pm", # "pm" or "harvest"
index_rate=0.75, # 0-1, voice retrieval strength
protect=0.33, # 0-0.5, voiceless consonant protection
api_name="/convert"
)
print(result) # (output_path, status_message)
# Beatrice v2 inference
result = client.predict(
source_audio=handle_file("voice.wav"),
model_type="Beatrice v2",
model_file=None, # Not used for Beatrice
index_file=None, # Not used for Beatrice
beatrice_model_file=handle_file("beatrice_model.pt.gz"),
beatrice_target_speaker=0, # Speaker index
beatrice_formant_shift=0.0, # -2 to 2
pitch_shift=0, # -12 to 12 semitones
f0_method="pm", # Ignored for Beatrice
index_rate=0.75, # Ignored for Beatrice
protect=0.33, # Ignored for Beatrice
api_name="/convert"
)
print(result) # (output_path, status_message)
```
### Python Client - Training (background job + tracking code)
Training is asynchronous: `/train` registers the job and returns a **tracking code**
immediately (it does not wait for training to finish). Poll `/check_training_status`
with that code to get progress/logs, and to fetch the model once it's done.
```python
from gradio_client import Client, handle_file
import time
client = Client("Luminia/rvc-voice-conversion")
# 1) Submit an RVC v2 training request — returns instantly
tracking_code, status_msg = client.predict(
trainer="RVC v2",
train_audio=handle_file("voice.wav"),
train_model_name="my_voice",
train_epochs=200, # 1-5000
train_batch=2, # Batch size
train_sr=40000, # 32000, 40000, or 48000
beatrice_epochs=30, # Ignored for RVC
beatrice_batch=8, # Ignored for RVC
beatrice_resume=False, # Ignored for RVC
api_name="/train"
)
print("Tracking code:", tracking_code)
# 2) Poll status any time later (even after restarting your script)
while True:
log, progress, model_path, index_path = client.predict(
tracking_code, api_name="/check_training_status"
)
print(log, progress)
if model_path or "خطا" in log or "لغو" in log:
break
time.sleep(15)
# 3) Cancel a queued/running job if needed
# client.predict(tracking_code, api_name="/cancel_training")
```
Beatrice v2 training uses the same `/train` call, just with `trainer="Beatrice v2"`
and the `beatrice_*` parameters filled in instead.
### MCP (Model Context Protocol)
This Space supports MCP for AI assistants (Claude Desktop, Cursor, VS Code).
1. Click **MCP** badge → **Add to MCP tools**
2. The `convert` and `train` tools become available
**MCP Config:**
```json
{
"mcpServers": {
"rvc": {"url": "https://luminia-rvc-voice-conversion.hf.space/gradio_api/mcp/"}
}
}
```
---
## CLI Usage
### Inference
```bash
# RVC v2
python app.py infer -i voice.wav -m model.pth -o output.wav
# Beatrice v2 (auto-detected from .pt.gz extension)
python app.py infer -i voice.wav -m beatrice_model.pt.gz -o output.wav
# With pitch shift
python app.py infer -i voice.wav -m model.pth -p 2 -o output.wav
# Beatrice with speaker/formant options
python app.py infer -i voice.wav -m beatrice.pt.gz --speaker 0 --formant-shift 1.0 -o output.wav
```
### Training
```bash
# RVC v2 training
python app.py train -a voice.mp3 -o ./my_model --epochs 100
# Beatrice v2 training
python app.py train-beatrice -a voice.mp3 -o ./beatrice_model --epochs 30
# Beatrice resume training
python app.py train-beatrice -a voice.mp3 -o ./beatrice_model --epochs 30 --resume
```
---
- Local real-time model usage: https://huggingface.co/wok000/vcclient000/tree/main
## Credits
Based on [RVC-Project](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI)-[Mangio-UI-Fork](https://github.com/Mangio621/Mangio-RVC-Fork), [Applio](https://github.com/IAHispano/Applio) data processing, and [Beatrice v2](https://huggingface.co/fierce-cats/beatrice-trainer)
|