Rvc-cpu / README.md
Transfer Bot
Moved to Hugging Face automatically
8733ee3
|
Raw History Blame Contribute Delete
8.98 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: RVC-v2 + Beatrice-v2 Voice Conversion + training
emoji: 🎤
colorFrom: red
colorTo: yellow
sdk: gradio
python_version: '3.10'
sdk_version: 6.3.0
app_file: app.py
pinned: false
license: mit
tags:
  - voice-conversion
  - rvc
  - beatrice
  - beatrice-v2
  - audio
  - mcp-server
short_description: RVC-v2 Beatrice-v2 - CPU inference + training

RVC + Beatrice Voice Conversion

CPU Inference + Training - Single-file app for HuggingFace Spaces.

Features

  • Voice Conversion - RVC v2 (.pth) + Beatrice v2 (.pt.gz)
  • Training - Train both model types (GPU recommended, CPU works but slow)
  • Single File - Everything in one app.py
  • CLI Support - Command-line interface for batch processing

Voice Conversion

  1. Upload source audio (any format)
  2. Select Model Type: RVC v2 or Beatrice v2
  3. Get a model into the inputs one of three ways:
    • Upload manually — .pth for RVC, .pt.gz for Beatrice.
    • Retrieve by tracking code — under "📥 بازیابی خودکار مدل", paste the tracking code from a finished Training job and click "🔄 بازیابی مدل". The already-trained model (and index, for RVC) is pulled straight from the server — nothing to upload.
    • Load from a URL — paste a direct link to a .zip (containing the .pth/.pt.gz and optionally .index) or to a raw model file, and click "⬇️ دانلود از لینک". The server downloads and unpacks it and drops the files into the right slots automatically — this doesn't use your own upload bandwidth at all, only the server's.
  4. Adjust pitch shift if needed
  5. Click Convert

Default model: audo/Benee-RVC (RVC v2)

Training

  1. Select Trainer: RVC v2 or Beatrice v2

  2. Upload training audio, either way:

    • ZIP upload (recommended for many files): zip up all your training audio files (any mix of .wav/.mp3/.flac/.ogg/.m4a/.aac/.wma/.opus/.aiff, subfolders ok) and upload the ZIP. It's auto-extracted, the file count is auto-detected, and every file is auto-added to the training slots — no manual per-file selection.
    • One-by-one: start with one file; after each upload, an "➕ افزودن فایل صوتی دیگر" (add another file) button appears right below it, letting you add up to 100 files total.

    All uploaded files (from either method) are combined (in order, with a short silence gap between them) into one training set.

  3. Enter model name

  4. Adjust epochs (up to 5000 for RVC v2) and batch size

  5. Click Start Training

Training runs as a background job on the server, not tied to your browser tab:

  • As soon as you click Start Training, you instantly get a tracking code (e.g. 4F9A2C) — you don't wait for training to finish.
  • You can close the tab or refresh the page; training keeps going on the server.
  • Go to the 📡 پیگیری آموزش / Track Training tab any time, enter your tracking code, and click Check Status to see live progress/logs.
  • Once training finishes, the same tab lets you download the trained model (and index file, for RVC).
  • You don't have to wait for training to finish. A checkpoint is saved periodically during the run (e.g. the 600-epoch checkpoint of a 1000-epoch target) — download it from the tracking tab any time, or go to 🎵 Voice Conversion → "📥 بازیابی خودکار مدل", enter the same tracking code, and test-convert with that in-progress checkpoint right away while training keeps going in the background.
  • You can also cancel a queued or running job from that tab using the same tracking code.

Note: CPU training works but is slow. For faster training, clone locally with CUDA GPU. Keep your tracking code — it stays valid as long as the Space itself hasn't restarted.

Compatibility

  • RVC v1 (256-dim HuBERT) - with f0 or no-f0
  • RVC v2 (768-dim HuBERT) - with f0 or no-f0
  • Beatrice v2 - 16kHz input, 24kHz output, per-speaker VQ
  • Index retrieval (.index files) for RVC voice matching

Model version and f0 flag are auto-detected from the checkpoint.

Find models: HuggingFace | Weights.gg


API

Python Client - Voice Conversion

from gradio_client import Client, handle_file

client = Client("Luminia/rvc-voice-conversion")

# RVC v2 inference
result = client.predict(
    source_audio=handle_file("voice.wav"),
    model_type="RVC v2",
    model_file=handle_file("model.pth"),
    index_file=None,                  # Optional .index file
    beatrice_model_file=None,         # Not used for RVC
    beatrice_target_speaker=0,        # Not used for RVC
    beatrice_formant_shift=0.0,       # Not used for RVC
    pitch_shift=0,                    # -12 to 12 semitones
    f0_method="pm",                   # "pm" or "harvest"
    index_rate=0.75,                  # 0-1, voice retrieval strength
    protect=0.33,                     # 0-0.5, voiceless consonant protection
    api_name="/convert"
)
print(result)  # (output_path, status_message)

# Beatrice v2 inference
result = client.predict(
    source_audio=handle_file("voice.wav"),
    model_type="Beatrice v2",
    model_file=None,                  # Not used for Beatrice
    index_file=None,                  # Not used for Beatrice
    beatrice_model_file=handle_file("beatrice_model.pt.gz"),
    beatrice_target_speaker=0,        # Speaker index
    beatrice_formant_shift=0.0,       # -2 to 2
    pitch_shift=0,                    # -12 to 12 semitones
    f0_method="pm",                   # Ignored for Beatrice
    index_rate=0.75,                  # Ignored for Beatrice
    protect=0.33,                     # Ignored for Beatrice
    api_name="/convert"
)
print(result)  # (output_path, status_message)

Python Client - Training (background job + tracking code)

Training is asynchronous: /train registers the job and returns a tracking code immediately (it does not wait for training to finish). Poll /check_training_status with that code to get progress/logs, and to fetch the model once it's done.

from gradio_client import Client, handle_file
import time

client = Client("Luminia/rvc-voice-conversion")

# 1) Submit an RVC v2 training request — returns instantly
tracking_code, status_msg = client.predict(
    trainer="RVC v2",
    train_audio=handle_file("voice.wav"),
    train_model_name="my_voice",
    train_epochs=200,                 # 1-5000
    train_batch=2,                    # Batch size
    train_sr=40000,                   # 32000, 40000, or 48000
    beatrice_epochs=30,               # Ignored for RVC
    beatrice_batch=8,                 # Ignored for RVC
    beatrice_resume=False,            # Ignored for RVC
    api_name="/train"
)
print("Tracking code:", tracking_code)

# 2) Poll status any time later (even after restarting your script)
while True:
    log, progress, model_path, index_path = client.predict(
        tracking_code, api_name="/check_training_status"
    )
    print(log, progress)
    if model_path or "خطا" in log or "لغو" in log:
        break
    time.sleep(15)

# 3) Cancel a queued/running job if needed
# client.predict(tracking_code, api_name="/cancel_training")

Beatrice v2 training uses the same /train call, just with trainer="Beatrice v2" and the beatrice_* parameters filled in instead.

MCP (Model Context Protocol)

This Space supports MCP for AI assistants (Claude Desktop, Cursor, VS Code).

  1. Click MCP badge → Add to MCP tools
  2. The convert and train tools become available

MCP Config:

{
  "mcpServers": {
    "rvc": {"url": "https://luminia-rvc-voice-conversion.hf.space/gradio_api/mcp/"}
  }
}

CLI Usage

Inference

# RVC v2
python app.py infer -i voice.wav -m model.pth -o output.wav

# Beatrice v2 (auto-detected from .pt.gz extension)
python app.py infer -i voice.wav -m beatrice_model.pt.gz -o output.wav

# With pitch shift
python app.py infer -i voice.wav -m model.pth -p 2 -o output.wav

# Beatrice with speaker/formant options
python app.py infer -i voice.wav -m beatrice.pt.gz --speaker 0 --formant-shift 1.0 -o output.wav

Training

# RVC v2 training
python app.py train -a voice.mp3 -o ./my_model --epochs 100

# Beatrice v2 training
python app.py train-beatrice -a voice.mp3 -o ./beatrice_model --epochs 30

# Beatrice resume training
python app.py train-beatrice -a voice.mp3 -o ./beatrice_model --epochs 30 --resume

Credits

Based on RVC-Project-Mangio-UI-Fork, Applio data processing, and Beatrice v2