OmniVoice Expressive v1 (Nepali Unified TTS)

OmniVoice Expressive v1 (mlwiseyak/omni_expressive_v1) is a multilingual neural text-to-speech (TTS) foundation model fine-tuned for high-fidelity Nepali (ne / npi), English, and Maithili speech synthesis with expressive stylistic control and zero-shot voice cloning.

The model was trained and aligned on 85 shards (approximately 52.4 hours) of verified Devanagari text, spontaneous conversations, formal broadcast audio, and acoustic emotional datasets.

Core Capabilities

  1. Native Nepali Personas:

    • Pooja (Female, Conversational): Natural, friendly conversational tone.
    • Sita (Female, Storyteller): Warm, calm storytelling and narration style.
    • Aayusha (Female, Energetic): Dynamic, youthful conversational delivery.
    • Aarav (Male, Professional Narrator): Standard professional broadcast and documentary delivery.
    • Bibek (Male, Broadcaster): Deep, authoritative news and announcement delivery.
    • Samir (Male, Casual): Relaxed podcast and interview style.
  2. Expressive Emotional Conditioning:

    • Supported emotions: happy, sad, angry, calm, urgent, excited, fear, laughter, sigh, and neutral.
    • Modulated through prompt tokens ([happy], [urgent], etc.) and natural language style descriptions.
  3. Zero-Shot Voice Cloning:

    • Clones speaker characteristics from a 3 to 8 second reference audio sample.
    • Preserves speaker timbre and prosodic characteristics.
  4. Inference Performance:

    • Diffusion-based flow matching architecture.
    • Real-time factor (RTF) below 0.25 on modern GPU hardware.

Inference Example

import re
import torch
import soundfile as sf
from omnivoice import OmniVoice
import omnivoice.models.omnivoice as ov_mod

# 1. Register emotion tags and instruction resolver
ov_mod._NONVERBAL_PATTERN = re.compile(
    r"\[(laughter|sigh|happy|sad|angry|calm|urgent|excited|fear|"
    r"confirmation-en|question-en|question-ah|question-oh|"
    r"question-ei|question-yi|surprise-ah|surprise-oh|surprise-wa|"
    r"surprise-yo|dissatisfaction-hnn)\]"
)

orig_resolve_instruct = ov_mod._resolve_instruct
def custom_resolve_instruct(instruct_str, use_zh=False):
    emotions = {"happy", "sad", "angry", "calm", "urgent", "excited", "fear", "laughter", "sigh", "neutral"}
    items = [x.strip().lower() for x in str(instruct_str).split(",") if x.strip()]
    valid_items = []
    for item in emotions:
        if item in items:
            valid_items.append(item)
    for item in items:
        if item not in emotions:
            try:
                valid_items.append(orig_resolve_instruct(item, use_zh=use_zh))
            except Exception:
                valid_items.append(item)
    return ", ".join(valid_items)
ov_mod._resolve_instruct = custom_resolve_instruct

# 2. Load model from repository
model = OmniVoice.from_pretrained(
    "mlwiseyak/omni_expressive_v1",
    device_map="cuda:0",
    dtype=torch.bfloat16,
)
model.eval()

# 3. Native Persona Synthesis with Style Conditioning
output = model.generate(
    text="[happy] आजको दिन निकै रमाइलो भयो, हाम्रो सबै काम समयमै सफलतापूर्वक सम्पन्न भएको छ।",
    instruct="female, young adult, moderate pitch, happy tone",
    num_step=32,
    speed=1.05,
)
audio_data = output[0].cpu().numpy()
sf.write("output_happy.wav", audio_data, 24000)

# 4. Zero-Shot Voice Cloning
output_cloned = model.generate(
    text="नेपालको प्राकृतिक सौन्दर्य र सांस्कृतिक विविधता विश्वमै अद्वितीय मानिन्छ।",
    ref_audio="reference_sample.wav",
    ref_text="Reference transcript corresponding to the audio clip",
    num_step=32,
)
cloned_audio = output_cloned[0].cpu().numpy()
sf.write("output_cloned.wav", cloned_audio, 24000)

Model Architecture

  • Text and Flow Transformer: Multi-head attention architecture with Rotary Position Embeddings (RoPE).
  • Acoustic Tokenizer: CosyVoice acoustic semantic codebook (12.5 Hz / 24 kHz).
  • Decoder: Flow-matching diffusion ODE sampler with classifier-free guidance support.

Organization

WiseYak AI / Machine Learning Team

Downloads last month
12
Safetensors
Model size
0.8B params
Tensor type
I64
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support