SteerTTS · Mamba-3 hybrid (steer-tts-m3hyb)

A small (9.4M parameters, plus the 13M Vocos vocoder), steerable, expressive English TTS from Whissle. You can steer the delivery in three ways, and combine them:

  1. A free-text prompt, e.g. "furious, shouting" or "soft and sad". A small adapter on MiniLM sentence embeddings maps it to a style tag plus controls.
  2. A style tag (33 tags), e.g. anger, whisper, happiness, fast or low_pitch.
  3. Explicit acoustic controls: 12 interpretable, speaker-relative z-scores (rate, loud, f0_med, f0_iqr, f0_end, pause, tilt, voiced, flat, centroid, cpps, h1h2). 0 is the voice's own neutral speech and ±1 is one standard deviation.

There's also per-word prosody editing (pitch in semitones, energy in dB, duration factor).

Architecture: a FastPitch-style non-autoregressive model with explicit duration, pitch and energy prediction and a jointly trained aligner. Each encoder and decoder block is a hybrid: self-attention followed by a bidirectional Mamba-3 branch (Lahoti et al., arXiv 2603.15569, in a pure-PyTorch SSD implementation). Every block is conditioned through adaptive LayerNorm on speaker + tag + control vector. The output is a 100-band 24 kHz log-mel, rendered by charactr/vocos-mel-24khz.

Status: this is a research preview, taken mid-training at step 116k of 150k. Voices are a fixed set of 287 training speakers; zero-shot cloning is not included in this checkpoint.

Quick start

pip install torch numpy safetensors huggingface_hub vocos g2p_en soundfile sentence-transformers
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("WhissleAI/steer-tts-m3hyb")
sys.path.insert(0, path)  # the repo ships its own self-contained `steertts` package

from steertts import SteerTTSPipeline
tts = SteerTTSPipeline(path)  # cuda if available, else cpu

wav = tts("I can't believe you actually did that.", voice="ears_p004", prompt="furious and shouting")
tts.save(wav, "angry.wav")

wav = tts("Come closer, I have a secret.", voice="expresso_ex02", tag="whisper", ctrl={"rate": -0.8})
wav = tts("Are you serious?", voice="libritts_89", ctrl={"f0_end": 2.0, "f0_iqr": 1.5})  # rising, lively
wav = tts("I never said she stole it.", voice="ears_p010", word_shift={2: {"pitch": 4, "energy": 6, "dur": 1.4}})

From the command line:

python inference.py --text "Hello from Whissle." --voice ears_p004 --prompt "cheerful and fast" --out hello.wav
python inference.py --list-voices

Or via the Whissle Python SDK (pip install "whissle-sdk[tts-prompt] @ git+https://github.com/WhissleAI/whissle-python"):

from whissle_sdk.tts import SteerTTS
tts = SteerTTS.from_pretrained()
tts.save(tts("Hello from Whissle.", voice="ears_p004", prompt="cheerful"), "hello.wav")

Voices and tags

  • Voices: ears_p001 … ears_p065 (EARS, the most expressive), expresso_ex01 … expresso_ex04 (Expresso), and 218 libritts_<id> voices (LibriTTS-R). The full list is tts.voices.
  • Tags: none neutral plus emotions (adoration amazement amusement anger confusion contentment cuteness desire disappointment disgust distress ecstasy embarrassment fear guilt happiness interest laughing pain pride realization relief sadness serenity) and styles (emphasis enunciated fast slow loud whisper high_pitch low_pitch). Emotion tags steer best on EARS and Expresso voices.

Evaluation (step 58k of the same run; updated numbers will follow)

set metric this model attention-only baseline ground truth
LibriTTS-R test-clean (500) WER whisper-large-v3 % 2.12 2.04 1.98
LibriTTS-R test-clean (500) UTMOS22 4.12 3.99 4.26
ParaSpeechCaps test (246 prompts) WER distil-whisper-large-v2 % 3.07 3.31 —

For reference, Parler-TTS scores 4.62 WER on ParaSpeechCaps under the same ASR. Pitch-level and rate attribute accuracy from text prompts (0.6) is still well below real speech (0.9), and that is the main known gap.

Training data and license

The model was trained on EARS (CC-BY-NC 4.0), Expresso (CC-BY-NC 4.0) and LibriTTS-R subsets (CC-BY 4.0), about 53 h in total. Because of the non-commercial datasets, these weights are released under CC-BY-NC 4.0. Contact Whissle for commercial licensing.

Limitations

  • English only, with a g2p_en front end (no explicit normalization of numbers or abbreviations beyond g2p_en's).
  • Fixed voice set; no cloning in this checkpoint.
  • The prompt adapter understands short style descriptions. It can miss modifiers buried in long prompts (e.g. "bored" → monotone).
  • Don't use it to impersonate real people or produce deceptive audio.
Downloads last month
10
Safetensors
Model size
9.44M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train WhissleAI/steer-tts-9m