SteerTTS · Mamba-3 hybrid (steer-tts-m3hyb)
A small (9.4M parameters, plus the 13M Vocos vocoder), steerable, expressive English TTS from Whissle. You can steer the delivery in three ways, and combine them:
- A free-text prompt, e.g.
"furious, shouting"or"soft and sad". A small adapter on MiniLM sentence embeddings maps it to a style tag plus controls. - A style tag (33 tags), e.g.
anger,whisper,happiness,fastorlow_pitch. - Explicit acoustic controls: 12 interpretable, speaker-relative z-scores (
rate,loud,f0_med,f0_iqr,f0_end,pause,tilt,voiced,flat,centroid,cpps,h1h2).0is the voice's own neutral speech and±1is one standard deviation.
There's also per-word prosody editing (pitch in semitones, energy in dB, duration factor).
Architecture: a FastPitch-style non-autoregressive model with explicit duration, pitch and energy prediction and a jointly trained aligner. Each encoder and decoder block is a hybrid: self-attention followed by a bidirectional Mamba-3 branch (Lahoti et al., arXiv 2603.15569, in a pure-PyTorch SSD implementation). Every block is conditioned through adaptive LayerNorm on speaker + tag + control vector. The output is a 100-band 24 kHz log-mel, rendered by charactr/vocos-mel-24khz.
Status: this is a research preview, taken mid-training at step 116k of 150k. Voices are a fixed set of 287 training speakers; zero-shot cloning is not included in this checkpoint.
Quick start
pip install torch numpy safetensors huggingface_hub vocos g2p_en soundfile sentence-transformers
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("WhissleAI/steer-tts-m3hyb")
sys.path.insert(0, path) # the repo ships its own self-contained `steertts` package
from steertts import SteerTTSPipeline
tts = SteerTTSPipeline(path) # cuda if available, else cpu
wav = tts("I can't believe you actually did that.", voice="ears_p004", prompt="furious and shouting")
tts.save(wav, "angry.wav")
wav = tts("Come closer, I have a secret.", voice="expresso_ex02", tag="whisper", ctrl={"rate": -0.8})
wav = tts("Are you serious?", voice="libritts_89", ctrl={"f0_end": 2.0, "f0_iqr": 1.5}) # rising, lively
wav = tts("I never said she stole it.", voice="ears_p010", word_shift={2: {"pitch": 4, "energy": 6, "dur": 1.4}})
From the command line:
python inference.py --text "Hello from Whissle." --voice ears_p004 --prompt "cheerful and fast" --out hello.wav
python inference.py --list-voices
Or via the Whissle Python SDK (pip install "whissle-sdk[tts-prompt] @ git+https://github.com/WhissleAI/whissle-python"):
from whissle_sdk.tts import SteerTTS
tts = SteerTTS.from_pretrained()
tts.save(tts("Hello from Whissle.", voice="ears_p004", prompt="cheerful"), "hello.wav")
Voices and tags
- Voices:
ears_p001…ears_p065(EARS, the most expressive),expresso_ex01…expresso_ex04(Expresso), and 218libritts_<id>voices (LibriTTS-R). The full list istts.voices. - Tags:
none neutralplus emotions (adoration amazement amusement anger confusion contentment cuteness desire disappointment disgust distress ecstasy embarrassment fear guilt happiness interest laughing pain pride realization relief sadness serenity) and styles (emphasis enunciated fast slow loud whisper high_pitch low_pitch). Emotion tags steer best on EARS and Expresso voices.
Evaluation (step 58k of the same run; updated numbers will follow)
| set | metric | this model | attention-only baseline | ground truth |
|---|---|---|---|---|
| LibriTTS-R test-clean (500) | WER whisper-large-v3 % | 2.12 | 2.04 | 1.98 |
| LibriTTS-R test-clean (500) | UTMOS22 | 4.12 | 3.99 | 4.26 |
| ParaSpeechCaps test (246 prompts) | WER distil-whisper-large-v2 % | 3.07 | 3.31 | — |
For reference, Parler-TTS scores 4.62 WER on ParaSpeechCaps under the same ASR. Pitch-level and rate attribute accuracy from text prompts (0.6) is still well below real speech (0.9), and that is the main known gap.
Training data and license
The model was trained on EARS (CC-BY-NC 4.0), Expresso (CC-BY-NC 4.0) and LibriTTS-R subsets (CC-BY 4.0), about 53 h in total. Because of the non-commercial datasets, these weights are released under CC-BY-NC 4.0. Contact Whissle for commercial licensing.
Limitations
- English only, with a g2p_en front end (no explicit normalization of numbers or abbreviations beyond g2p_en's).
- Fixed voice set; no cloning in this checkpoint.
- The prompt adapter understands short style descriptions. It can miss modifiers buried in long prompts (e.g. "bored" → monotone).
- Don't use it to impersonate real people or produce deceptive audio.
- Downloads last month
- 10