Amami, the voice of Surogate Speech 🗣️
Amami does the talking in Surogate Speech, with voices per language, and Romanian goes first. It reads the MiniMax Romanian test set at 2.02% word error rate, judged by Whisper large-v3. It runs on a plain CPU, no GPU and no PyTorch needed.
amami-357m-ro is a 357M Magpie-TTS generator adapted to Romanian, with three voices built in: Doina, Tudor and Radu. It reads sums, dates, IBANs and abbreviations the way a person would say them.
It comes four ways:
cpu/: native runtime for Linux x86-64, any CPU with AVX2gpu/: the same runtime for NVIDIA GPUs from Ampere to Blackwell (A100, L4, RTX 30/40/50, H100, B200), about 25× real time on an RTX 5090macos/: Apple Silicon with Metal, about real time on an 8 GB Apple M3nemo/: the PyTorch package, for anywhere else
Why this one
An agent that talks to people has to be understood first. A misread sum or a wrong date is a failed call, and it doesn't matter how warm the voice sounds. It also has to start talking fast, and it has to run on whatever box the rest of the agent runs on, which usually has no GPU.
How intelligible
Amami reads the sentence, an independent recogniser writes down what it heard, and we count the wrong words.
MiniMax multilingual test set, Romanian, 100 sentences, judge Whisper large-v3:
| System | WER |
|---|---|
| Surogate Amami 357M, mean of 3 voices | 2.02% |
| Doina / Tudor / Radu | 2.65 / 1.75 / 1.66% |
| Facebook MMS-TTS Romanian | 6.53% |
| ElevenLabs Multilingual v2 (their paper, voice cloning) | 1.35% |
| MiniMax-Speech (their paper, voice cloning) | 2.88% |
The published systems clone the set's two reference speakers. Amami reads with its own voices, so the WER column compares and a speaker-similarity column wouldn't.
We also built a stress set: 331 sentences written to trip TTS up, with sums, dates, IBANs, articles of law, abbreviations and English loanwords. It ships in surogate-speech with the harness, so you can run it against Amami or any other Romanian TTS.
Run it
surogate serve --tts surogate/amami-357m-ro --device 0
curl http://localhost:8000/v1/audio/speech -H "Content-Type: application/json" \
-d '{"model": "surogate/amami-357m-ro", "voice": "Doina", "input": "Bună ziua! Cu ce vă pot ajuta?"}' -o doina.wav
pip install "surogate-speech[tts] @ git+https://github.com/invergent-ai/surogate-speech"
surogate-speech speak "Plata de 327,45 lei a fost programată pentru 3 octombrie." --voice Tudor -o tudor.wav
Three fixed voices, no voice cloning in this release; custom voices come from Invergent. About 1 in 200 two-sentence inputs drops its second sentence, so check long outputs. WER tells you it's understood, not that it sounds natural, so listen before you ship it. Don't use synthetic speech to pass as a real person.
Weights are CC-BY-NC-4.0, commercial licensing through Invergent. Built on NVIDIA's Magpie-TTS Multilingual 357M and NanoCodec, which keep their own licenses. Raw outputs: surogate-speech-evals.

