Surogate Speech | Training and Serving Engine | Toolkit and Evals |
License: CC-BY-NC-4.0 | Authors: Invergent

Amami 357M (Romanian)

Amami is Surogate's text-to-speech family. This is the Romanian model: a 357M-parameter Magpie-TTS generator with three built-in voices, Doina, Tudor and Radu, that reads Romanian text aloud, numbers, dates and abbreviations included, at 22.05 kHz.

It is served fastest by the surogate engine, which runs the generator and the audio codec natively on CPU or GPU and streams audio as it is synthesized. Every number on this page was measured, and the harnesses are named, so you can reproduce them.

Listen

Doina Tudor Radu

Accuracy

Intelligibility is measured the standard way: Amami reads the sentences, an independent speech recognizer transcribes the audio, and we report the word error rate against the text. Lower is better.

MiniMax multilingual TTS test set, Romanian (100 sentences):

System WER, judge Whisper large-v3 WER, judge Canary 1B v2
Surogate Amami 357M, mean of 3 voices 2.02 2.96
Amami voices Doina / Tudor / Radu 2.65 / 1.75 / 1.66 2.99 / 2.86 / 3.03
Facebook MMS-TTS Romanian (our run) 6.53 5.12
ElevenLabs Multilingual v2 (published, voice cloning) 1.35
MiniMax-Speech (published, voice cloning) 2.88

The published systems clone the test set's two reference speakers; Amami reads with its own fixed voices, so only the word error rate compares, not speaker similarity. Audio and transcripts for every sentence are in surogate-speech-evals.

A harder test, the Surogate stress set (331 sentences of sums, dates, IBANs, legal and banking phrasing), ships in surogate-speech with the harness, to run against any TTS.

Speed

Package Hardware Real-time factor
gpu/ 1× RTX 5090 about 25× faster than real time
macos/ Apple M3, 8 GB, Metal about real time

Running it

With surogate (recommended)

Needs surogate 1.5.4 or newer.

docker pull ghcr.io/invergent-ai/surogate:1.5.4
docker run --gpus all -p 8000:8000 ghcr.io/invergent-ai/surogate:1.5.4 \
  serve --tts surogate/amami-357m-ro --device 0 --host 0.0.0.0 --port 8000

curl http://localhost:8000/v1/audio/speech -H "Content-Type: application/json" \
  -d '{"model": "surogate/amami-357m-ro", "voice": "Doina", "input": "Bună ziua! Cu ce vă pot ajuta?"}' -o doina.wav

The API is OpenAI-compatible and can stream audio as it is generated. Drop --gpus all and --device 0 to run on CPU. Server options are in the engine's TTS docs.

With surogate-speech (local)

pip install "surogate-speech[tts] @ git+https://github.com/invergent-ai/surogate-speech"
surogate-speech speak "Bună ziua! Cu ce vă pot ajuta?" --voice Tudor -o tudor.wav

It picks the native package for your machine, with no PyTorch: macos/ on Apple Silicon, cpu/ on Linux x86-64, and gpu/ with --device gpu on an NVIDIA GPU (--device gpu also turns on Metal on a Mac). Anywhere else, add --backend nemo to run the NeMo package with PyTorch.

Files

folder size what it is
cpu/ 1.2 GB native runtime for Linux x86-64 (x86-64-v3: AVX2, FMA): generator amami-357m-ro.gguf, codec codec.gguf, runtime libraries, voices.json
gpu/ 1.5 GB the same model and voices with the runtime built for CUDA 13, NVIDIA Ampere to Blackwell (A100, A30, L4, L40, RTX 30/40/50, H100, H200, B200)
macos/ 1.2 GB the same model and voices for Apple Silicon, Metal and Accelerate
nemo/ 1.4 GB the NeMo (PyTorch) package: model.ckpt, config.yaml, text-processing assets; the audio codec is NVIDIA's NanoCodec, fetched from its own repository
samples/ 1 MB one sentence per voice

cpu/, gpu/ and macos/ share the generator and codec byte for byte. The generator file holds the 286M parameters used to speak; the full 357M model, including parts used only in training, is in nemo/. voices.json in each folder lists every file's SHA-256, and the engine verifies them before loading. The runtime is NVIDIA NeMo-Speech.cpp at 07003daa with the long-form patch included in each folder.

Limitations

Romanian only, three fixed voices; this release does not clone voices. Custom voices are available from Invergent. Long codes, unusual acronyms and rare foreign names can still be misread, and very long inputs are split into sentences before synthesis. About 1 in 200 two-sentence inputs ends early and drops its second sentence. Word error rate measures intelligibility, not naturalness or prosody. Synthetic speech must not be used to impersonate real people or mislead listeners.

License

CC-BY-NC-4.0: free for research and other non-commercial use; for commercial use, contact Invergent. Built on nvidia/magpie_tts_multilingual_357m and NanoCodec by NVIDIA, which keep their own licenses; the native runtime is NeMo-Speech.cpp (Apache-2.0) with GGML (MIT), whose notices ship in each folder.

Downloads last month
328
GGUF
Model size
0.3B params
Architecture
magpietts
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for surogate/amami-357m-ro

Quantized
(5)
this model

Collection including surogate/amami-357m-ro