Instructions to use florianvoss/supertonic-3-sima with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Supertonic
How to use florianvoss/supertonic-3-sima with Supertonic:
from supertonic import TTS tts = TTS(auto_download=True) style = tts.get_voice_style(voice_name="M1") text = "The train delay was announced at 4:45 PM on Wed, Apr 3, 2024 due to track maintenance." wav, duration = tts.synthesize(text, voice_style=style) tts.save_audio(wav, "output.wav")
- Notebooks
- Google Colab
- Kaggle
Supertonic 3 for SiMa Modalix
This repository contains SiMa Modalix-compiled vector-field and vocoder models for Supertone/Supertonic 3. The duration predictor, text encoder, and text processing remain host CPU components. The complete graph-surgery and verification source is available at florianvoss-commit/supertonic-sima.
Files
| File | Purpose | SHA-256 |
|---|---|---|
supertonic_vector_field_sima_mpk.tar.gz |
Batch-1 compiled vector-field MPK | 2f6b8c918e0c402453e48bd2686dbea429e6ce1dd98151c940d88229980e8dd2 |
supertonic_runtime_data.npz |
Timestep rows, length-indexed RoPE bank, and CFG constants | 81fc7a131dd0eafe6fa7062b23f49e0dc9e3fb178f5473bf8d404ab555bcf5aa |
artifact_manifest.json |
Batch-1 compiler provenance, sizes, hashes, and validation metrics | See file |
supertonic_vocoder_sima_bf16_mpk.tar.gz |
Full-BF16 batch-1 vocoder MPK | 90b2d6a089c8527826dd1d0cb5b557316ac703045f422d6e1332deeabd84e0cb |
vocoder_bf16_manifest.json |
Vocoder graph, compiler, and accuracy provenance | See file |
validation/vocoder_bfloat16_vs_onnx.json |
Real-latent BF16-versus-ONNX metrics | 797adefed9e36b2ab403655cca53fef2f3cdf25b59f0e43b57aa30be86c95129 |
Each tarball contains its model MPK JSON and MLA ELF. The application passes the tarball to the SiMa model runtime; the ELF does not need to be downloaded separately.
Compilation profile
- Target: SiMa Modalix
- Model Compiler:
2.1.3 - Batch size:
1 - Text and latent profile:
192 - Activations: BF16 internally
- Vector-field weights: symmetric per-channel INT8
- Vocoder weights: BF16
- Application boundary: FP32
- Calibration: compiler-generated random data
- MLA tessellation: enabled
- Model placement: exactly one
MLA_0compute partition - EV74: eight input casts, one input pack transform, and one output cast
- A65/APU model partitions: zero
The vocoder also compiles into exactly one MLA_0 compute partition. Its only
EV74 plugins are the expected FP32-to-BF16 input cast and BF16-to-FP32 output
cast. Its compiler cycle estimate is 11,625,345.
Download
conda activate hf
hf download florianvoss/supertonic-3-sima \
supertonic_vector_field_sima_mpk.tar.gz \
supertonic_vocoder_sima_bf16_mpk.tar.gz \
supertonic_runtime_data.npz \
artifact_manifest.json \
vocoder_bf16_manifest.json \
--local-dir models/supertonic-3-sima
The remaining CPU components come from the pinned upstream revision:
hf download Supertone/supertonic-3 \
--revision 724fb5abbf5502583fb520898d45929e62f02c0b \
--local-dir models/supertonic-3
The pinned upstream onnx/vector_estimator.onnx SHA-256 is
883ac868ea0275ef0e991524dc64f16b3c0376efd7c320af6b53f5b780d7c61c.
The pinned upstream onnx/vocoder.onnx SHA-256 is
085de76dd8e8d5836d6ca66826601f615939218f90e519f70ee8a36ed2a4c4ba.
Compiled tensor contract
Inputs are ordered, contiguous FP32 tensors. The compiled vector estimator has a fixed batch size of 1.
| Index | Name | Shape |
|---|---|---|
| 0 | noisy_latent |
[1, 144, 1, 192] |
| 1 | text_emb |
[1, 256, 1, 192] |
| 2 | style_ttl |
[1, 256, 1, 50] |
| 3 | style_key |
[1, 256, 1, 50] |
| 4 | latent_mask |
[1, 1, 1, 192] |
| 5 | text_mask |
[1, 192, 1, 1] |
| 6 | time_sinusoidal |
[1, 64, 1, 1] |
| 7 | rope_tables |
[1, 128, 1, 192] |
Output: velocity, shape [1, 144, 1, 192], FP32.
The runtime NPZ contains eight time_sinusoidal rows and a RoPE bank indexed
by effective sequence length from 1 through 192. For each invocation, pack the
RoPE input in this channel order:
latent_sin, latent_cos, text_sin, text_cos
Compiled vocoder contract
The vocoder uses contiguous FP32 application buffers and full BF16 internally:
| Direction | Name | Logical shape | Public MPK shape |
|---|---|---|---|
| Input | latent |
[1, 144, 1, 192] |
[1, 1, 192, 144] |
| Output | wav_frames |
[1, 512, 1, 1152] |
[1, 1, 1152, 512] |
The public output is already time-major: 1,152 frames with 512 consecutive samples per frame. Read the returned FP32 bytes directly as a flat 589,824 sample waveform; do not transpose them. Crop to the natural latent length times 3,072 samples.
Denoising loop
Invoke the batch-1 graph twice per denoising step: once with conditional text/style context and once with unconditional context. Reuse the same noisy latent, masks, timestep row, and RoPE table for both invocations. Apply guidance and the Euler update on the host in FP32:
latent = (latent + (4 * conditional - 3 * unconditional) / total_steps) * latent_mask
After eight steps, zero-pad the natural latent to 192 positions and invoke the
compiled BF16 vocoder once. Crop its time-major output to
natural_latent_length * 3072 samples.
This requires 16 vector-estimator MLA submissions for the full eight-step loop.
Processed text and predicted latent length must each be at most 192. This latent profile represents at most approximately 13.37 seconds at 44.1 kHz. Live applications should split at sentence boundaries and use clause boundaries for unusually long sentences.
Validation
Graph surgery was validated over five reference cases, eight steps, and both CFG branches: 80 comparisons with zero observed difference from the pre-externalized ONNX graph.
The real production loop was also evaluated with the exact saved quantized network:
| Case | Audio | Step-1 cosine | Final-latent cosine | Waveform cosine |
|---|---|---|---|---|
| Short English | 4.14 s | 0.999926 | 0.979563 | 0.835558 |
| Near-capacity English | 12.66 s | 0.999931 | 0.968573 | 0.479883 |
The sample-aligned waveform metrics strongly penalize small timing and phase
differences. Listening samples and the complete reports are included under
samples/ and validation/.
The static vocoder surgery itself is near-exact against the released dynamic
ONNX model: minimum cosine 0.999999999968 and maximum relative L2
8.03e-6. Full-BF16 AFE execution gives:
| Case | Natural latent | Waveform cosine | Relative L2 | Mean absolute error |
|---|---|---|---|---|
| Short English | 60 | 0.992113 | 0.125363 | 0.002408 |
| Near-capacity English | 182 | 0.995686 | 0.093048 | 0.002207 |
INT8 vocoder weights are not published because their error changes a sparse
positive activation mask before a PReLU with slope 0.0001946, degrading
waveform cosine to 0.420–0.444.
Status
Quantization, one-segment MLA compilation, ONNX production-loop verification, and saved-AFE-model verification are complete for the batch-1 vector field and BF16 vocoder.
License and attribution
Supertonic 3 model weights are distributed under the OpenRAIL-M license. This repository includes the upstream license and identifies the original model and pinned revision above.
Model tree for florianvoss/supertonic-3-sima
Base model
Supertone/supertonic-3