Supertonic 3 for SiMa Modalix

This repository contains SiMa Modalix-compiled vector-field and vocoder models for Supertone/Supertonic 3. The duration predictor, text encoder, and text processing remain host CPU components. The complete graph-surgery and verification source is available at florianvoss-commit/supertonic-sima.

Files

File Purpose SHA-256
supertonic_vector_field_sima_mpk.tar.gz Batch-1 compiled vector-field MPK 2f6b8c918e0c402453e48bd2686dbea429e6ce1dd98151c940d88229980e8dd2
supertonic_runtime_data.npz Timestep rows, length-indexed RoPE bank, and CFG constants 81fc7a131dd0eafe6fa7062b23f49e0dc9e3fb178f5473bf8d404ab555bcf5aa
artifact_manifest.json Batch-1 compiler provenance, sizes, hashes, and validation metrics See file
supertonic_vocoder_sima_bf16_mpk.tar.gz Full-BF16 batch-1 vocoder MPK 90b2d6a089c8527826dd1d0cb5b557316ac703045f422d6e1332deeabd84e0cb
vocoder_bf16_manifest.json Vocoder graph, compiler, and accuracy provenance See file
validation/vocoder_bfloat16_vs_onnx.json Real-latent BF16-versus-ONNX metrics 797adefed9e36b2ab403655cca53fef2f3cdf25b59f0e43b57aa30be86c95129

Each tarball contains its model MPK JSON and MLA ELF. The application passes the tarball to the SiMa model runtime; the ELF does not need to be downloaded separately.

Compilation profile

  • Target: SiMa Modalix
  • Model Compiler: 2.1.3
  • Batch size: 1
  • Text and latent profile: 192
  • Activations: BF16 internally
  • Vector-field weights: symmetric per-channel INT8
  • Vocoder weights: BF16
  • Application boundary: FP32
  • Calibration: compiler-generated random data
  • MLA tessellation: enabled
  • Model placement: exactly one MLA_0 compute partition
  • EV74: eight input casts, one input pack transform, and one output cast
  • A65/APU model partitions: zero

The vocoder also compiles into exactly one MLA_0 compute partition. Its only EV74 plugins are the expected FP32-to-BF16 input cast and BF16-to-FP32 output cast. Its compiler cycle estimate is 11,625,345.

Download

conda activate hf
hf download florianvoss/supertonic-3-sima \
  supertonic_vector_field_sima_mpk.tar.gz \
  supertonic_vocoder_sima_bf16_mpk.tar.gz \
  supertonic_runtime_data.npz \
  artifact_manifest.json \
  vocoder_bf16_manifest.json \
  --local-dir models/supertonic-3-sima

The remaining CPU components come from the pinned upstream revision:

hf download Supertone/supertonic-3 \
  --revision 724fb5abbf5502583fb520898d45929e62f02c0b \
  --local-dir models/supertonic-3

The pinned upstream onnx/vector_estimator.onnx SHA-256 is 883ac868ea0275ef0e991524dc64f16b3c0376efd7c320af6b53f5b780d7c61c. The pinned upstream onnx/vocoder.onnx SHA-256 is 085de76dd8e8d5836d6ca66826601f615939218f90e519f70ee8a36ed2a4c4ba.

Compiled tensor contract

Inputs are ordered, contiguous FP32 tensors. The compiled vector estimator has a fixed batch size of 1.

Index Name Shape
0 noisy_latent [1, 144, 1, 192]
1 text_emb [1, 256, 1, 192]
2 style_ttl [1, 256, 1, 50]
3 style_key [1, 256, 1, 50]
4 latent_mask [1, 1, 1, 192]
5 text_mask [1, 192, 1, 1]
6 time_sinusoidal [1, 64, 1, 1]
7 rope_tables [1, 128, 1, 192]

Output: velocity, shape [1, 144, 1, 192], FP32.

The runtime NPZ contains eight time_sinusoidal rows and a RoPE bank indexed by effective sequence length from 1 through 192. For each invocation, pack the RoPE input in this channel order:

latent_sin, latent_cos, text_sin, text_cos

Compiled vocoder contract

The vocoder uses contiguous FP32 application buffers and full BF16 internally:

Direction Name Logical shape Public MPK shape
Input latent [1, 144, 1, 192] [1, 1, 192, 144]
Output wav_frames [1, 512, 1, 1152] [1, 1, 1152, 512]

The public output is already time-major: 1,152 frames with 512 consecutive samples per frame. Read the returned FP32 bytes directly as a flat 589,824 sample waveform; do not transpose them. Crop to the natural latent length times 3,072 samples.

Denoising loop

Invoke the batch-1 graph twice per denoising step: once with conditional text/style context and once with unconditional context. Reuse the same noisy latent, masks, timestep row, and RoPE table for both invocations. Apply guidance and the Euler update on the host in FP32:

latent = (latent + (4 * conditional - 3 * unconditional) / total_steps) * latent_mask

After eight steps, zero-pad the natural latent to 192 positions and invoke the compiled BF16 vocoder once. Crop its time-major output to natural_latent_length * 3072 samples.

This requires 16 vector-estimator MLA submissions for the full eight-step loop.

Processed text and predicted latent length must each be at most 192. This latent profile represents at most approximately 13.37 seconds at 44.1 kHz. Live applications should split at sentence boundaries and use clause boundaries for unusually long sentences.

Validation

Graph surgery was validated over five reference cases, eight steps, and both CFG branches: 80 comparisons with zero observed difference from the pre-externalized ONNX graph.

The real production loop was also evaluated with the exact saved quantized network:

Case Audio Step-1 cosine Final-latent cosine Waveform cosine
Short English 4.14 s 0.999926 0.979563 0.835558
Near-capacity English 12.66 s 0.999931 0.968573 0.479883

The sample-aligned waveform metrics strongly penalize small timing and phase differences. Listening samples and the complete reports are included under samples/ and validation/.

The static vocoder surgery itself is near-exact against the released dynamic ONNX model: minimum cosine 0.999999999968 and maximum relative L2 8.03e-6. Full-BF16 AFE execution gives:

Case Natural latent Waveform cosine Relative L2 Mean absolute error
Short English 60 0.992113 0.125363 0.002408
Near-capacity English 182 0.995686 0.093048 0.002207

INT8 vocoder weights are not published because their error changes a sparse positive activation mask before a PReLU with slope 0.0001946, degrading waveform cosine to 0.420–0.444.

Status

Quantization, one-segment MLA compilation, ONNX production-loop verification, and saved-AFE-model verification are complete for the batch-1 vector field and BF16 vocoder.

License and attribution

Supertonic 3 model weights are distributed under the OpenRAIL-M license. This repository includes the upstream license and identifies the original model and pinned revision above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for florianvoss/supertonic-3-sima

Finetuned
(13)
this model