Supertonic 3 for Core ML (Rytability)

A Core ML conversion of Supertonic 3 by Supertone Inc., prepared for on-device Read Aloud in the Rytability app on iPhone, iPad and Mac.

The original model, its weights and the voice styles are the work of Supertone Inc. and are distributed here under the original BigScience OpenRAIL-M License (see LICENSE), including its use-based restrictions (Attachment A), which apply to this conversion and to anything made with it.

What is in this repository

Path What
DurationPredictor.mlmodelc duration_predictor.onnx, converted
TextEncoder.mlmodelc text_encoder.onnx, converted
VectorEstimator.mlmodelc vector_estimator.onnx, converted, plus one extra output (below)
Vocoder.mlmodelc vocoder.onnx, converted
config/tts.json, config/unicode_indexer.json unchanged from the original onnx/ folder
voice_styles/*.json the ten original preset voices (F1-F5, M1-M5), unchanged
LICENSE the original OpenRAIL-M license, unchanged
NOTICE.md attribution and the list of modifications

Total download: about 381 MB.

How it was converted

The four ONNX graphs were translated directly into Core ML ML Program (MIL) operations (iOS 18 / macOS 15, float32), keeping the text length and latent length dynamic, so the models run at the true input lengths with no padding. (Supertonic's text encoder and vector estimator are not padding-invariant, so fixed padded shapes would have changed the output.)

Validation against ONNX Runtime on the original graphs:

  • every stage matches to a relative error of at most 3e-6, at text lengths 12 to 159 and latent lengths 19 to 145;
  • the Core ML CPU path runs every text length from 8 to 600 characters and every latent length from 1 to 600 frames;
  • speech intelligibility (Whisper large-v3-turbo, 8 sentences in 8 languages): CER 0.005.

Run the models with MLComputeUnits.cpuOnly. The dynamic shapes are not supported on the Neural Engine, and CPU execution keeps generation working while an iOS app is in the background.

The extra alignment output

VectorEstimator returns one additional tensor, alignment [1, L, T]: a weighted mean of four of the model's own text cross-attention heads (block 15 head 5 at 0.4; block 9 head 6, block 21 heads 2 and 7 at 0.2 each). These heads follow the text monotonically in every voice and language tested. A monotonic Viterbi path through the final step's map gives each character a latent frame (3072 samples, about 69.7 ms at 44.1 kHz). Against an independent forced aligner (torchaudio MMS_FA) the attention leads the acoustic word onset by a constant ~77 ms; with that offset applied, word start times agree to 15 ms mean absolute error (23 ms at the 90th percentile) over 102 words in English, Spanish, French and German. The audio output is unaffected.

Pipeline

Identical to the official helper.py: preprocess the text and wrap it in <lang>...</lang> (or <na>...</na> for language-agnostic reading), map characters through unicode_indexer.json, run the duration predictor and text encoder, draw Gaussian noise of shape [1, 144, ceil(duration * 44100 / 3072)], run the vector estimator for total_step (8) steps, then the vocoder. Output is 44.1 kHz mono.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rybib/rytability-supertonic3-coreml

Finetuned
(13)
this model