Supertonic 3 for Core ML (Rytability)
A Core ML conversion of Supertonic 3 by Supertone Inc., prepared for on-device Read Aloud in the Rytability app on iPhone, iPad and Mac.
The original model, its weights and the voice styles are the work of Supertone Inc. and are
distributed here under the original BigScience OpenRAIL-M License (see LICENSE), including
its use-based restrictions (Attachment A), which apply to this conversion and to anything made with it.
- Original model: https://huggingface.co/Supertone/supertonic-3 (revision
3cadd1ee6394adea1bd021217a0e650ede09a323) - Archive mirror: https://huggingface.co/supertone-oss-archive/supertonic-3 (revision
aafc6e32416a594460b32413efc49d7fe4ce6d46) - Source code: https://github.com/supertone-oss-archive/supertonic (MIT)
What is in this repository
| Path | What |
|---|---|
DurationPredictor.mlmodelc |
duration_predictor.onnx, converted |
TextEncoder.mlmodelc |
text_encoder.onnx, converted |
VectorEstimator.mlmodelc |
vector_estimator.onnx, converted, plus one extra output (below) |
Vocoder.mlmodelc |
vocoder.onnx, converted |
config/tts.json, config/unicode_indexer.json |
unchanged from the original onnx/ folder |
voice_styles/*.json |
the ten original preset voices (F1-F5, M1-M5), unchanged |
LICENSE |
the original OpenRAIL-M license, unchanged |
NOTICE.md |
attribution and the list of modifications |
Total download: about 381 MB.
How it was converted
The four ONNX graphs were translated directly into Core ML ML Program (MIL) operations (iOS 18 / macOS 15, float32), keeping the text length and latent length dynamic, so the models run at the true input lengths with no padding. (Supertonic's text encoder and vector estimator are not padding-invariant, so fixed padded shapes would have changed the output.)
Validation against ONNX Runtime on the original graphs:
- every stage matches to a relative error of at most 3e-6, at text lengths 12 to 159 and latent lengths 19 to 145;
- the Core ML CPU path runs every text length from 8 to 600 characters and every latent length from 1 to 600 frames;
- speech intelligibility (Whisper large-v3-turbo, 8 sentences in 8 languages): CER 0.005.
Run the models with MLComputeUnits.cpuOnly. The dynamic shapes are not supported on the Neural
Engine, and CPU execution keeps generation working while an iOS app is in the background.
The extra alignment output
VectorEstimator returns one additional tensor, alignment [1, L, T]: a weighted mean of four of
the model's own text cross-attention heads (block 15 head 5 at 0.4; block 9 head 6, block 21 heads 2
and 7 at 0.2 each). These heads follow the text monotonically in every voice and language tested.
A monotonic Viterbi path through the final step's map gives each character a latent frame
(3072 samples, about 69.7 ms at 44.1 kHz). Against an independent forced aligner (torchaudio MMS_FA)
the attention leads the acoustic word onset by a constant ~77 ms; with that offset applied, word start
times agree to 15 ms mean absolute error (23 ms at the 90th percentile) over 102 words in English,
Spanish, French and German. The audio output is unaffected.
Pipeline
Identical to the official helper.py: preprocess the text and wrap it in <lang>...</lang>
(or <na>...</na> for language-agnostic reading), map characters through unicode_indexer.json,
run the duration predictor and text encoder, draw Gaussian noise of shape [1, 144, ceil(duration * 44100 / 3072)],
run the vector estimator for total_step (8) steps, then the vocoder. Output is 44.1 kHz mono.
Model tree for Rybib/rytability-supertonic3-coreml
Base model
Supertone/supertonic-3