Paradee-8M Core ML

Core ML conversion of Paradee-8M v1.0 by Sahil Mahendrakar: Kokoro-82M distilled into an 8.07M-parameter, single-voice (af_heart) English TTS model, 24 kHz. Paper: arXiv 2610.06817.

Runs ~150× real time on the CPU of an M5 Pro (93 s of audio in 0.62 s, text side + acoustic side), against 20× for the upstream fp32 ONNX on one onnxruntime thread. The int8 build has a 12 MB weight footprint.

Used by FluidAudio (ParadeeManager).

Files

Path What
int8/ Default. ParadeeText.mlmodelc (3.1 MB) + ParadeeAcoustic.mlmodelc (9.1 MB), int8 per-channel weights; the phase-lock filter's DFT bases stay fp32
fp32/ Same graphs in fp32 (12 + 22 MB). Matches PyTorch to rounding
*/vocab.json Phoneme → id map (identical to Kokoro's)
mlpackage/ Source .mlpackages for both variants
config.json Vocab, sample rate, shape limits
samples/ The model card's five held-out sentences, rendered by int8/

Pipeline

text -> misaki-style en-US phonemes -> ids = [0, ...vocab ids..., 0]       (<= 512 total)
ParadeeText      input_ids [1,T] int32 -> duration [1,T], d [1,224,T], asr_tok [1,512,T]
host             n_i = max(1, round(duration_i / speed));  F = sum(n)
                 en  = d       with column i repeated n_i times   [1,224,F]
                 asr = asr_tok with column i repeated n_i times   [1,512,F]
                 noise ~ N(0,1)                                    [1,1,600F]
ParadeeAcoustic  en, asr, noise -> audio [1,600F] float32, 24 kHz
  • F may be 1…4000 (100 s at speed 1). Split longer input into chunks of ≤ 510 phonemes.
  • noise is the harmonic source's Gaussian noise, passed in so the graph is deterministic. It is equivalent to the upstream in-graph noise (9 per-harmonic channels mixed by a linear layer = one scaled channel).
  • Phonemes must be written the way misaki writes them (Kokoro's G2P); that is all Paradee saw in training. That includes misaki's last step, flap ɾ → T and glottal stop Ê” → t: lexicon-only frontends that skip it garble words like "kittens" and "satellite".
  • Use CPU_ONLY or CPU_AND_NE. Do not use ALL / CPU_AND_GPU: the LSTMs abort in MPSGraph (GPURNNOps … JIT not supported), the same failure Kokoro's prosody stage has.

Accuracy (M5 Pro, macOS 27)

10 inputs: the README quick-start line, the five held-out sentences, a numbers/abbreviation sentence, a 28.6 s paragraph, and one sentence at speed 0.8 and 1.3.

  • fp32 vs PyTorch with the same noise: 0 duration rounding differences, identical lengths. Log-mel L1 against the upstream ONNX equals the ONNX's own run-to-run difference in every case (e.g. 0.123 vs 0.123).
  • Whisper large-v3-turbo transcripts of Core ML fp32 and upstream ONNX fp32 are identical (WER 3.81 %; every miss is text normalization: "Ia", "favourite", "Zzyzx").
  • int8: log-mel L1 to fp32 ONNX 0.18–0.20 vs 0.21–0.31 for upstream paradee_int8.onnx; WER 3.81 % vs 2.97 % (one word of ~236).

FluidAudio tts-benchmark --corpus minimax-english (100 phrases, Parakeet TDT round trip, Swift, CPU, includes the English G2P frontend): WER 1.20 % / CER 0.14 % for both int8 and fp32, RTFx ~100, 77 ms median per phrase.

Conversion notes: the 1024-point phase-lock STFT is a gather + matmul + overlap-add instead of strided convolutions (the conv form ran ~40× slower on the Core ML CPU), and kokoro's random harmonic start phases are omitted because the upstream graph never reads them.

License

Apache-2.0, same as Paradee and Kokoro-82M. Paradee © Sahil Mahendrakar.

@misc{mahendrakar2026paradee,
  title  = {Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model},
  author = {Mahendrakar, Sahil},
  year   = {2026},
  url    = {https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0}
}
Downloads last month
41
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/paradee-8m-coreml

Quantized
(1)
this model

Paper for FluidInference/paradee-8m-coreml