Paradee-8M for Glade
Mixed-precision conversion of Paradee-8M v1.0,
a single-voice English distillation of Kokoro. It accepts phonemes and generates
mono 24 kHz Float PCM with its learned af_heart voice.
Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac qualification is included; iPhone and Watch qualification is not claimed. Download weights and configuration from one snapshot.
Bundle and execution
INT8 text/prosody/acoustic weights use FP16 activations on ANE. The GPU waveform generator retains FP16 weights. CPU weights are stored as FP16 and expanded for FP32 recurrence and signal processing. The original 33-frame phase-locking filter remains enabled. A compatible client phonemizer is required.
metadata.json describes the 96-token/192-acoustic-frame capacities, masking,
compute preferences and phase filter. Four source .aimodel files plus
cpu.f16/cpu.json supply the model. The mixed bundle is approximately 11.9 MB.
Over-capacity input reports an error; clients partition long phoneme input at
appropriate pauses. Utterance-wide normalization excludes padding. No compiled
cache, generated audio or private voice sample is included.
Measured Mac performance
M3 MacBook Air, 16 GB, macOS 27.0.1; warmed Release runtime. Input is the complete reference text of JFK's “We choose to go to the Moon” speech.
| Runtime | Warm synthesis | RTFx |
|---|---|---|
| Glade mixed-precision profile | 8.43 s | 95.6× |
| Unchanged FluidAudio default on the same Mac | 47.37 s | 17.0× |
RTFx uses each generated recording's duration. Timings include the text/acoustic runtime and native signal processing, excluding preparation, G2P and WAV writing. Both complete outputs pass a Cohere recognition check against the supplied text; that establishes coverage/coherence for this sample, not general MOS or voice similarity. Matched numerical controls qualify reduced precision separately. Traces show ANE execution in text, prosody and acoustic calls, with GPU generation.
Source assets specialize on the device. No isolated pristine-cache specialization time is claimed for the packaged mixed profile. Source locations are stripped without changing graph signatures or operation statistics. No AoT cache is supplied.
Attribution
Weights revision: f662642d44c03c17588e4176469c54d462c0b623.
Original implementation revision: 9c8b4d7504cbee7e64de2d0341bb690f0b0ab708.
Original model/source use Apache-2.0; see LICENSE and the upstream card.
Conversion performs no additional training.
Model tree for coder543/paradee-8m-glade
Base model
yl4579/StyleTTS2-LJSpeech