Kokoro-82M for Glade
Native-context conversion of Kokoro-82M v1.0.
Glade accepts English text or explicit phonemes and generates mono 24 kHz Float PCM. The bundle includes
af_heart and am_adam; measurements use af_heart at speed 1.
Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac and iPhone qualification are included; Watch qualification is not claimed. Download weights and configuration from one snapshot.
Bundle and execution
The profile retains up to 510 phoneme symbols, with 512 text positions including boundaries and 1,280 acoustic frames (32 seconds). It preserves full-utterance normalization with padding excluded. Over-capacity input reports an error rather than truncating or independently windowing an utterance.
Text runs on ANE; prosody/acoustic graphs and waveform generation use GPU. Native
CPU code handles bidirectional recurrence, harmonic synthesis and FFT/overlap-add.
Glade provides native American English pronunciation lookup and text planning.
Explicit phonemes remain available for controlled comparisons. Runtime preferences, capacities,
vocabulary and voice inventory are in metadata.json; CPU weights/layout are
cpu.f32/cpu.json. There are four source .aimodel assets and two voice tables.
No compiled specialization cache or generated audio is included.
Measured Mac performance
M3 MacBook Air, 16 GB, macOS 27.0.1; warmed Release runtime. The complete text of JFK's “We choose to go to the Moon” speech uses the same 41 original Kokoro/Misaki chunks, up to 509 phonemes, in both implementations.
| Runtime | Generated audio | Synthesis | RTFx |
|---|---|---|---|
| Glade native-context profile | 12 min 13 s | 34.44 s | 21.3× |
| MLX-Audio FP16, same native chunk plan | 12 min 13 s | 53.37 s | 13.7× |
RTFx is generated audio divided by synthesis time. G2P, preparation and WAV writing are excluded; recurrence, signal work, cache cleanup and neural calls are included. Sampled simultaneous client footprint plus neural accounting peaked at 1.31 GiB. Only one Mac/input/voice is represented. Source RNG distributions match, but RNG sequences differ. Numerical controls and full-text recognition checks do not establish general perceptual or voice-similarity equivalence.
Source assets require device specialization. No isolated pristine-cache preparation measurement is available for this packaged profile. Packaging removes source-location metadata and checks unchanged graph signatures and operation statistics; it does not distribute an AoT cache or imply an all-ANE vocoder.
Measured iPhone performance
iPhone 15 Pro Max (A17 Pro, 8 GB),iOS 27.0.1; Release, the same complete Moon-speech phoneme plan, voice and seed policy as the Mac control above. Three measured passes generate 12 min 13 s each. The median is 93.01 s / 7.9× RTFx; the first pass is 68.39 s / 10.7×. Thermal state rises from nominal to serious across these sustained runs, so 7.9× is a sustained result, not a nominal-temperature median. The GPU vocoder accounts for most of the runtime.
A separate device trace confirms ANE text execution and GPU prosody, acoustic and waveform stages, matching the bundle configuration. Initial observed preparation is 5.56 s, with a later cached load around 0.32 s; underlying compiler caches were not purged. Sampled client peak is 1095 MiB, including about 70 MB of benchmark PCM. Client/Metal ledgers exclude unattributed kernel and compiler/service resources. Repeated phone PCM is byte-identical, and a full-output recognition check confirms speech coverage for this sample; this does not establish general perceptual or voice-similarity equivalence.
Attribution
Upstream weights/source revision: f3ff3571791e39611d31c381e3a41a3af07b4987.
Weights and original Kokoro code use Apache-2.0; see LICENSE and the upstream card.
Kokoro acknowledges StyleTTS2 and iSTFTNet. No additional training is performed.