Pocket TTS for Glade

Specialized conversion of Kyutai Pocket TTS. This profile supports English text with the predefined alba voice and returns mono 24 kHz Float PCM, including incremental audio.

Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac qualification is included; iPhone and Watch qualification is not claimed. Download weights and configuration from one snapshot.

Bundle and execution

The approximately 117.2 MB bundle uses W8A16 language weights and an FP16 Mimi codec. Both source .aimodel assets prefer ANE. The language graph contains six causal Transformer layers and the original one-step LSD flow head; each step produces an 80 ms latent. Mimi retains causal state across codec calls.

The 512-position language cache includes alba's 126-position prefix. Prefill uses 32 positions. The codec shares weights across 1/2/4/8 consecutive-latent entrypoints; its native 250-position acoustic attention context is preserved. Tokenization, sentence splitting, 50-token target, EOS conventions and short fade-in match the source. Unsupported excess capacity fails explicitly instead of dropping history.

metadata.json supplies the runtime geometry and sampling conventions. Tokenizer, embeddings/projections and alba conditioning arrays are included. Arbitrary voice cloning is not implemented. No reference recordings or compiler caches are included.

Measured Mac performance

M3 MacBook Air, 16 GB, macOS 27.0.1; warmed Release runtime. The complete 2,201-word reference transcript of JFK's “We choose to go to the Moon” speech uses all 90 original text chunks. Sampling is temperature 0.3, one-step LSD.

Runtime Generated audio Warm synthesis RTFx
Glade selected W8/FP16 profile 10 min 39 s 26.48 s 24.1×
Kyutai original PyTorch CPU, one thread 10 min 37 s 78.30 s 8.1×

Timings include tokenization, all language/codec calls and host state work; preparation, warmup and WAV writing are excluded. Gaussian PRNG sequences differ, so equal seeds do not imply identical waveforms. A full traced pass contains ANE predictions in all 9,325 neural calls and zero GPU intervals. This establishes placement, not ALU occupancy or energy savings. Complete-text Cohere checks pass; perceptual quality and other hardware remain unqualified.

Observed cached preparation for the short sentence was 0.385 s. An isolated pristine-cache specialization measurement is unavailable for this packaged profile. Authoring source locations are stripped, with graph signatures/statistics unchanged.

Attribution and licenses

Weights: Kyutai, revision 983151f13aaeab1b13c1e5e3c2c383d49a9edf3f, CC-BY-4.0 (LICENSE). Original code: revision 41cbc84af539ea78a804ffca5f9c6edc1a22ce44, MIT (CODE-LICENSE). Tokenizer revision: 00eac05ed3d16bdc3f6b5d598874019c34a89214. Alba conditioning revision: 1e08e6a23401048648a9fdcfde2f89348215c2a7. See the original model card for its usage conditions. Conversion performs no training.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/pocket-tts-glade

Quantized
(52)
this model

Collection including coder543/pocket-tts-glade