Instructions to use coder543/pocket-tts-glade with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use coder543/pocket-tts-glade with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("coder543/pocket-tts-glade") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket TTS for Glade
Specialized conversion of Kyutai Pocket TTS. This profile supports English text with the predefined alba voice and returns mono 24 kHz Float PCM, including incremental audio.
Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac qualification is included; iPhone and Watch qualification is not claimed. Download weights and configuration from one snapshot.
Bundle and execution
The approximately 117.2 MB bundle uses W8A16 language weights and an FP16 Mimi
codec. Both source .aimodel assets prefer ANE. The language graph contains six
causal Transformer layers and the original one-step LSD flow head; each step
produces an 80 ms latent. Mimi retains causal state across codec calls.
The 512-position language cache includes alba's 126-position prefix. Prefill uses 32 positions. The codec shares weights across 1/2/4/8 consecutive-latent entrypoints; its native 250-position acoustic attention context is preserved. Tokenization, sentence splitting, 50-token target, EOS conventions and short fade-in match the source. Unsupported excess capacity fails explicitly instead of dropping history.
metadata.json supplies the runtime geometry and sampling conventions. Tokenizer,
embeddings/projections and alba conditioning arrays are included. Arbitrary voice
cloning is not implemented. No reference recordings or compiler caches are included.
Measured Mac performance
M3 MacBook Air, 16 GB, macOS 27.0.1; warmed Release runtime. The complete 2,201-word reference transcript of JFK's “We choose to go to the Moon” speech uses all 90 original text chunks. Sampling is temperature 0.3, one-step LSD.
| Runtime | Generated audio | Warm synthesis | RTFx |
|---|---|---|---|
| Glade selected W8/FP16 profile | 10 min 39 s | 26.48 s | 24.1× |
| Kyutai original PyTorch CPU, one thread | 10 min 37 s | 78.30 s | 8.1× |
Timings include tokenization, all language/codec calls and host state work; preparation, warmup and WAV writing are excluded. Gaussian PRNG sequences differ, so equal seeds do not imply identical waveforms. A full traced pass contains ANE predictions in all 9,325 neural calls and zero GPU intervals. This establishes placement, not ALU occupancy or energy savings. Complete-text Cohere checks pass; perceptual quality and other hardware remain unqualified.
Observed cached preparation for the short sentence was 0.385 s. An isolated pristine-cache specialization measurement is unavailable for this packaged profile. Authoring source locations are stripped, with graph signatures/statistics unchanged.
Attribution and licenses
Weights: Kyutai, revision 983151f13aaeab1b13c1e5e3c2c383d49a9edf3f, CC-BY-4.0 (LICENSE).
Original code: revision 41cbc84af539ea78a804ffca5f9c6edc1a22ce44, MIT (CODE-LICENSE).
Tokenizer revision: 00eac05ed3d16bdc3f6b5d598874019c34a89214.
Alba conditioning revision: 1e08e6a23401048648a9fdcfde2f89348215c2a7.
See the original model card for its usage conditions. Conversion performs no training.
- Downloads last month
- -
Model tree for coder543/pocket-tts-glade
Base model
kyutai/pocket-tts