Instructions to use coder543/moss-tts-v1.5-glade with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use coder543/moss-tts-v1.5-glade with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download coder543/moss-tts-v1.5-glade --local-dir moss-tts-v1.5-glade
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
MOSS-TTS-v1.5 for Glade
Hybrid conversion of MOSS-TTS-v1.5 and MOSS Audio Tokenizer. The 8.49B autoregressive language model runs on GPU through native Swift MLX; seventeen Core AI codec groups render mono 24 kHz PCM on ANE. English controls are qualified; other upstream languages are not qualified here.
Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac qualification is included; iPhone and Watch qualification is not claimed. Download weights and configuration from one snapshot.
Layout and downloads
language/: mapped quantized language weights, tensor layout, configuration and tokenizer. Four-bit affine matrix groups, eight-bit audio heads and FP16 audio embeddings preserve the selected quantized runtime. No further training.codec/: seventeen FP16.aimodelgroups plus original FP32 codebooks and boundary projections inboundary.f32/boundary.json.voice-encoder/: optional original FP32 reference-audio encoder for importing client-supplied voices. Encoded conditioning can be cached and reused. Download this only when reference voice cloning is needed. No voice recording is included.
hf download coder543/moss-tts-v1.5-glade \
--include 'language/*' 'codec/*' LICENSE README.md --local-dir ./moss
Glade's optional GladeMOSS product consumes language/ and codec/ directories.
It supplies text cleanup, prompt construction, dynamic continuous batching,
ordered incremental output, optional phase switching and MFA-assisted runout
recovery. The planner targets five minutes per text unit with a hard limit of
5,000 generated positions; prompt positions are additional. Completed lanes
leave the actual GPU batch. Memory budgets can reduce admitted lanes.
Mapped weights avoid repeated loading copies; clean pages are reclaimable, but mapping does not make the approximately 4.69 GiB of language weights disappear. The app must keep a mapped file unchanged while loaded. Generation retains each unit's full language history; codec blocks preserve their causal acoustic caches. Reference voice import can load the encoder before the language model and retain only encoded conditioning afterward. Neither path requires Python.
Measured Mac controls
M3 MacBook Air, 16 GB, macOS 27.0.1; Release runtime. These are bounded English functional controls, not a full Moon-speech benchmark or general perceptual study.
- A longer GPU-language/ANE-codec control generated 48.40 s of audio in 37.91 s (1.28×), excluding preparation and WAV writing. Its language frames exactly matched the separate GPU-codec control.
- Two complete 197-word units produced 157.36 s of audio. Source-ordered online phased and resident rendering produced identical PCM and language tokens.
- A deliberately forced runout retained verified audio through a sentence at 25.63 s after 20.16 s was delivered, then generated only the unspoken suffix. The other batched sequence remained intact. Cohere recovered all 394 normalized words from both this output and the ordinary control.
- The original FP32 codec comparison on a 78.48-second token stream differs by 5.60% relative waveform RMS. Streaming/phase policies match each other exactly; the converted codec is not numerically lossless against FP32.
Initial observed preparation of the seventeen codec assets took 30.51 s; subsequent cached loading took 0.24 s. Related compiler caches may have been reused. These values exclude language loading/prefill, and source-location stripping changes artifact identity. They are not guaranteed fresh-device startup times. Mapped load time can also defer work to the first GPU call. Memory accounting must include mapped clean pages, GPU allocations, codec allocations and KV caches. No phone memory allowance, general MOS or power-efficiency claim is made.
iPhone memory qualification
This bundle is not qualified on the 8 GB iPhone 15 Pro Max. On iOS 27.0.1,
Release, first observed preparation succeeds in 45.89 s (later cached preparation
about 1.5 s), but GPU warmup is killed by the OS before a timed synthesis pass.
Both phase-switched tests use one decoder lane and a 256 MiB KV allowance;
reducing prefill from 128 to 16 positions does not avoid the failure. The reports
identify system-wide vm-pageshortage, not a returned Swift error.
The runtime verifies that Metal wraps the mapped weights rather than silently copying them. That check passes. Before the kill, logical MLX allocation is near the 5.04 GB language file, with only about 80 MB additional peak allocation and zero client neural footprint after codec release. These ledgers do not establish that the GPU/system working set fits, and do not fully attribute kernel resources. No iPhone end-to-end RTFx or complete MOSS device-placement claim is made.
Attribution
Language source revision: cdd3b911b1585e3f2dbc7775ef10f9926f58850a.
Codec/voice source revision: 3cd226ba2947efa357ef453bcad111b6eafba782.
Original weights/source use Apache-2.0; see LICENSE and the two upstream cards.
Authoring source locations are stripped from Core AI assets while graph signatures
and operation statistics are checked unchanged. No compiled specialization cache
is distributed.
Quantized
Model tree for coder543/moss-tts-v1.5-glade
Base model
OpenMOSS-Team/MOSS-Audio-Tokenizer