MOSS-TTS-v1.5 for Glade

Hybrid conversion of MOSS-TTS-v1.5 and MOSS Audio Tokenizer. The 8.49B autoregressive language model runs on GPU through native Swift MLX; seventeen Core AI codec groups render mono 24 kHz PCM on ANE. English controls are qualified; other upstream languages are not qualified here.

Model assets and runtime configuration for Glade. Python is export/measurement tooling, not an application dependency. Mac qualification is included; iPhone and Watch qualification is not claimed. Download weights and configuration from one snapshot.

Layout and downloads

  • language/: mapped quantized language weights, tensor layout, configuration and tokenizer. Four-bit affine matrix groups, eight-bit audio heads and FP16 audio embeddings preserve the selected quantized runtime. No further training.
  • codec/: seventeen FP16 .aimodel groups plus original FP32 codebooks and boundary projections in boundary.f32/boundary.json.
  • voice-encoder/: optional original FP32 reference-audio encoder for importing client-supplied voices. Encoded conditioning can be cached and reused. Download this only when reference voice cloning is needed. No voice recording is included.
hf download coder543/moss-tts-v1.5-glade \
  --include 'language/*' 'codec/*' LICENSE README.md --local-dir ./moss

Glade's optional GladeMOSS product consumes language/ and codec/ directories. It supplies text cleanup, prompt construction, dynamic continuous batching, ordered incremental output, optional phase switching and MFA-assisted runout recovery. The planner targets five minutes per text unit with a hard limit of 5,000 generated positions; prompt positions are additional. Completed lanes leave the actual GPU batch. Memory budgets can reduce admitted lanes.

Mapped weights avoid repeated loading copies; clean pages are reclaimable, but mapping does not make the approximately 4.69 GiB of language weights disappear. The app must keep a mapped file unchanged while loaded. Generation retains each unit's full language history; codec blocks preserve their causal acoustic caches. Reference voice import can load the encoder before the language model and retain only encoded conditioning afterward. Neither path requires Python.

Measured Mac controls

M3 MacBook Air, 16 GB, macOS 27.0.1; Release runtime. These are bounded English functional controls, not a full Moon-speech benchmark or general perceptual study.

  • A longer GPU-language/ANE-codec control generated 48.40 s of audio in 37.91 s (1.28×), excluding preparation and WAV writing. Its language frames exactly matched the separate GPU-codec control.
  • Two complete 197-word units produced 157.36 s of audio. Source-ordered online phased and resident rendering produced identical PCM and language tokens.
  • A deliberately forced runout retained verified audio through a sentence at 25.63 s after 20.16 s was delivered, then generated only the unspoken suffix. The other batched sequence remained intact. Cohere recovered all 394 normalized words from both this output and the ordinary control.
  • The original FP32 codec comparison on a 78.48-second token stream differs by 5.60% relative waveform RMS. Streaming/phase policies match each other exactly; the converted codec is not numerically lossless against FP32.

Initial observed preparation of the seventeen codec assets took 30.51 s; subsequent cached loading took 0.24 s. Related compiler caches may have been reused. These values exclude language loading/prefill, and source-location stripping changes artifact identity. They are not guaranteed fresh-device startup times. Mapped load time can also defer work to the first GPU call. Memory accounting must include mapped clean pages, GPU allocations, codec allocations and KV caches. No phone memory allowance, general MOS or power-efficiency claim is made.

iPhone memory qualification

This bundle is not qualified on the 8 GB iPhone 15 Pro Max. On iOS 27.0.1, Release, first observed preparation succeeds in 45.89 s (later cached preparation about 1.5 s), but GPU warmup is killed by the OS before a timed synthesis pass. Both phase-switched tests use one decoder lane and a 256 MiB KV allowance; reducing prefill from 128 to 16 positions does not avoid the failure. The reports identify system-wide vm-pageshortage, not a returned Swift error.

The runtime verifies that Metal wraps the mapped weights rather than silently copying them. That check passes. Before the kill, logical MLX allocation is near the 5.04 GB language file, with only about 80 MB additional peak allocation and zero client neural footprint after codec release. These ledgers do not establish that the GPU/system working set fits, and do not fully attribute kernel resources. No iPhone end-to-end RTFx or complete MOSS device-placement claim is made.

Attribution

Language source revision: cdd3b911b1585e3f2dbc7775ef10f9926f58850a. Codec/voice source revision: 3cd226ba2947efa357ef453bcad111b6eafba782. Original weights/source use Apache-2.0; see LICENSE and the two upstream cards. Authoring source locations are stripped from Core AI assets while graph signatures and operation statistics are checked unchanged. No compiled specialization cache is distributed.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/moss-tts-v1.5-glade

Quantized
(6)
this model

Collection including coder543/moss-tts-v1.5-glade