File size: 4,089 Bytes
2afe215
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4557e6c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2afe215
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
---
library_name: coreml
base_model:
- ResembleAI/chatterbox
license: mit
pipeline_tag: text-to-speech
tags:
- coreml
- text-to-speech
- apple-silicon
- ios
- macos
- multilingual
language:
- en
- fr
- de
- es
- it
- pt
- pl
- tr
- ru
- nl
- cs
- ar
- zh
- ja
- hu
- ko
- hi
- da
- el
- fi
- he
- ms
- no
- sv
- sw
---

# Chatterbox Multilingual — CoreML

CoreML export of [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox)
**multilingual** (23 languages, `t3_mtl23ls_v2` + `s3gen`) for Apple platforms,
converted by [FluidInference](https://github.com/FluidInference)
(conversion toolkit: [mobius PR #89](https://github.com/FluidInference/mobius/pull/89)).

Each model ships as both `.mlpackage` (source) and compiled `.mlmodelc`.

## Models

| File | Size (fp16) | Role | Compute |
|---|---:|---|---|
| `T3-Prefill-T256-M1024-fp16` | 977 MB | Llama-520M prefill over ≤256-token context (CFG batch 2), initializes 1024-slot KV cache | CPU+GPU |
| `T3-Decode-M1024-fp16` | 977 MB | Single-step AR decode, KV cache via I/O tensors (38 ms/step) | CPU+GPU |
| `T3-Decode-M1024-fp16-stateful` | 977 MB | Single-step AR decode, KV cache in `MLState` (**16.8 ms/step**; macOS 15+/iOS 18+) | CPU+GPU |
| `Flow-N500-fp16` | 229 MB | S3Gen flow: 500-token bucket → 1000 mel frames, 10-step CFG Euler in-graph | CPU+GPU |
| `HiFT-T1000-fp16` | 40 MB | HiFTNet vocoder: mel → 24 kHz waveform (0.09 s/call) | CPU+GPU |
| `tables/tables.safetensors` | 34 MB | text/speech embedding + learned positional tables (host applies) |
| `tables/voice-default.safetensors` | 0.1 MB | precomputed built-in voice conditioning (T3 cond embeds + S3Gen ref dict) |
| `tokenizer/grapheme_mtl_merged_expanded_v1.json` | | 23-language grapheme tokenizer |

⚠️ Do **not** load the T3 packages with `.cpuOnly` — prediction hard-crashes
(also independently reported by other Chatterbox CoreML ports). Use
`.cpuAndGPU` or `.all`.

## Samples

[`samples/`](./tree/main/samples) has CoreML end-to-end renders (`e2e_*.wav`)
next to stock PyTorch renders (`baseline_*.wav`) for en/de/fr, all using the
built-in voice.

To synthesize locally without the upstream checkpoint (Apple silicon):

```bash
git clone -b feat/chatterbox-mtl-coreml https://github.com/FluidInference/mobius
cd mobius/models/tts/chatterbox/coreml
uv sync
uv run python verify/e2e_coreml.py --lang en   # models auto-download from this repo
```

## Runtime boundary

The graphs cover T3 prefill/decode (with the multilingual alignment-analyzer
attention rows as outputs), the S3Gen flow, and the HiFT vocoder. The host
runtime must provide:

- text normalization + tokenization (`tokenizer/`)
- embedding prep from `tables.safetensors` (text/speech + positional; the
  stock prefill context ends with **two** BOS embeds — replicate exactly)
- CFG combine `cond + w*(cond-uncond)`, repetition penalty, min-p/top-p,
  sampling, EOS handling
- the `AlignmentStreamAnalyzer` heuristics, fed by the exported `align_attn`
  rows (reference port: `verify/analyzer_port.py` in the conversion toolkit)
- SineGen randomness (`phase_vec`, `noise` inputs to HiFT) and CFM noise `z`
- flow bucket padding/cropping; MLState seeding from prefill KV for the
  stateful decode

Voice cloning from a reference wav additionally needs the VoiceEncoder /
S3TokenizerV2 / CAMPPlus encoders, which are not converted here; voices can
be prepared offline in Python (`export-tables.py --ref-wav`) and shipped as
`voice-*.safetensors`.

## Parity (vs upstream PyTorch)

| Check | Result |
|---|---|
| T3 wrappers vs stock (fp32) | logits 3.8e-05, alignment rows exact |
| T3 CoreML fp16 | logits 2.6e-02 (range ±15), align 1.4e-03 |
| Flow CoreML fp16 | mel max 2.4e-02, mean 2.7e-03 |
| HiFT CoreML fp16 | wav max 1.7e-02, mean 2.6e-04 |
| e2e ASR round-trip (en) | exact transcript, matches PyTorch baseline |

## License

MIT, following upstream
[ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox).
Upstream embeds Resemble's Perth watermarker in its Python pipeline; this
CoreML export does not include a watermarking stage.