|
Download README.md from FluidInference/chatterbox-multilingual-coreml: direct link, hf CLI and curl.
- Browser
- Download file 4.09 kB
-
https://huggingface.co/FluidInference/chatterbox-multilingual-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/chatterbox-multilingual-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/chatterbox-multilingual-coreml/resolve/main/README.md
4.09 kB
| library_name: coreml | |
| base_model: | |
| - ResembleAI/chatterbox | |
| license: mit | |
| pipeline_tag: text-to-speech | |
| tags: | |
| - coreml | |
| - text-to-speech | |
| - apple-silicon | |
| - ios | |
| - macos | |
| - multilingual | |
| language: | |
| - en | |
| - fr | |
| - de | |
| - es | |
| - it | |
| - pt | |
| - pl | |
| - tr | |
| - ru | |
| - nl | |
| - cs | |
| - ar | |
| - zh | |
| - ja | |
| - hu | |
| - ko | |
| - hi | |
| - da | |
| - el | |
| - fi | |
| - he | |
| - ms | |
| - no | |
| - sv | |
| - sw | |
| # Chatterbox Multilingual — CoreML | |
| CoreML export of [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox) | |
| **multilingual** (23 languages, `t3_mtl23ls_v2` + `s3gen`) for Apple platforms, | |
| converted by [FluidInference](https://github.com/FluidInference) | |
| (conversion toolkit: [mobius PR #89](https://github.com/FluidInference/mobius/pull/89)). | |
| Each model ships as both `.mlpackage` (source) and compiled `.mlmodelc`. | |
| ## Models | |
| | File | Size (fp16) | Role | Compute | | |
| |---|---:|---|---| | |
| | `T3-Prefill-T256-M1024-fp16` | 977 MB | Llama-520M prefill over ≤256-token context (CFG batch 2), initializes 1024-slot KV cache | CPU+GPU | | |
| | `T3-Decode-M1024-fp16` | 977 MB | Single-step AR decode, KV cache via I/O tensors (38 ms/step) | CPU+GPU | | |
| | `T3-Decode-M1024-fp16-stateful` | 977 MB | Single-step AR decode, KV cache in `MLState` (**16.8 ms/step**; macOS 15+/iOS 18+) | CPU+GPU | | |
| | `Flow-N500-fp16` | 229 MB | S3Gen flow: 500-token bucket → 1000 mel frames, 10-step CFG Euler in-graph | CPU+GPU | | |
| | `HiFT-T1000-fp16` | 40 MB | HiFTNet vocoder: mel → 24 kHz waveform (0.09 s/call) | CPU+GPU | | |
| | `tables/tables.safetensors` | 34 MB | text/speech embedding + learned positional tables (host applies) | | |
| | `tables/voice-default.safetensors` | 0.1 MB | precomputed built-in voice conditioning (T3 cond embeds + S3Gen ref dict) | | |
| | `tokenizer/grapheme_mtl_merged_expanded_v1.json` | | 23-language grapheme tokenizer | | |
| ⚠️ Do **not** load the T3 packages with `.cpuOnly` — prediction hard-crashes | |
| (also independently reported by other Chatterbox CoreML ports). Use | |
| `.cpuAndGPU` or `.all`. | |
| ## Samples | |
| [`samples/`](./tree/main/samples) has CoreML end-to-end renders (`e2e_*.wav`) | |
| next to stock PyTorch renders (`baseline_*.wav`) for en/de/fr, all using the | |
| built-in voice. | |
| To synthesize locally without the upstream checkpoint (Apple silicon): | |
| ```bash | |
| git clone -b feat/chatterbox-mtl-coreml https://github.com/FluidInference/mobius | |
| cd mobius/models/tts/chatterbox/coreml | |
| uv sync | |
| uv run python verify/e2e_coreml.py --lang en # models auto-download from this repo | |
| ``` | |
| ## Runtime boundary | |
| The graphs cover T3 prefill/decode (with the multilingual alignment-analyzer | |
| attention rows as outputs), the S3Gen flow, and the HiFT vocoder. The host | |
| runtime must provide: | |
| - text normalization + tokenization (`tokenizer/`) | |
| - embedding prep from `tables.safetensors` (text/speech + positional; the | |
| stock prefill context ends with **two** BOS embeds — replicate exactly) | |
| - CFG combine `cond + w*(cond-uncond)`, repetition penalty, min-p/top-p, | |
| sampling, EOS handling | |
| - the `AlignmentStreamAnalyzer` heuristics, fed by the exported `align_attn` | |
| rows (reference port: `verify/analyzer_port.py` in the conversion toolkit) | |
| - SineGen randomness (`phase_vec`, `noise` inputs to HiFT) and CFM noise `z` | |
| - flow bucket padding/cropping; MLState seeding from prefill KV for the | |
| stateful decode | |
| Voice cloning from a reference wav additionally needs the VoiceEncoder / | |
| S3TokenizerV2 / CAMPPlus encoders, which are not converted here; voices can | |
| be prepared offline in Python (`export-tables.py --ref-wav`) and shipped as | |
| `voice-*.safetensors`. | |
| ## Parity (vs upstream PyTorch) | |
| | Check | Result | | |
| |---|---| | |
| | T3 wrappers vs stock (fp32) | logits 3.8e-05, alignment rows exact | | |
| | T3 CoreML fp16 | logits 2.6e-02 (range ±15), align 1.4e-03 | | |
| | Flow CoreML fp16 | mel max 2.4e-02, mean 2.7e-03 | | |
| | HiFT CoreML fp16 | wav max 1.7e-02, mean 2.6e-04 | | |
| | e2e ASR round-trip (en) | exact transcript, matches PyTorch baseline | | |
| ## License | |
| MIT, following upstream | |
| [ResembleAI/chatterbox](https://huggingface.co/ResembleAI/chatterbox). | |
| Upstream embeds Resemble's Perth watermarker in its Python pipeline; this | |
| CoreML export does not include a watermarking stage. | |