Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta, 2026-06).
Audio8-TTS-Preview-0.6b — Core AI
🤗 mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI · Apache-2.0 · base Edge0/Audio8-TTS-Preview-0.6b
Edge0/Audio8-TTS-Preview-0.6b (Edge0, Apache-2.0, 601M + a
337M codec) is a DualAR text-to-speech model of the Fish Audio S2 Pro design: a Qwen2.5-shaped slow AR (24
layers, 896 wide) predicts one semantic token per 46 ms frame, a 4-layer fast AR predicts the frame's nine other
codec codebooks one after another, and a 44.1 kHz DAC-style codec turns the ten codebooks into audio. Eleven
languages (Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish) and
zero-shot voice cloning from a 0.5–30 s reference recording and its transcript. This is the zoo's first DualAR /
semantic-token TTS and the first port with the sampler inside the graph. On 2026-09-28 the Hub had MLX builds
(mlx-community/Audio8-TTS-Preview-0.6b-bf16, 8-bit and 4-bit conversions), the publisher's ONNX INT4 build for CPUs
and a GGUF, and no Core AI or Core ML port of the TTS (a Core ML port exists of the publisher's ASR sibling); this port
runs the model on iPhone and Mac through Apple's runtime.
Pipeline
text (+ reference transcript + reference codes) ──(host: the publisher's prompt segments, encoded one at a time)──▶ prompt [11, P]
1. dualar.aimodel prefill(codes [1,11,32], pos) in windows ─▶ logits [32,4097] · hidden [32,896]
frame(codes [1,11,1], pos, noise_slow [2,4097], window [10], noise_fast [9,4096], forced, use_forced)
= the slow step + the semantic draw (top-k 50 / top-p 0.9 / T 0.7, Gumbel-max, RAS) + the fast AR's
ten rows with nine codebook draws — ONE call per frame ─▶ semantic · codes [10]
first_frame(logits, hidden, …) for the frame right after the prefill
host: the 10-token RAS window, the uniform draws (a seeded stream), stop at eos 151645 or 512 frames
2. codec.aimodel codes [1,10,160] ─▶ wav [1, 327680] at 44.1 kHz (every op causal: 160-frame windows, keep the last 32)
voice registration: wav 44.1 kHz [1,1,442368] ─▶ encoder.aimodel ─▶ codes [1,10,216] (0.5–10 s reference; the codes + transcript are the voice)
A frame is 2,048 samples (21.5 frames/s). The prompt is the publisher's: <|im_start|>system\n · "convert the provided
text to speech" (or, with a voice, "… reference to the following:\n\nText:\n" · <|speaker:0|> + transcript ·
"\n\nSpeech:\n" · the reference's codebook-0 codes as semantic ids) · <|im_end|>\n<|im_start|>user\n · text ·
<|im_end|>\n<|im_start|>assistant\n<|voice|> — segments encoded separately, asserted id for id against the publisher's
processor on all 18 fixtures, in Python and in the Swift host.
Graph contracts
dualar prefill in codes[1,11,32] i32 (row 0 ids, rows 1-10 codebooks; pad 151643 / 0) · pos[1] i32 out logits[32,4097] f16 · hidden[32,896] f16
frame in codes[1,11,1] i32 (the previous frame) · pos[1] · noise_slow[2,4097] f32 · window[10] i32 (-1 = none)
noise_fast[9,4096] f32 · forced[11] i32 · use_forced[1] f32 out semantic[1] i32 · codes[10] i32 · logits · hidden · fast_logits[9,4096] · sampled_*
first_frame in logits[4097] f16 · hidden[896] f16 · the same noise / window / forced out as frame
state k_cache / v_cache [24,1,2,2048,64] f16 (prefill and frame)
precision slow AR linears int8 (weight-only, per-block-32, symmetric with clipping); embeddings, the 4,097-row head, norms, the fast AR fp16
codec main in codes[1,10,160] i32 (right-pad 0; codebook 0 < 4096, codebooks 1-9 < 1024, clamped in-graph) out wav[1,327680] f16
encoder main in audio[1,1,442368] f32 mono 44.1 kHz (right-pad 0) out codes[1,10,216] i32 (keep ceil(samples/2048) frames)
The head is the embedding table's 4,097 rows (the semantic ids and eos) instead of 155,776: the publisher's sampler sets every other logit to -inf before drawing, so nothing is lost (7 MB instead of 279 MB).
What made it work
- One graph call per frame, the sampler inside. The publisher's ONNX runtime — and this port's first export —
runs a frame as a slow-AR call plus nine fast-AR calls with sampling on the host; on the raw Core AI runtime path
that is ~48 ms of engine time per 46 ms frame on an M4 Max, a call costing milliseconds before any arithmetic. The
shipped
framefunction does the slow step, the semantic draw and the fast AR's ten rows in one call, with the publisher's top-k / top-p / temperature / Gumbel-max / RAS sampler written without a sort (topk(50)+logsumexp+ a cumulative sum); the uniform draws are inputs, so a recorded draw replays the oracle's choice exactly. Details:knowledge/audio8-tts-port.md. - int8 on the slow AR, fp16 on the fast AR. On 15,093 teacher-forced codebook draws the fp16 fast AR reproduces the oracle's draw 99.5 % of the time; int8 drops to 95.8 % for 53 MB. The slow AR takes int8 well (1,671 / 1,695 semantic draws, 21 of the 24 misses at a Gumbel margin under 0.13; fp16: 1,688).
- The codec encoder needs fp32 arithmetic. Its vector quantizers pick nearest codebook entries; the whole-fp16 export flips those choices and the nine residual books cascade (44–121 of ~200 frames exact). fp16 weights with fp32 arithmetic (the Fun-ASR encoder recipe) gives 691 / 712 frames exact and codebook 0 exact on every frame.
- Kept from the publisher's code: bfloat16-rounded RoPE tables (interleaved pairs), RMSNorm eps 1e-6 in fp32, the RAS window that skips the first frame's token, the clamp of codebooks 1–9 to 1,023. The slow AR's residual peaks at 12,827 — a fifth of float16's range — so no rescale was needed.
Numerics gate (18 fixtures: 6 ja + 6 en + 2 zh sentences of our own, 4 voice-clone fixtures on FLEURS test clips)
Oracle = the publisher's modeling_arktts.py in fp32 on the CPU, seeded (generate reproduced call for call, every
uniform draw recorded). Teacher-forced = the oracle's tokens fed to the graphs; a draw comparison asks whether the
in-graph sampler, given the port's logits and the oracle's noise, makes the oracle's choice.
| stage (Mac GPU) | result |
|---|---|
| tokenizer-only prompt vs the publisher's processor | 18 / 18 prompts identical (Python and Swift) |
| plain-torch re-author, eager fp32, teacher-forced | every argmax and every replayed draw equal to the oracle; codec wav cos 1.000000 |
| dualar int8 (ship), teacher-forced, 1,695 slow steps | logits cos min 0.99985, argmax 1,668, draw 1,671 / 1,695 |
| same, 15,093 fast-AR codebook draws | logits cos min 0.99805, argmax 14,705, draw 14,711 / 15,093 |
| codec decoder on the oracle's codes, 18 utterances | wav cos min 0.99978, log-mel cos min 0.99876 |
| codec encoder (fp16w32) on the 4 reference clips, 712 frames | 691 frames exact; codebook 0 exact on every frame, every codebook ≥ 97.7 % |
| free run (the port's own loop, the oracle's draws) | reached eos 18 / 18; Swift host (Mac) == Python engine on 1,637 / 1,688 frames (10 / 18 utterances identical end to end); the iPhone's fp16 GPU agrees with the Mac's on 385 / 1,672 frames — it diverges earlier, so its audio is gated by the rows below |
| ASR round trip vs the fixture text (Fun-ASR fp32): ja CER / en WER / zh CER | Mac 2.3 % / 0.9 % / 0.0 % · iPhone 18 Pro 3.2 % / 0.0 % / 0.0 % · the oracle's own audio 6.5 / 0.0 / 0.0 |
| speaker cosine to the reference clip (WavLM-Base-Plus-SV), 4 clone fixtures | Mac mean 0.855 / min 0.561 · iPhone 0.842 / 0.572 · oracle 0.854 / 0.582 |
A miss is fp16 GPU or int8 arithmetic moving a draw that sat within a hair of the runner-up; under sampling, one flip changes every later frame, so end-to-end identity is not the bar — the ASR and speaker rows are. They are eight to fourteen sentences per language, our own measurement with one normalizer: the port's speech is as intelligible and as speaker-faithful as the publisher's fp32 output on these fixtures, not more (the Japanese "errors" are the ASR's number normalization — 三十 → 30 — on both arms, plus one real word error on the phone).
Speed
Measured with apps/Audio8Gate (the kit's Audio8TTS in a Release build, the 18 fixtures,
78 s of audio; medians; 2026-09-28; raw runs in device/, the four runs side by side in
device/tables.md). Both assets are JIT: the first load specializes them, later loads read the
cache. Two paths: streaming (a 160-frame codec window every 32 frames, audio starts after ~1.5 s of frames) and
whole utterance (synthesize: the codec once at the end — one window for anything up to 7.4 s, a fifth of the codec
work).
| frame (slow step + sampling + fast AR) | codec | time to first audio, streaming | RTF streaming, median / p90 | RTF whole utterance (bench, 5 s sentence) | first load (cold) → later loads | footprint | |
|---|---|---|---|---|---|---|---|
| M4 Max (GPU, macOS 27 26A428, GPU lock held, another conversion running: load average 3–4) | 28 ms | 0.17 s per 160-frame window | 1.2 s | 0.81 / 0.84 | 0.70 | 5.4 s (cache 0 → 1.27 GB) → 0.65 s | 0.85–0.99 GB |
| iPhone 18 Pro (GPU, iOS 27 24A437, device JIT, h19p), fresh install, thermal nominal | 33–36 ms | 0.83 s per window | 2.0 s | 1.53 / 1.83 | 0.92 (nominal, after a 145 s cool-down; 0.93 on a thermally serious phone) | 6.1 s (cache 0 → 1.27 GB) → 0.5 s | 0.59–0.72 GB |
A 46 ms frame costs the same ~33 ms of engine time on the phone as on the Mac — the frame is dispatch-bound, not compute-bound (a 512-slot cache, AOT compilation and Apple's composite RMSNorm/RoPE change it by 0–10 %). What separates the two devices is the codec: 0.17 s per 160-frame window on the Mac, 0.83 s on the phone, so streaming in 32-frame chunks — which decodes each window five times over — is real-time on the Mac and 1.5× real time on the phone, while a whole utterance decodes on the phone at about real time. Apple's pipelined engine drives a same-sized Qwen3-0.6B at 2.8 ms per token; moving the slow step onto it is the lever for a several-times faster frame — a separate round. Back-to-back synthesis heats the phone: the third run in eight minutes reached serious and its streaming frames slowed from 33 to 39 ms, but the whole-utterance bench reads the same at serious (0.93) as at nominal (0.92) — the frame is dispatch-bound either way.
Use it
CoreAIKit Audio8TTS; catalog id audio8-tts-preview-0.6b (coreai-kit#62):
import CoreAIKit
let tts = try await Audio8TTS(paths: .standard(root: modelDir)) // dualar + codec .aimodel, tokenizer/
let audio = try await tts.synthesize("明日の午後、駅前の喫茶店で待ち合わせましょう。") // 44.1 kHz mono [Float]
// streaming: tts.synthesizeStreaming(text) { chunk in play(chunk) } ~1.5 s chunks
// voice cloning: Audio8Voice(referenceText:codes:) from register_voice.py / the encoder graph
let cloned = try await tts.synthesize("Turn left at the second traffic light.", voice: voice)
⬇️ Bundle
mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI (revision 72e1c935, 2026-09-28) —
audio8_dualar_int8_cl2048_w32.aimodel/ (876 MB) + audio8_codec_decoder_fp16_t160.aimodel/ (261 MB) +
audio8_codec_encoder_fp16w32_t216.aimodel/ (416 MB, voice registration only) + tokenizer/; one subtree for macOS and
iOS (JIT .aimodels). Apache-2.0: LICENSE + NOTICE from the publisher's repository. Mirror:
coreai-community/Audio8-TTS-Preview-0.6b-CoreAI.
Convert yourself — conversion/audio8_tts/: oracle_audio8.py → parity_audio8.py →
export_audio8_frame.py --mode int8, export_audio8.py --part codec, audio8_encoder.py --frames 216 --dtype fp16w32,
gated by gate_audio8_frame.py (recipe: recipe.toml). Generated speech can be misused for
impersonation; the publisher asks for consent before cloning a voice and disclosure of synthetic audio, and so does
this port.
- Downloads last month
- 16
Model tree for mlboydaisuke/Audio8-TTS-Preview-0.6b-CoreAI
Base model
Edge0/Audio8-TTS-Preview-0.6b