Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
This model has no row on DeviceMark, the on-device LLM leaderboard.
VibeVoice-Realtime-0.5B β Core AI
microsoft/VibeVoice-Realtime-0.5B
(MIT) converted to Apple Core AI β the zoo's
first multi-speaker / dialogue (podcast-style) TTS. iPhone + Mac, all-fp16.
Not a "first on-device VibeVoice" claim β other CoreML/GGUF ports exist. What this is: the zoo's first multi-speaker TTS, app-integrated, and the other half of a generate -> diarize loop with the zoo's Streaming Sortformer diarizer.
Architecture
Dual Qwen2.5 LM (4-layer text context LM, norm = Identity + 20-layer speech trunk) ->
per-frame next-token diffusion (4-layer adaLN head, DDPM cosine, v-prediction, DPMSolver++
5-step, CFG 1.5) -> causal-conv acoustic VAE decoder (7.5 Hz latent -> 24 kHz, 3200 samples/frame).
The LM predicts one latent per frame; the diffusion head denoises it; the VAE decoder renders audio.
Multi-speaker output is host turn-switching: each Speaker N: turn is generated from its own
voice preset and the turns are concatenated β no multi-speaker prefill, no acoustic encoder.
Contents
| path | what |
|---|---|
macos/vibevoice_mainlm_fp16_decode_cl512.aimodel |
context LM, q=1 decode, KV-stateful, cache 512 (114 MB) |
macos/vibevoice_ttslm_fp16_decode_cl512.aimodel |
speech trunk, same shape; also drives the CFG-negative stream (569 MB) |
macos/vibevoice_diffusion_head_fp16.aimodel |
prediction head, (noisy[2,64], t[2]) -> [2,64] (80 MB) |
macos/vibevoice_connector_fp16.aimodel |
acoustic connector, latent[1,1,64] -> embed[1,1,896] (1.7 MB) |
macos/vibevoice_decoder_fp16_t64.aimodel |
acoustic VAE decoder, latents[1,64,64] -> audio[1,1,204800] (656 MB) |
ios/*.aimodel |
the same five JIT bundles, byte for byte; every iPhone generation specializes them on its first load |
ios-h18p/*.h18p.aimodelc |
the same five, compiled ahead of time for the iPhone 17 Pro (h18p), that phone only, GPU; moved from ios/ in revision b4480f87 (2026-09-26) |
voices/*.pt |
25 upstream voice presets (EN/ZH/β¦): pre-computed prefill KV, so no acoustic encoder is shipped |
coreai_host/voices/<name>/ |
the same 25 presets as flat fp16 blobs a Swift host reads directly (no torch) |
coreai_host/glue/ |
type embeddings, EOS classifier, DPMSolver++ schedule, Qwen2.5 tokenizer |
coreai_host/embed/embed_tokens_fp16.bin |
(151936, 896) fp16 embedding table, mmapped for the host token lookup |
device_bundle/ |
compact host inputs + golden.f32 for the on-device self-test |
fp16 is required. int8 LMs diverge inside the speech feedback loop (min cos 0.187, early EOS); the diffusion head is fp16-sensitive too (pure-torch fp16 collapses to 0.79 β Core AI keeps the RMSNorm/adaLN reductions in fp32, so the host DDPM reference must run fp32).
Fixed shapes only. Every graph is static (q=1 decode, fixed-T decoder), so the runtime must
not be given the expectFrequentReshapes hint on iOS. With the h18p AOT bundles on the iPhone
17 Pro, the hint made the runtime skip the AOT specialization and compile on device. That compile
segfaulted inside the MPSGraph AICode compiler.
Gates
| gate | result |
|---|---|
| diffusion head / connector, engine vs oracle | cos 0.999999 / 1.000000 |
| acoustic decoder (T=30 / T=64), engine vs non-stream golden | cos 1.000004 / 1.000005 |
| main LM / tts LM decode, engine vs torch | cos 0.999999 / 0.999996 |
| Python E2E on all 5 engines vs the upstream streamed wav | latent min cos 0.999198, wav cos 0.999479 |
| iPhone 17 Pro (A19 Pro, AOT h18p, GPU) vs the golden | cos 0.998308 |
On the iPhone 17 Pro with the h18p AOT bundles: 6 graph loads in 2.6 s (warm), 24 latents / 3.20 s of audio in 2.3 s = 10.6 tok/s ~ 1.4x real-time.
Measured 2026-09-26 on an iPhone 18 Pro (iPhone19,2, iOS 27.0 build 24A437, h19p) with the zoo's
DecideGate app in its load-only mode, without the increased-memory entitlement. Each first load was the
graph's first in an app container that held no specialization of it; the app had run other graphs
before. The call is one run on all-zero inputs; a graph with a KV state was loaded only. Each first load
wrote a specialization of about the bundle's size into the app container, and the load after a relaunch
reused it. One measurement per graph
(knowledge/jit-distribution.md).
The same phone refuses an h18p bundle with incompatibleCompiledAssetArchitecture.
JIT bundle in ios/ |
MB | first load | first call | load after relaunch |
|---|---|---|---|---|
vibevoice_connector_fp16.aimodel |
1.7 | 0.04 s | 93 ms | 0.01 s |
vibevoice_diffusion_head_fp16.aimodel |
84 | 0.17 s | 415 ms | 0.06 s |
vibevoice_mainlm_fp16_decode_cl512.aimodel |
120 | 0.20 s | β (KV state) | 0.07 s |
vibevoice_ttslm_fp16_decode_cl512.aimodel |
597 | 1.36 s | β (KV state) | 0.50 s |
vibevoice_decoder_fp16_t64.aimodel |
688 | 1.78 s | 1,183 ms | 0.08 s |
Use it
import CoreAIKit
let dialogue = try await KitDialogue(catalog: "vibevoice-realtime-0.5b")
let (audio, turns) = try await dialogue.perform("""
Speaker 1: Did you know this runs entirely on the phone?
Speaker 2: No cloud at all? That is wild.
""")
Swift host reference: ondevice/VibeVoiceRunner (Mac) and VibeVoiceSelfTest.swift in the zoo's
coreai-audio app β
raw Core AI stateful-KV loop + a Swift DPMSolver++ sampler. Python host + conversion recipe:
conversion/vibevoice
(host_e2e.py = the full generate loop, host_multispeaker.py = the dialogue demo).
Base model: microsoft/VibeVoice-Realtime-0.5B (MIT). EN/ZH. Community port β not an Apple model.
More models in this format: Core AI Model Zoo β 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request β free, open weights only; the export and its measured numbers get published publicly.
- Downloads last month
- 82