Confucius4-R2T2 · Core ML for the Apple Neural Engine
English
Core ML conversion of NetEase Youdao Confucius4-R2T2 (Qwen3-ASR-1.7B architecture), made by OpenInfinity from mlx-community/Confucius4-R2T2-8bit at revision 2d6d997c3e09c65a65b1b2576b6b9b7728df8eab. It lets streaming speech recognition run on the Apple Neural Engine instead of the GPU, at much lower power.
Disclaimer (License §4.1a). Any modifications made to the original model in this Derivative Work are not endorsed, warranted, or guaranteed by the original right-holder of the original model, and the original right-holder disclaims all liability related to this Derivative Work.
该衍生品对原模型所作的任何改动与原模型原始权利人无关,原始权利人对该衍生品不背书、不担保、不承担责任。
License
These files are a Derivative Work of Confucius4-R2T2 and are governed by the NetEase Youdao Model Use License Agreement (LICENSE; the Chinese text in MODEL_LICENSE_zh prevails). By downloading or using them you accept that agreement. If you redistribute them or anything derived from them, you must keep the license and notices (NOTICE) and bind your recipients to the same agreement. The agreement prohibits high-risk uses (§4.2), such as medical diagnosis, autonomous driving, military applications and critical-infrastructure control. It also requires a separate license from NetEase Youdao above 100 million monthly active users or RMB 1 billion annual revenue (§2.2), and limits using the model to improve other AI models (§3.3c).
Contents
The repository holds two independent packages. Each package carries its own copy of the license, notice and disclaimer.
| Folder | Model | Runs on | Precision | Size |
|---|---|---|---|---|
encoder/ |
R2T2EncoderFront.mlmodelc: convolutional front end, one 1 s chunk of log-mel (100 frames) per call |
CPU | fp32 | 0.66 GB (both) |
R2T2EncoderBody.mlmodelc: audio transformer over an 8 s window (104 positions) |
Neural Engine | fp16 | ||
decoder/ |
R2T2Decoder_{0,7,14,21}.mlmodelc: 4 stateful parts × 7 text-decoder layers, 64 positions per call, KV cache of 704 positions |
Neural Engine | int8 weights, fp16 activations | 1.72 GB (all) |
R2T2Head.mlmodelc: final norm, LM head and argmax |
Neural Engine | int8 weights |
The two packages support two placements:
- Encoder on the Neural Engine (
encoder/only). The text decoder stays on MLX (GPU) from the MLX repository. - Everything on the Neural Engine (
encoder/anddecoder/).
Both placements still need the MLX repository at the revision above, for the tokenizer, the 8-bit token-embedding table (model.safetensors) and the audio front-end configuration.
Interface
| Model | Inputs | Output |
|---|---|---|
| EncoderFront | mel fp32 [1, 128, 100]; mask1 fp32 [1, 1, 1, 50] and mask2 fp32 [1, 1, 1, 25] (valid frames after each stride-2 stage, for a final partial chunk) |
embeddings (13 positions per second) |
| EncoderBody | embeddings fp16 [1, 104, 1024] (eight chunks); key_bias fp16 [1, 1, 1, 104] (0 for valid positions, a large negative value for padding) |
tokens (audio tokens for the decoder) |
| Decoder_n | hidden fp16 [1, 64, 2048]; rotary cos and sin fp16 [1, 1, 64, 128]; mask fp16 [1, 1, 64, 704] (additive causal mask over cache slots); onehot fp16 [64, 704] (cache slot written by each input position); states k_cache and v_cache fp16 [7, 8, 704, 128] |
hidden_out fp16 [1, 64, 2048] |
| Head | hidden fp16 [1, 64, 2048] |
ids: greedy token id per position |
Run the four decoder parts in order (layers 0–6, 7–13, 14–20, 21–27), each with its own MLState, then the head. The caller embeds text tokens with the MLX embedding table and places the audio tokens at the audio positions of the R2T2 prompt. Writing the KV cache through the onehot input keeps the update on the Neural Engine. Inputs shorter than 64 positions are padded and masked.
Load the front end with MLComputeUnits.cpuOnly; it loses precision on the Neural Engine. Load all other models with .cpuAndNeuralEngine. The first load compiles for the Neural Engine and takes about 5 s for the encoder and 20 s for the decoder; later loads are cached by Core ML.
Conversion notes
- The decoder's
down_projinput is scaled by 1/256 and its output by 256. The layer-2 MLP activation (about 15 000) otherwise overflows fp16 accumulation on the Neural Engine. - Uniform LUT palettization (4 and 6 bit) was evaluated and rejected because accuracy dropped too far.
Evaluation
All measurements are streaming recognition on an Apple M3 Max with macOS 26.7, using a 320 ms update cadence. The baseline is the original MLX 8-bit model on the GPU.
| MLX on GPU (baseline) | Encoder on Neural Engine | All on Neural Engine | |
|---|---|---|---|
| Extra power over idle (12 real-time clips) | 9.52 W | 5.03 W (−47 %) | 1.93 W (−80 %) |
| Update latency p95 | 209 ms | 177 ms | 560 ms |
| First text | 1.04 s | 1.04 s | 1.42 s |
| Chinese CER (42 clips) | 9.83 % | 10.01 % | 10.01 % |
| Agreement with baseline output | — | ≥ 99.4 % | ≥ 99.4 % |
Outputs are not bit-identical to the baseline: small fp16 differences on the Neural Engine flip near-ties, mostly in punctuation. System memory is about the same in all three placements; Neural Engine weights are counted as wired memory rather than process memory.
Requirements
- Apple silicon and macOS 15 or later (Core ML stateful models, iOS 18 operation set).
- Tested on macOS 26.7.
中文
网易有道 Confucius4-R2T2(Qwen3-ASR-1.7B 架构)的 Core ML 转换版,由 OpenInfinity 基于 mlx-community/Confucius4-R2T2-8bit(版本 2d6d997c3e09c65a65b1b2576b6b9b7728df8eab)转换。它让流式语音识别在 Apple 神经网络引擎上运行,而不是 GPU,耗电低得多。
免责声明(许可第 4.1(a) 条)。 该衍生品对原模型所作的任何改动与原模型原始权利人无关,原始权利人对该衍生品不背书、不担保、不承担责任。
Any modifications made to the original model in this Derivative Work are not endorsed, warranted, or guaranteed by the original right-holder of the original model, and the original right-holder disclaims all liability related to this Derivative Work.
许可
这些文件是 Confucius4-R2T2 的衍生品,受《网易有道模型使用许可协议》约束(LICENSE;以 MODEL_LICENSE_zh 中文版为准)。下载或使用即表示接受该协议。再分发这些文件或其衍生品时,必须保留许可和声明(NOTICE),并要求接收方同样遵守该协议。
协议还有以下限制:
- 禁止用于高风险场景(第 4.2 条),如医疗诊断、自动驾驶、军事和关键基础设施控制。
- 月活跃用户超过 1 亿或年收入超过 10 亿元人民币时,需另向网易有道申请授权(第 2.2 条)。
- 限制用该模型改进其他 AI 模型(第 3.3(c) 条)。
内容
仓库包含两个相互独立的包,每个包都附带一份许可、声明和免责声明。
| 文件夹 | 模型 | 运行在 | 精度 | 大小 |
|---|---|---|---|---|
encoder/ |
R2T2EncoderFront.mlmodelc:卷积前端,每次处理 1 秒的 log-mel(100 帧) |
CPU | fp32 | 0.66 GB(两个合计) |
R2T2EncoderBody.mlmodelc:音频 Transformer,处理 8 秒窗口(104 个位置) |
神经网络引擎 | fp16 | ||
decoder/ |
R2T2Decoder_{0,7,14,21}.mlmodelc:4 段有状态模型,每段 7 层文本解码器;每次处理 64 个位置,KV 缓存 704 个位置 |
神经网络引擎 | int8 权重,fp16 激活 | 1.72 GB(全部合计) |
R2T2Head.mlmodelc:最终归一化、语言模型头和 argmax |
神经网络引擎 | int8 权重 |
两个包对应两种运行方式:
- 编码器用神经网络引擎:只用
encoder/,文本解码器仍用 MLX 仓库里的模型在 GPU 上运行。 - 全部用神经网络引擎:同时使用
encoder/和decoder/。
两种方式都还需要上述版本的 MLX 仓库,用其中的分词器、8 位词嵌入表(model.safetensors)和音频前端配置。
接口
| 模型 | 输入 | 输出 |
|---|---|---|
| EncoderFront | mel:fp32 [1, 128, 100]。mask1:fp32 [1, 1, 1, 50]。mask2:fp32 [1, 1, 1, 25]。两个掩码标出每级 2 倍下采样后的有效帧,用于最后一个不满 1 秒的块。 |
embeddings(每秒 13 个位置) |
| EncoderBody | embeddings:fp16 [1, 104, 1024](8 个块)。key_bias:fp16 [1, 1, 1, 104],有效位置为 0,填充位置为很大的负数。 |
tokens(给解码器的音频 token) |
| Decoder_n | hidden:fp16 [1, 64, 2048]。旋转位置编码 cos 和 sin:fp16 [1, 1, 64, 128]。mask:fp16 [1, 1, 64, 704],按缓存槽位加的因果掩码。onehot:fp16 [64, 704],每个输入位置写入哪个缓存槽位。状态 k_cache 和 v_cache:fp16 [7, 8, 704, 128]。 |
hidden_out:fp16 [1, 64, 2048] |
| Head | hidden:fp16 [1, 64, 2048] |
ids:每个位置的贪心 token id |
按顺序运行 4 段解码器(第 0–6、7–13、14–20、21–27 层),每段使用自己的 MLState,最后运行输出头。
- 调用方用 MLX 词嵌入表把文本 token 转成向量,并把音频 token 放到 R2T2 提示词中的音频位置。
- KV 缓存通过
onehot输入写入,所以更新也在神经网络引擎上完成。 - 不足 64 个位置的输入要补齐并用掩码屏蔽。
前端要用 MLComputeUnits.cpuOnly 加载,因为它在神经网络引擎上精度不够;其余模型用 .cpuAndNeuralEngine 加载。第一次加载要为神经网络引擎编译:编码器约 5 秒,解码器约 20 秒。之后 Core ML 会缓存编译结果。
转换要点
- 解码器
down_proj的输入乘以 1/256、输出乘以 256。否则第 2 层 MLP 的激活值(约 15000)会在神经网络引擎的 fp16 累加中溢出。 - 试过均匀 LUT 调色板量化(4 位和 6 位),准确率下降太多,没有采用。
评测
所有数据都在 Apple M3 Max(macOS 26.7)上测得,测试方式是流式识别,每 320 毫秒更新一次。基准是原版 MLX 8 位模型在 GPU 上运行。
| MLX 在 GPU 上(基准) | 编码器用神经网络引擎 | 全部用神经网络引擎 | |
|---|---|---|---|
| 比空闲时多出的功耗(12 段实时音频) | 9.52 W | 5.03 W(−47%) | 1.93 W(−80%) |
| 每次更新延迟 p95 | 209 ms | 177 ms | 560 ms |
| 首字出现时间 | 1.04 s | 1.04 s | 1.42 s |
| 中文字错率(42 段) | 9.83% | 10.01% | 10.01% |
| 与基准输出一致率 | — | ≥ 99.4% | ≥ 99.4% |
输出与基准并非逐位相同:神经网络引擎上 fp16 的细小数值差异会改变几乎持平的候选,差别主要在标点。三种方式的系统内存占用基本相同;神经网络引擎的权重计入常驻内存(wired),而不是进程内存。
要求
- Apple 芯片,macOS 15 或更高版本(Core ML 有状态模型,iOS 18 操作集)。
- 已在 macOS 26.7 上测试。