--- license: apache-2.0 base_model: Qwen/Qwen3-ASR-0.6B tags: [coreml, neural-engine, asr, qwen3] --- # Qwen3-ASR-0.6B text decoder for the Apple Neural Engine The text decoder of [Qwen/Qwen3-ASR-0.6B](https://huggingface.co/Qwen/Qwen3-ASR-0.6B), converted to Core ML with [ANEMLL](https://github.com/Anemll/Anemll) 0.3.5. - `decoder.mlmodelc`: 28 layers, two functions over one KV cache (context 512): `prefill` reads 64 positions per call, `infer` writes one. - `lm_head.mlmodelc`: final norm is in the decoder; the head returns the best token of each of 16 vocabulary slices (`argmax_idx`, `argmax_val`). - Weights: LUT8 palettized, 8 channels per group. Needs iOS 18 / macOS 15. Pair it with the audio encoder and token embeddings from [aufklarer/Qwen3-ASR-CoreML](https://huggingface.co/aufklarer/Qwen3-ASR-CoreML) (`encoder.mlmodelc`, `embedding.mlmodelc`). Audio embeddings go into `prefill` as hidden states; the chat template and tokenizer are the base model's. Measured on an M3 Max Neural Engine against the fixed 128-position decoder in aufklarer/Qwen3-ASR-CoreML. Decode times come from live captioning replayed in real time (10 minutes of Mandarin video, each decode stretched 1.1x to match an iPhone 16 Pro). Error rates are against Qwen3-ASR-1.7B's transcript of about 23 minutes of speech. | | fixed 128 | this build | |---|---|---| | preview decode, median | 506 ms | 245 ms | | final decode, median | 1641 ms | 634 ms | | decoder busy | 62% | 32% | | character error rate | 8.14% | 8.19% | License: Apache 2.0, as the base model.