whisper large-v3-turbo Core ML encoder, 8-bit (for whisper.cpp)
A Core ML version of the encoder of OpenAI's whisper large-v3-turbo, with weights palettized to 8 bits. It's for use with whisper.cpp (and whisper.rn) built with Core ML support, next to a ggml-large-v3-turbo*.bin model, so the encoder runs on the Apple Neural Engine.
Compared with the fp16 encoder published in ggerganov/whisper.cpp (ggml-large-v3-turbo-encoder.mlmodelc.zip), the weights are half the size (637,609,152 vs 1,273,969,152 bytes). On a Mac, the Neural Engine memory it holds drops from about 1,247 MB to 646 MB, which lets it run on 4 GB iPhones.
File
ggml-large-v3-turbo-encoder.mlmodelc.zip: unzips to the ggml-large-v3-turbo-encoder.mlmodelc/ folder (the name whisper.cpp looks for next to ggml-large-v3-turbo-q5_0.bin).
- Size: 570,709,613 bytes
- SHA-256:
5b06af945696ec364544e1c9633224f6c381a2ca89bebc61097011e8b16369c7 - MD5:
89303dd190c60078a8716767c2b8e943
How it was built
- Converter: whisper.cpp v1.9.3's own
models/convert-whisper-to-coreml.py(--encoder-only True --optimize-ane True --quantize True), with a minimum deployment target of iOS 16 and float32 inputs and outputs. - Compression: coremltools 9.0
OpPalettizerConfig(mode="kmeans", nbits=8, granularity="per_tensor")on all weight tensors, then compiled withxcrun coremlc compile. - Pins: torch 2.7.0, openai-whisper 20250625, ane_transformers 0.1.3, Python 3.12.
- Reproducibility: the regenerated fp16 weights are byte-identical to the published fp16 encoder.
Accuracy (Google FLEURS da_dk test, all 930 utterances, whisper.cpp v1.9.3, Apple M4 Pro)
| Encoder | WER | CER |
|---|---|---|
| ggml (Metal) | 14.11 % | 4.86 % |
| Core ML fp16 | 14.02 % | 4.83 % |
| Core ML 8-bit (this) | 13.93 % | 4.81 % |
The differences are within noise. Built for the meetnotes app (on-device meeting transcription).
License
MIT, following the original whisper weights (© OpenAI).
Model tree for carl1758/meetnotes-models
Base model
openai/whisper-large-v3