Gemma 4 E2B for the Apple Neural Engine
google/gemma-4-E2B-it converted to Core ML to run on the iPhone's Neural Engine, which keeps working while an app is in the background. Text only, with a 16,384-token context. Built for Kagami Local, an on-device assistant app.
What changed from Google's release
- The transformer layers are split into four Core ML programs (
chunk_1tochunk_4), with weights quantized to 4 bits. Chunks 1 and 2 keep the key/value cache as Core ML state. - The token embeddings are stored outside the programs as 8-bit tables with one scale per row
(
embed_tokens_*), and the rotary position tables as NumPy files. - The output head returns only the most likely next token, so generation is greedy.
Each chunk has two functions: infer reads one token and prefill_b64 reads 64.
model_config.json gives the model's shape. The tokenizer, chat template, and configuration files
are Google's, from revision 3e22461f65e89153144f8adb70e3b8c2cc9845a7.
The conversion uses CoreML-LLM (MIT License).
License
Gemma 4 E2B is released by Google under the Apache License 2.0, and so is this conversion.
- Downloads last month
- 12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support