Gemma 4 E2B for the Apple Neural Engine

google/gemma-4-E2B-it converted to Core ML to run on the iPhone's Neural Engine, which keeps working while an app is in the background. Text only, with a 16,384-token context. Built for Kagami Local, an on-device assistant app.

What changed from Google's release

  • The transformer layers are split into four Core ML programs (chunk_1 to chunk_4), with weights quantized to 4 bits. Chunks 1 and 2 keep the key/value cache as Core ML state.
  • The token embeddings are stored outside the programs as 8-bit tables with one scale per row (embed_tokens_*), and the rotary position tables as NumPy files.
  • The output head returns only the most likely next token, so generation is greedy.

Each chunk has two functions: infer reads one token and prefill_b64 reads 64. model_config.json gives the model's shape. The tokenizer, chat template, and configuration files are Google's, from revision 3e22461f65e89153144f8adb70e3b8c2cc9845a7.

The conversion uses CoreML-LLM (MIT License).

License

Gemma 4 E2B is released by Google under the Apache License 2.0, and so is this conversion.

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ericlmtn/gemma-4-e2b-it-neural-engine

Finetuned
(392)
this model