--- license: other license_name: apache-2.0-and-gemma license_link: LICENSE.md base_model: - google/gemma-4-E4B-it - google/embeddinggemma-300m tags: - coreml - apple-neural-engine - on-device --- ## Gemma4E4B gemma-4-E4B-it for the chunked engine of [john-rocky/CoreML-LLM](https://github.com/john-rocky/CoreML-LLM): four decode chunks plus batched prefill (N = 2048), context 4096. This conversion differs from `mlboydaisuke/gemma-4-E4B-coreml` in three ways: - **RoPE tables fixed.** `cos_full.npy` / `sin_full.npy` implement Gemma 4's proportional RoPE for full attention: only the first 64 of the 256 frequencies rotate. Upstream rotates all of them, so the model loses track of anything more than ~512 tokens back. - **Band mask for prefill.** Prefill applies the 512-token band mask to the sliding-window layers. - **Multifunction chunks.** Each `chunkN.mlmodelc` holds two functions, `decode_q1` (the default) and `prefill`, which share one copy of the weights. The `verify_qK` functions are not included. The tokenizer adds no BOS token, so prepend `` yourself. ## EmbeddingGemma EmbeddingGemma-300M from [erjigit17/embeddinggemma-300m-ane-coreml](https://huggingface.co/erjigit17/embeddinggemma-300m-ane-coreml), compiled to `model.mlmodelc`. - Inputs: `input_ids` and `attention_mask`, each int32 `[1, 128]`. - Output: `embedding`, `[1, 768]`. ## Licenses - **Gemma4E4B:** Apache 2.0, the license of `google/gemma-4-E4B-it`. - **EmbeddingGemma:** Gemma is provided under and subject to the Gemma Terms of Use found at [ai.google.dev/gemma/terms](https://ai.google.dev/gemma/terms). Use is also subject to the [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy). See `EmbeddingGemma/NOTICE`.