Configuration Parsing Warning:In config.json: "num_experts" must be a number

Gemma 4 E4B Text for MLX (4-bit)

google/gemma-4-E4B-it as a text-only model for Apple's MLX, compressed to 4 bits. It runs on device on Apple silicon: Mac, and iPhone or iPad with 8 GB of memory.

Gemma 4 E4B can also read images and audio. This version keeps only the language model, so it is 4.2 GB instead of 5.2 GB and writes exactly the same text.

All credit for the model goes to Google. This repository only changes its file format.

At a glance

Detail Value
Base model google/gemma-4-E4B-it (instruction-tuned)
Input and output Text only
Version 4-bit: weights compressed from 16 to 4 bits (affine, groups of 64)
Size 4.2 GB (the multimodal 4-bit MLX version is 5.2 GB)
Model type gemma4_text
Runtime mlx-lm (Python) or mlx-swift-lm
License Apache 2.0, the same as the original model

Usage

pip install -U mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("UIDUser-NSB/Gemma4-E4B-Text-MLX")
messages = [{"role": "user", "content": "Summarize this meeting in three points: ..."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512))

Pass enable_thinking=False for a direct answer. Without it, Gemma 4 writes its reasoning first.

Check

With greedy decoding, this model and mlx-community/gemma-4-e4b-it-4bit (the multimodal version) write the same text for the same prompt.

How it was converted

  1. Source: google/gemma-4-E4B-it, revision ee0ef6023621cff504d758262d4e04895a5af4a2.
  2. Text only: the language model's weights are kept and renamed from model.language_model.* to model.*. The vision and audio encoders and their embedders are removed.
  3. Config: model_type is gemma4_text. The text model's settings (text_config) are also written at the top level, where mlx-lm and mlx-swift-lm read them for this model type.
  4. Compression: python -m mlx_lm convert -q --q-bits 4 --q-group-size 64 with mlx-lm 0.32.0.

Credits

Gemma 4 by Google DeepMind. See the original model card for training data, evaluations, intended use and limitations.

Downloads last month
-
Safetensors
Model size
7B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for UIDUser-NSB/Gemma4-E4B-Text-MLX

Quantized
(387)
this model