Configuration Parsing Warning:In config.json: "num_experts" must be a number

▚ gemma4-e4b-coder · W4A16

GPTQ · int4 g128 · vLLM

> stock vLLM, no forks, no custom kernels · calibrated on 15 M tokens of its own agentic distribution

◈ 3.65 GiB · 2.48× smaller ◈ compressed-tensors · pack-quantized ◈ calibrated on 15,001,898 tok @ ctx 32,768 ◈ 131,072 ctx on an 8 GB card ◈ vocab 65,536 · tokenizer bundled ◈ MTP drafter bundled · 2.29× decode

▶ What this is

A compressed-tensors W4A16 (GPTQ, int4 group-128) build of gemma4-e4b-coder — trained on ~1.1 B tokens of real agentic-coding sessions — that stock vLLM serves directly, with a bundled MTP drafter for 2.3× faster decoding. Calibrated on 15 M tokens of its own training distribution at 32,768 context. For llama.cpp / Ollama / LM Studio use the GGUF sibling.

▶ Load the tokenizer from this repo

The vocabulary is pruned to 65,536 tokens and the ids are not Gemma's canonical ids. Pointing this at the stock google/gemma-4-E4B-it tokenizer produces fluent-looking nonsense, with no error. There is no runtime check that catches it.

▚ Quick start

vLLM with the bundled MTP drafter (2.3× faster decoding):

python3 -m venv .venv && . .venv/bin/activate
pip install vllm==0.26.0 transformers==5.10.1
hf download pearsonkyle/gemma4-e4b-coder-gptq-w4a16 --local-dir gemma4-e4b-coder-gptq-w4a16

vllm serve gemma4-e4b-coder-gptq-w4a16 \
    --max-model-len 131072 \
    --max-num-seqs 8 \
    --gpu-memory-utilization 0.90 \
    --enable-auto-tool-choice --tool-call-parser gemma4 \
    --reasoning-parser gemma4 \
    --speculative-config '{"method":"mtp","model":"gemma4-e4b-coder-gptq-w4a16/drafter","num_speculative_tokens":7}'

Then, from a second shell, python3 gemma4-e4b-coder-gptq-w4a16/deploy/smoke_test.py checks tool calls, the reasoning split, decode speed and that the drafter is accepting tokens, and prints PASS.

  • Keep the gemma4 parsers. The chat template emits Gemma 4's native <|tool_call>call:name{…}<tool_call|> format and a separate thinking channel; with pythonic the tool calls come back as plain text.
  • Tested from a clean venv on one RTX 4060 Ti 16 GB (Python 3.13, CUDA 13); the full environment is pinned in deploy/requirements.txt. On a smaller card, lower --max-model-len or --max-num-seqs (see Memory below).
  • transformers loads it too (with compressed-tensors), but dequantizes on every matmul, so use vLLM to serve.

▚ Speed

RTX 4060 Ti 16 GB, one stream decode tok/s tokens per target pass
no drafter 93.6 1.00
Google's E4B assistant, vocab remapped 193.5 2.52
drafter/ (remapped + fine-tuned) 214.4 2.80

vLLM 0.26.0, greedy, --max-num-seqs 1, 5 coding prompts × 400 tokens, 7 draft tokens (4 gives 193 tok/s, 10 gives 197). Raw numbers: eval/drafter_bench.json. The speedup depends on what is generated and shrinks as concurrent requests fill the GPU.

▶ How the drafter was built

drafter/ is Google's Gemma-4-E4B MTP assistant (Gemma4AssistantForCausalLM, 4 layers, 55 MB). Every token in this model's pruned 65,536-token vocabulary exists in Gemma's original one, so the assistant's output head was rebuilt by copying rows by token string — Google's original drafter predicts Gemma's canonical ids and does not fit this checkpoint. It was then fine-tuned for 1,000 steps on 13 M tokens of this model's own training data, against this checkpoint's W4A16 hidden states, reproducing exactly how vLLM drafts: input embed(token_{t+1}) + h_t, target K/V up to t only, distilled on the target's own next-token distribution, 3 draft steps unrolled. 8-bit AdamW with fp32 master weights, lr 6e-5 cosine, ~4 h on one RTX 4060 Ti.

Held-out first-draft acceptance (vs. the target's greedy token, 3,062 positions): 60.6 % with the remapped head alone, 65.8 % after fine-tuning. Provenance: drafter/drafter_train.json; code: quant_tuner.drafter in Quant-Tuner.


▚ Memory — how many tokens fit

Only 4 of 42 layers grow their KV cache with context (the rest reuse shared KV or run a 512-token sliding window), so a sequence costs 16 KiB × ctx + 20 MiB at bf16 — 2.02 GiB at 131,072. Context length and concurrency trade against one budget:

cardctx 8,192ctx 32,768ctx 131,072
8 GB17 seq · 0.14 M tok4 seq · 0.13 M tok1 seq · 0.13 M tok
12 GB42 seq · 0.34 M tok11 seq · 0.36 M tok3 seq · 0.39 M tok
16 GB67 seq · 0.55 M tok18 seq · 0.59 M tok4 seq · 0.52 M tok
24 GB117 seq · 0.96 M tok32 seq · 1.05 M tok8 seq · 1.05 M tok
32 GB167 seq · 1.37 M tok46 seq · 1.51 M tok11 seq · 1.44 M tok
48 GB266 seq · 2.18 M tok74 seq · 2.42 M tok19 seq · 2.49 M tok
96 GB565 seq · 4.63 M tok157 seq · 5.14 M tok40 seq · 5.24 M tok

Derived from config.json (scripts/kv_budget.py in Quant-Tuner), not measured — vLLM's startup line GPU KV cache size: … tokens is the real number. Reaching these needs --max-num-seqs raised to match. --kv-cache-dtype fp8_e4m3 halves the KV cost but is uncalibrated here.


▚ Quality

Tool-calling on 107 held-out turns, bf16 vs W4A16 in the same harness and settings, so the rows can be compared pair by pair:

metric · 107 held-out turnsbf16W4A16paired test
Schema-valid tool calls0.86920.87859 discordant (4/5) · McNemar p = 1.00
Tool-selection accuracy0.62620.588814 discordant (9/5) · McNemar p = 0.42
Parameter accuracy0.46860.4407Δ −0.028 · 95% CI [−0.086, +0.028]

No difference survives a paired test — every interval contains zero; the 3.7-point tool-selection gap is 14 disagreements split 9/5. That does not make W4A16 lossless: at n = 107 the suite resolves about ±0.09 on parameter accuracy. The bf16 arm reproduces the base model's published numbers to within one turn.

▶ How it was made — quantization settings, what stayed bf16, calibration corpus

GPTQ W4A16 via llm-compressor 0.13.0 → compressed-tensors 0.18.0: int4, group 128, symmetric, pack-quantized; 258 Linear modules quantized across 42 layers. --pipeline basic is required for Gemma 4 — the default sequential pipeline breaks on its cross-layer shared KV (KeyError: 'sliding_attention').

Kept bf16 on purpose — 132 modules:

held outcountwhy
embed_tokens2wide vocab table; quantizing it is the classic rare-token failure
per_layer129Gemma 4's per-layer input embedding table
lm_head1tied to embed_tokens

config.json's ignore list holds the 86 of those that are Linear modules; quant_tuner_ptq.json records all 132. Vision/audio patterns matched nothing (those towers were removed at vocabulary pruning), and no checkpoint tensors were dropped at export.

Calibration: 15,001,898 tokens / 11,982 conversations / 458 sequences at ctx 32,768, token-balanced with no source above 6 %, drawn only from the train slice so the evaluation holdouts stay clean; 12,175 tool calls and 2,940 tool-schema declarations survive templating. Corpus SHA-256, config groups and ignore counts are in quant_tuner_ptq.json. Built with Quant-Tuner.

▶ Limits
  • Text only — vision and audio towers were removed at vocabulary pruning.
  • Not for multilingual use — non-Latin scripts fall back toward per-byte tokens under the pruned vocabulary.
  • Drafter speedups were measured on coding prompts; other text may accept fewer draft tokens.
  • Inherits the base model's limits, including that its held-out decision accuracy never improved during training — only format did.

▚ Family

repofor
gemma4-e4b-coderbf16 merged weights + stage-1 resume state
gemma4-e4b-coder-gptq-w4a16this repo — vLLM
gemma4-e4b-coder-ggufllama.cpp / Ollama / LM Studio
gemma4-e4b-stage0-32k-v65536stage-0 base, for attaching the adapter

▚ License

Gemma Terms of Use, inherited from google/gemma-4-E4B-it.

Downloads last month
203
Safetensors
Model size
5B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pearsonkyle/gemma4-e4b-coder-gptq-w4a16

Quantized
(2)
this model

Collection including pearsonkyle/gemma4-e4b-coder-gptq-w4a16