Gemma-4 E2B decision model (OptiQ 4-bit)

gemma-4-e2b-it-OptiQ-4bit turned into a decision model with optiq lora train --decision: a LoRA on the backbone and a new joint schema head that scores noul, choice and score questions in one pass. It runs on Apple Silicon through MLX and OptiQ (optiq serve exposes /v1/systemone).

Typed Decisions result

Scored on the full 400-case test split (2,000 decisions) with the benchmark's own metrics.

Model Accuracy KL from gold ↓ Brier ↓ ECE ↓ p50 latency
This model 0.802 0.072 0.039 0.181 0.99 s
Qwen3.5-0.8B-decision 0.772 0.097 0.054 0.177 0.34 s
clef-OptiQ-4bit (27B) 0.719 0.189 0.102 0.118 9.7 s
Always answer the train-split average 0.489 0.327 0.181 – –

Latency is one request per case with all five questions, on an Apple M3 Max.

This model was fitted on the benchmark's own train split and nothing else: 1,080 records, 3 epochs, LoRA rank 16. The test cases are disjoint from it, but they come from the same four workflows, so expect less on decisions of a different kind. The Clef models were trained on other data. Top-label calibration (ECE) is no better than Clef's.

Use

pip install -U mlx-optiq
optiq serve --model mlx-community/gemma-4-e2b-it-OptiQ-4bit-decision

Size on disk is 5.4 GB (4-bit backbone, vision sidecar, LoRA adapter and head).

Trained and scored with OptiQ.

Downloads last month
21
Safetensors
Model size
5B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/gemma-4-e2b-it-OptiQ-4bit-decision

Quantized
(1)
this model

Evaluation results

  • LocalLLaMA/typed-decisions leaderboard
  • Accuracy View evaluation results
    source
    Fitted by us on the benchmark's own train split only (1,080 records, LoRA rank 16 on the 4-bit backbone plus a new schema head, 3 epochs), then scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.8 *
  • Kl From Gold View evaluation results
    source
    Fitted by us on the benchmark's own train split only (1,080 records, LoRA rank 16 on the 4-bit backbone plus a new schema head, 3 epochs), then scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.07 *
  • Brier View evaluation results
    source
    Fitted by us on the benchmark's own train split only (1,080 records, LoRA rank 16 on the 4-bit backbone plus a new schema head, 3 epochs), then scored on the full 400-case test split (2,000 decisions): one request per case with all five questions, text only, the benchmark's own metrics, MLX on an Apple M3 Max.
    0.04 *