MLX port for Apple silicon (System 1): Avicennasis/GEV-26B-Decide-mlx-8bit / -4bit

#1
by Avicennasis - opened

Thanks for releasing GEV-26B-Decide. There was no MLX build, so we converted System 1 for Apple silicon:

How each build was made:

  1. The adapter/ LoRA is merged into the backbone with peft's merge_and_unload, and the original config.json is kept.
  2. The merged model is converted with mlx_lm.convert (affine, group size 64).
  3. Your head.safetensors, judge_config.json and calibration files are copied unchanged.

A small torch-free loader, gev_mlx.py, runs the bare-v1 read-out and your tournament for more than 16 options. It
also serves /v1/decide (System 1 only) and /v1/systemone.

We checked parity against your transformers + peft path (bf16, PyTorch MPS) on 121 decisions from a private set:

  • The top answer is the same on 118/121 at both quantizations.
  • The three that differ sat at 0.41–0.59 confidence on both sides.
  • The median per-decision max |Δp| is 0.006 at 8-bit and 0.012 at 4-bit.
  • One decision takes about 0.2 s on an M1 Max.

A note for anyone else converting: transformers 5.17 re-saves Gemma 4's global_head_dim and
num_global_key_value_heads as a per_layer_config block, which mlx-lm 0.32 cannot read. After a merge, keep the
original config.json.

Adaptive thinking is not ported. If a link from your card would be useful, feel free to add one.

AutoTrust AI Lab org

thank you very much

AutoTrust AI Lab org

Hi @Avicennasis

Thank you for the MLX builds of GEV-26B-Decide. Apple silicon was a gap in our release, and you filled it carefully: merge checks that confirm the adapter actually loaded, the head and calibration files left untouched, and a real parity run against our transformers + peft path rather than a smoke test.

The results are what we would hope for. At about 0.2 s per decision on an M1 Max, System 1 becomes practical on a laptop for exactly the kind of guardrail and triage work you tested it on.

Your config.json note will save others real time. We'll link both builds from the GEV-26B-Decide card, along with that tip.

We would welcome more collaboration. A few areas where help would matter most:

  • Adaptive thinking (System 2) on MLX, so the full System 1 + System 2 path runs locally
  • The vision tower, for multimodal decisions on Apple silicon
  • Batched decisions in gev_mlx.py

If any of the decisions the reference got wrong on your set can be shared, even redacted, that would be valuable input for the next version. We're also glad to give you a heads-up before future releases so MLX builds can land alongside them.

Reply here to Cloud Yu (Yu Hai) or me if you'd like to work together more closely. Thanks again.

Josh Liu

Thanks Josh and Cloud! 😺 I'd be happy to help with more of this. Batching looks like the best place to start, and it'll be useful for my own evaluation runs too.

For System 2, there's one wrinkle: the current MLX builds have the System 1 adapter merged in, so they can't switch back to the unadapted backbone the way your reference server does. I'd like to work out that separation properly rather than just enable generation on the merged weights.

I'm interested in vision as well. mlx-vlm already has Gemma 4 support, so there's a starting point, though the decision-head integration and image parity still need testing. My current plan is batching first, then System 2, then vision, without putting dates on those yet.

Where would you prefer the code contributions to land? If you have reference cases you'd like us to use for batching or multimodal checks, those would help. I'd also appreciate a heads-up on future releases.

I'll go through the cases the reference got wrong and see which I can share in sanitized form.

Sign up or log in to comment