Instructions to use autotrust/GEV-26B-Decide with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use autotrust/GEV-26B-Decide with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="autotrust/GEV-26B-Decide")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("autotrust/GEV-26B-Decide") model = AutoModelForMultimodalLM.from_pretrained("autotrust/GEV-26B-Decide", device_map="auto") - Notebooks
- Google Colab
- Kaggle
MLX port for Apple silicon (System 1): Avicennasis/GEV-26B-Decide-mlx-8bit / -4bit
Thanks for releasing GEV-26B-Decide. There was no MLX build, so we converted System 1 for Apple silicon:
- https://huggingface.co/Avicennasis/GEV-26B-Decide-mlx-8bit (25 GB)
- https://huggingface.co/Avicennasis/GEV-26B-Decide-mlx-4bit (13 GB)
How each build was made:
- The
adapter/LoRA is merged into the backbone with peft'smerge_and_unload, and the originalconfig.jsonis kept. - The merged model is converted with
mlx_lm.convert(affine, group size 64). - Your
head.safetensors,judge_config.jsonand calibration files are copied unchanged.
A small torch-free loader, gev_mlx.py, runs the bare-v1 read-out and your tournament for more than 16 options. It
also serves /v1/decide (System 1 only) and /v1/systemone.
We checked parity against your transformers + peft path (bf16, PyTorch MPS) on 121 decisions from a private set:
- The top answer is the same on 118/121 at both quantizations.
- The three that differ sat at 0.41–0.59 confidence on both sides.
- The median per-decision max |Δp| is 0.006 at 8-bit and 0.012 at 4-bit.
- One decision takes about 0.2 s on an M1 Max.
A note for anyone else converting: transformers 5.17 re-saves Gemma 4's global_head_dim andnum_global_key_value_heads as a per_layer_config block, which mlx-lm 0.32 cannot read. After a merge, keep the
original config.json.
Adaptive thinking is not ported. If a link from your card would be useful, feel free to add one.
thank you very much
Hi @Avicennasis
Thank you for the MLX builds of GEV-26B-Decide. Apple silicon was a gap in our release, and you filled it carefully: merge checks that confirm the adapter actually loaded, the head and calibration files left untouched, and a real parity run against our transformers + peft path rather than a smoke test.
The results are what we would hope for. At about 0.2 s per decision on an M1 Max, System 1 becomes practical on a laptop for exactly the kind of guardrail and triage work you tested it on.
Your config.json note will save others real time. We'll link both builds from the GEV-26B-Decide card, along with that tip.
We would welcome more collaboration. A few areas where help would matter most:
- Adaptive thinking (System 2) on MLX, so the full System 1 + System 2 path runs locally
- The vision tower, for multimodal decisions on Apple silicon
- Batched decisions in gev_mlx.py
If any of the decisions the reference got wrong on your set can be shared, even redacted, that would be valuable input for the next version. We're also glad to give you a heads-up before future releases so MLX builds can land alongside them.
Reply here to Cloud Yu (Yu Hai) or me if you'd like to work together more closely. Thanks again.
Josh Liu
Thanks Josh and Cloud! 😺 I'd be happy to help with more of this. Batching looks like the best place to start, and it'll be useful for my own evaluation runs too.
For System 2, there's one wrinkle: the current MLX builds have the System 1 adapter merged in, so they can't switch back to the unadapted backbone the way your reference server does. I'd like to work out that separation properly rather than just enable generation on the merged weights.
I'm interested in vision as well. mlx-vlm already has Gemma 4 support, so there's a starting point, though the decision-head integration and image parity still need testing. My current plan is batching first, then System 2, then vision, without putting dates on those yet.
Where would you prefer the code contributions to land? If you have reference cases you'd like us to use for batching or multimodal checks, those would help. I'd also appreciate a heads-up on future releases.
I'll go through the cases the reference got wrong and see which I can share in sanitized form.