Why no mlx inference support for decision models like this?

#3
by Narutoouz - opened
  • please make a mlx model for this and share it mlx inference engines to support this ideal mlx model made by mlx community
AutoTrust AI Lab org

https://huggingface.co/autotrust/GEV-26B-Decide/discussions/1 @Narutoouz - @Avicennasis just provided a great MLX build for this decision model. We are short of hands and would appreciate if the community can contribute to this MLX gap. BTW, we are just releasing a quantized version of GEV-26B-Decide https://huggingface.co/autotrust/GEV-26B-Decide-NVFP4. You can try this out as well.

@Narutoouz Related, though not a JEV build: I built Seb-9B, a smaller (9B) decision model that already runs on Apple silicon through mlx-lm: https://huggingface.co/ironbcc/seb-9b

The card's MLX snippet reads the probability of each option key from a single forward pass. On an M5 Max, text decisions measured p50 88 ms in bf16 and 104 ms with an 8-bit copy (8.9 GB). That MLX snippet covers text; for image decisions on a Mac, the GGUF build with its vision projector in llama.cpp measured p50 255 ms on 100 image rows. I haven't compared it head-to-head with JEV-27B-VL, so this is an option for local Mac use, not a replacement claim.

Sign up or log in to comment