First MLX port

#2
by Strikesure5555 - opened

There's no MLX/Apple Silicon support for this model anywhere, at least not that I could find (I accept the possibility of pilot error — probably why I am not a pilot). The official GGUFs run only through Flower Labs' own llama.cpp fork, since the architecture isn't in mainline llama.cpp. I built a native MLX port instead: real architecture implementation, not a trust_remote_code workaround.

One thing worth knowing if you're loading the original repo directly: trust_remote_code requires transformers>=5.4.0 — works fine from there through current, but breaks on anything older (tokenizer/cache-API issues below 5.4).

Validated against the reference PyTorch implementation: fp32 layerwise parity, bf16 agreement, greedy-decode exact match, KV-cache self-consistency. Plus negative controls, confirming the checks actually catch a wrong implementation.

Weights:

Sign up or log in to comment