Dev-4B-MLX-8bit

The 8-bit MLX build of Qwen3-4B-Instruct-2507 that Dev-4B runs on, for Apple-silicon Macs. It is the unchanged base model, quantised: it contains none of Dev-4B's add-ons. Download it alongside Dev-4B to skip downloading the 8 GB original and converting it yourself.

  • Source: Qwen/Qwen3-4B-Instruct-2507 at revision cdbee75f17c01a7cc42f958dc650907174af0554, the revision Dev-4B was trained on.
  • Conversion: mlx_lm convert -q --q-bits 8 --q-group-size 64 with mlx 0.32.3 and mlx-lm 0.32.0 (affine quantisation, group size 64). About 4 GB.
  • Fidelity with Dev-4B's add-ons: on sampled test items, decision probabilities differ from the bf16 PyTorch model by 0.006 on average, with the same accuracy (84.0% vs 84.0% on trained task types).

Use it with Dev-4B:

hf download suhaas-teja/Dev-4B --local-dir Dev-4B && cd Dev-4B
hf download suhaas-teja/Dev-4B-MLX-8bit --local-dir mlx-8bit
pip install -e ".[serve,mac]"
python -m serve.server_mlx --model mlx-8bit --artifacts .

It also works on its own as a plain chat model with mlx_lm.generate --model suhaas-teja/Dev-4B-MLX-8bit.

Licence: Apache-2.0, as the original Qwen3-4B-Instruct-2507. All credit for the model goes to the Qwen team.

Downloads last month
12
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for suhaas-teja/Dev-4B-MLX-8bit

Quantized
(329)
this model