rapidchat MLX 4-bit

Quantized builds of rapidchat, the 9B airline-support fine-tune. These builds were not separately benchmarked: the 86.5 tau2-bench airline pass^1 is for the full bf16 model, and 4-bit quantization usually costs a little accuracy. Training data: rapidchat-data.

For iPhone, iPad and Mac (Apple silicon). Converted with mlx-lm (4-bit); not test-run on Apple hardware yet.

run it

pip install mlx-lm
mlx_lm.generate --model badr7/rapidchat-MLX-4bit --prompt 'Hi, I need to change my flight.'
Downloads last month
19
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for badr7/rapidchat-MLX-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(2)
this model