XYZ-Aquila-mini-OptiQ-4bit

An OptiQ mixed-precision quantization of XYZAILab/XYZ-Aquila-mini, built for Apple Silicon with MLX.

OptiQ measures each layer's sensitivity and assigns a per-layer bit-width (4-bit or 8-bit here) instead of quantizing every layer the same, so the layers that need the precision keep it.

  • Method: OptiQ data-driven mixed precision (optiq_mixed_precision)
  • Average: ~5.0 bits per weight (94 layers at 8-bit, 296 at 4-bit)
  • Architecture: qwen3_5_moe sparse MoE
  • Reasoning: append :think / :no-think to the model id to control the thinking channel

Use

pip install mlx-optiq
optiq serve --model mlx-community/XYZ-Aquila-mini-OptiQ-4bit

This is a large mixture-of-experts model. optiq serve streams the experts off SSD automatically, so it runs with a small resident footprint (about 4 GB here) rather than loading all weights at once. See mlx-optiq.com for the CLI, the local Lab UI, and the OptiQ Code agent.

Downloads last month
21
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/XYZ-Aquila-mini-OptiQ-4bit

Quantized
(1)
this model