Qwen3.5-2B OneCompression 4-bit MLX

Text-only MLX checkpoint produced from Qwen/Qwen3.5-2B for AyaneSDK's on-device conversation example.

  • OneCompression GPTQ with quantization-error propagation (QEP)
  • 4-bit weights, group size 128
  • 256 Japanese dialogue calibration samples of 512 tokens
  • 186 quantized linear layers
  • Token embedding quantized separately to asymmetric MLX 4-bit
  • Vision weights are not included

The packed GPTQ-v1 linear weights were converted losslessly to MLX's row-major affine representation. The model configuration retains an onecompression_source_quantization audit record.

Source

Base model: Qwen/Qwen3.5-2B

Quantizer: FujitsuResearch/OneCompression

Downloads last month
114
Safetensors
Model size
2B params
Tensor type
F32
路
U32
路
BF16
路
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(475)
this model