Qwen3.6-27B Mini - IQ4_XS (GGUF)

An optimally sized quantized version of Qwen3.6-27B was created to fit 16GB with usable MTP.

Model Details

Quick test

  • Wikitext-2-raw PPL: 7.0516 ± 0.04664
  • Decoding speed RTX4070 Ti Super : 80 t/s
  • Chessboard Test 😃

image

Hardware Requirements

  • Fully fits on 16GB VRAM with 92K context using MTP + q4_0 KV cache.
  • Can push even higher context by reducing KV cache further with TurboQuant or Kvarn.

Summary of tensor counts and bpw per qtype

QTYPE		Count	BPW	Assigned GiB	% Assigned	Max GiB (all)
+f32       	353	32    	  0.01 GiB	-		-
q8_0      	6  	8.5   	  0.00 GiB	 0.01%		26.61
q6_K      	101	6.5625	  0.06 GiB	 0.30%		20.55
q5_1      	0  	6     	  0.00 GiB	 0.00%		18.78
q5_K      	20 	5.5   	  0.08 GiB	 0.49%		17.22
q5_0      	0  	5.5   	  0.00 GiB	 0.00%		17.22
q4_1      	0  	5     	  0.00 GiB	 0.00%		15.65
q4_K      	0  	4.5   	  0.00 GiB	 0.00%		14.09
q4_0      	0  	4.5   	  0.00 GiB	 0.00%		14.09
iq4_nl    	0  	4.5   	  0.00 GiB	 0.00%		14.09
iq4_xs    	282	4.25  	  8.86 GiB	66.62%		13.31
q3_K      	0  	3.4375	  0.00 GiB	 0.00%		10.76
iq3_s     	89 	3.4375	  3.51 GiB	32.59%		10.76
iq3_xxs   	0  	3.0625	  0.00 GiB	 0.00%		9.59
q2_K      	0  	2.625 	  0.00 GiB	 0.00%		8.22
iq2_xs    	0  	2.3125	  0.00 GiB	 0.00%		7.24
iq2_xxs   	0  	2.0625	  0.00 GiB	 0.00%		6.46
iq1_m     	0  	1.75  	  0.00 GiB	 0.00%		5.48
iq1_s     	0  	1.5625	  0.00 GiB	 0.00%		4.92

Average BPW: 4.0012
Downloads last month
5,462
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tooltd/Qwen3.6-27B-mini-IQ4-XS-MTP-16GB-VRAM-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(647)
this model