Qwen3.8 Flash Next — AutoRound 3 bpw + MTP

Mixed-precision quantization of Qwen/Qwen3.8-Flash-Next: 2/3-bit routed experts (~2.998 effective bpw), 8-bit main projections, BF16 sensitive paths, and 4-bit MTP experts. All 13 safetensors shards are required.

KLD vs original weights: 0.1208.

GitHub: full vLLM patch and setup guide

SSD offloading with vLLM

The 95.37 GiB BF16 PLE table stays on SSD, using asynchronous Linux AIO; loaded model weights use approximately 47.3 GiB VRAM. Tested on a 64 GiB SM80 GPU at 180 W with 15 GiB RAM, 31 GiB swap, and enterprise SATA storage.

Apply the complete patch to vLLM a5a30471ff2bb7f0824f2da10e358af98d304472 and build ple_ssd_io.so. Required additions: metadata-only PLE loading, async row reads/read-ahead, CUDA graph host breaks, mixed-bit expert overrides, and MTP layer remapping. Installation and screen guide uses ~/vllm/.venv.

Performance

Workload Throughput
Fresh 8K prefill 2,636 prompt tok/s
Single-request decode 111 tok/s
Aggregate output, 16 requests 667 tok/s

MTP=1, 2K prefill chunks. Three-round medians; chat uses 128 output tokens, temperature zero, thinking disabled, and EOS ignored. Aggregate throughput includes prefill; decode excludes first-token latency. Full measurements

Experimental runtime: an arithmetic consistency issue remains. Context is advertised as 262K; varied-token timing was tested through 32K.

License

Qwen Community License 1.0, inherited from the base model.

Downloads last month
670
Safetensors
Model size
54B params
Tensor type
I64
·
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for klee100/Qwen3.8-Flash-Next-AutoRound-3bpw-MTP

Quantized
(273)
this model