basho / serving /README.md
com-junkawasaki's picture
Add pinned GLM-5.3 vLLM runtime patches
9e9e40a verified
|
Raw History Blame Contribute Delete
3.78 kB

serving/ — the matched-pair runtime

This checkpoint stores attention, shared/dense-MLP, and lm_head weights as block-FP8 (F8_E4M3 + FP32 weight_scale_inv). The public glm5_next vLLM images cannot load that: they build every attention layer unquantized (the official checkpoints keep attention BF16, so the model code strips quant_config at layer construction), and they have no FP8 ParallelLMHead loader. The result on an unpatched image is a KeyError on the first attention scale tensor, or a vocab-embedding shape assert.

The fix is four files on top of cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1 (vLLM 0.1.dev20051+g487ecf187), applied by the Dockerfile here:

file target path in image semantic change
kda.py vllm/models/glm5next/nvidia/kda.py remove the save/None/restore strip of vllm_config.quant_config around KDA base-class construction (lines 168–175 upstream) — the layer now sees the real quant config; the checkpoint's ignore list still keeps anything BF16 that should stay BF16
model.py vllm/models/glm5next/nvidia/model.py Glm5NextDecoderLayer.__init__: MLA attention gets quant_config=quant_config instead of quant_config=None (line 329 upstream); one site covers the 11 MLA layers and the MTP draft layer. The vision-tower quant_config=None site is deliberately untouched (upstream warns quantizing the tower NaNs image features; this checkpoint ships the tower BF16)
modelopt.py vllm/model_executor/layers/quantization/modelopt.py ModelOptMixedPrecisionConfig: FP8_BLOCK128 / FP8_BLOCK64 / FP8_BLOCK32 per-layer dispatch to dynamic-activation block Fp8LinearMethod, fused-module name resolution, MTP draft-prefix aliases, and a ParallelLMHead method whose weight_scale_inv loader shards scale rows by vocab shard / 128 (exact: 154880 and the 77440-row TP2 shard divide by 128)
configs/N=12576,K=4096,…block_shape=[32,32].json vllm/model_executor/layers/quantization/utils/configs/ autotuned triton tile configs for the KDA fused in_proj GEMM on RTX PRO 6000 Blackwell (SM120) — vLLM ships zero SM120 block-FP8 configs; the default config costs 313 µs/call vs 22.3 µs tuned (43.8% of decode GPU time → ~3%)

Safety property (measured): this patched image is bitwise identical to the vendor image for BF16-attention checkpoints — the parent NVFP4 checkpoint produces 100.00% teacher-forced agreement with |Δlogprob| = 0.00000 under it, because its quantization_config.ignore already excludes every attention module. The patch only allows quantized attention when a checkpoint's manifest asks for it.

All three .py files are Apache-2.0 vLLM-tree files (SPDX headers retained) carrying local modifications; they are not MIT-licensed model weights. Credit: the vLLM project; the official glm5_next per-model image as packaged by chriswritescode-dev/glm-5.3-flash-sm120 (our base-image lineage). The base image is pulled from Docker Hub at build time and is not redistributed in this repo.

Parallel-work credit: local-inference-lab independently landed the equivalent MLA quant-config passthrough in their public vLLM fork (dev/jovian-judgement@8590bf9c, 2026-08-29), paired with MXFP8 attention via the b12x SM120 kernels. This patch set does the same for the official per-model-image code path and adds the block-[32,32] FP8 dispatch for the fused KDA in_proj plus the block-FP8 lm_head loader.

Numerics note for the [32,32] block: the tuned configs keep BLOCK_SIZE_K=32, the same k-split as the default config, so tuned vs untuned kernels are bit-identical in reduction order.