# serving/ — the matched-pair runtime This checkpoint stores attention, shared/dense-MLP, and `lm_head` weights as block-FP8 (`F8_E4M3` + FP32 `weight_scale_inv`). The public `glm5_next` vLLM images cannot load that: they build every attention layer **unquantized** (the official checkpoints keep attention BF16, so the model code strips `quant_config` at layer construction), and they have no FP8 `ParallelLMHead` loader. The result on an unpatched image is a `KeyError` on the first attention scale tensor, or a vocab-embedding shape assert. The fix is four files on top of `cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1` (vLLM `0.1.dev20051+g487ecf187`), applied by the Dockerfile here: | file | target path in image | semantic change | |---|---|---| | `kda.py` | `vllm/models/glm5next/nvidia/kda.py` | remove the save/`None`/restore strip of `vllm_config.quant_config` around KDA base-class construction (lines 168–175 upstream) — the layer now sees the real quant config; the checkpoint's `ignore` list still keeps anything BF16 that should stay BF16 | | `model.py` | `vllm/models/glm5next/nvidia/model.py` | `Glm5NextDecoderLayer.__init__`: MLA attention gets `quant_config=quant_config` instead of `quant_config=None` (line 329 upstream); one site covers the 11 MLA layers and the MTP draft layer. The vision-tower `quant_config=None` site is deliberately untouched (upstream warns quantizing the tower NaNs image features; this checkpoint ships the tower BF16) | | `modelopt.py` | `vllm/model_executor/layers/quantization/modelopt.py` | `ModelOptMixedPrecisionConfig`: `FP8_BLOCK128` / `FP8_BLOCK64` / `FP8_BLOCK32` per-layer dispatch to dynamic-activation block `Fp8LinearMethod`, fused-module name resolution, MTP draft-prefix aliases, and a `ParallelLMHead` method whose `weight_scale_inv` loader shards scale rows by vocab shard / 128 (exact: 154880 and the 77440-row TP2 shard divide by 128) | | `configs/N=12576,K=4096,…block_shape=[32,32].json` | `vllm/model_executor/layers/quantization/utils/configs/` | autotuned triton tile configs for the KDA fused in_proj GEMM on RTX PRO 6000 Blackwell (SM120) — vLLM ships zero SM120 block-FP8 configs; the default config costs 313 µs/call vs 22.3 µs tuned (43.8% of decode GPU time → ~3%) | Safety property (measured): this patched image is **bitwise identical** to the vendor image for BF16-attention checkpoints — the parent NVFP4 checkpoint produces 100.00% teacher-forced agreement with |Δlogprob| = 0.00000 under it, because its `quantization_config.ignore` already excludes every attention module. The patch only *allows* quantized attention when a checkpoint's manifest asks for it. All three `.py` files are Apache-2.0 vLLM-tree files (SPDX headers retained) carrying local modifications; they are **not** MIT-licensed model weights. Credit: the vLLM project; the official `glm5_next` per-model image as packaged by [chriswritescode-dev/glm-5.3-flash-sm120](https://github.com/chriswritescode-dev/glm-5.3-flash-sm120) (our base-image lineage). The base image is pulled from Docker Hub at build time and is not redistributed in this repo. Parallel-work credit: [local-inference-lab](https://github.com/local-inference-lab/vllm) independently landed the equivalent MLA quant-config passthrough in their public vLLM fork (`dev/jovian-judgement@8590bf9c`, 2026-08-29), paired with MXFP8 attention via the b12x SM120 kernels. This patch set does the same for the official per-model-image code path and adds the block-[32,32] FP8 dispatch for the fused KDA in_proj plus the block-FP8 lm_head loader. Numerics note for the [32,32] block: the tuned configs keep `BLOCK_SIZE_K=32`, the same k-split as the default config, so tuned vs untuned kernels are bit-identical in reduction order.