basho / serving /README.md
com-junkawasaki's picture
Add pinned GLM-5.3 vLLM runtime patches
9e9e40a verified
|
Raw History Blame Contribute Delete
3.78 kB
# serving/ — the matched-pair runtime
This checkpoint stores attention, shared/dense-MLP, and `lm_head` weights as block-FP8
(`F8_E4M3` + FP32 `weight_scale_inv`). The public `glm5_next` vLLM images cannot load that:
they build every attention layer **unquantized** (the official checkpoints keep attention BF16,
so the model code strips `quant_config` at layer construction), and they have no FP8
`ParallelLMHead` loader. The result on an unpatched image is a `KeyError` on the first attention
scale tensor, or a vocab-embedding shape assert.
The fix is four files on top of `cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1`
(vLLM `0.1.dev20051+g487ecf187`), applied by the Dockerfile here:
| file | target path in image | semantic change |
|---|---|---|
| `kda.py` | `vllm/models/glm5next/nvidia/kda.py` | remove the save/`None`/restore strip of `vllm_config.quant_config` around KDA base-class construction (lines 168–175 upstream) — the layer now sees the real quant config; the checkpoint's `ignore` list still keeps anything BF16 that should stay BF16 |
| `model.py` | `vllm/models/glm5next/nvidia/model.py` | `Glm5NextDecoderLayer.__init__`: MLA attention gets `quant_config=quant_config` instead of `quant_config=None` (line 329 upstream); one site covers the 11 MLA layers and the MTP draft layer. The vision-tower `quant_config=None` site is deliberately untouched (upstream warns quantizing the tower NaNs image features; this checkpoint ships the tower BF16) |
| `modelopt.py` | `vllm/model_executor/layers/quantization/modelopt.py` | `ModelOptMixedPrecisionConfig`: `FP8_BLOCK128` / `FP8_BLOCK64` / `FP8_BLOCK32` per-layer dispatch to dynamic-activation block `Fp8LinearMethod`, fused-module name resolution, MTP draft-prefix aliases, and a `ParallelLMHead` method whose `weight_scale_inv` loader shards scale rows by vocab shard / 128 (exact: 154880 and the 77440-row TP2 shard divide by 128) |
| `configs/N=12576,K=4096,…block_shape=[32,32].json` | `vllm/model_executor/layers/quantization/utils/configs/` | autotuned triton tile configs for the KDA fused in_proj GEMM on RTX PRO 6000 Blackwell (SM120) — vLLM ships zero SM120 block-FP8 configs; the default config costs 313 µs/call vs 22.3 µs tuned (43.8% of decode GPU time → ~3%) |
Safety property (measured): this patched image is **bitwise identical** to the vendor image for
BF16-attention checkpoints — the parent NVFP4 checkpoint produces 100.00% teacher-forced
agreement with |Δlogprob| = 0.00000 under it, because its `quantization_config.ignore` already
excludes every attention module. The patch only *allows* quantized attention when a checkpoint's
manifest asks for it.
All three `.py` files are Apache-2.0 vLLM-tree files (SPDX headers retained) carrying local
modifications; they are **not** MIT-licensed model weights. Credit: the vLLM project; the
official `glm5_next` per-model image as packaged by
[chriswritescode-dev/glm-5.3-flash-sm120](https://github.com/chriswritescode-dev/glm-5.3-flash-sm120)
(our base-image lineage). The base image is pulled from Docker Hub at build time and is not
redistributed in this repo.
Parallel-work credit: [local-inference-lab](https://github.com/local-inference-lab/vllm)
independently landed the equivalent MLA quant-config passthrough in their public vLLM fork
(`dev/jovian-judgement@8590bf9c`, 2026-08-29), paired with MXFP8 attention via the b12x SM120
kernels. This patch set does the same for the official per-model-image code path and adds the
block-[32,32] FP8 dispatch for the fused KDA in_proj plus the block-FP8 lm_head loader.
Numerics note for the [32,32] block: the tuned configs keep `BLOCK_SIZE_K=32`, the same k-split
as the default config, so tuned vs untuned kernels are bit-identical in reduction order.