Download serving/README.md from awai-network/basho: direct link, hf CLI and curl.
- Browser
- Download file 3.78 kB
-
https://huggingface.co/awai-network/basho/resolve/main/serving/README.md
- Command line
-
hf download hf://awai-network/basho/serving/README.md
-
curl -L -o README.md https://huggingface.co/awai-network/basho/resolve/main/serving/README.md
serving/ — the matched-pair runtime
This checkpoint stores attention, shared/dense-MLP, and lm_head weights as block-FP8
(F8_E4M3 + FP32 weight_scale_inv). The public glm5_next vLLM images cannot load that:
they build every attention layer unquantized (the official checkpoints keep attention BF16,
so the model code strips quant_config at layer construction), and they have no FP8
ParallelLMHead loader. The result on an unpatched image is a KeyError on the first attention
scale tensor, or a vocab-embedding shape assert.
The fix is four files on top of cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1
(vLLM 0.1.dev20051+g487ecf187), applied by the Dockerfile here:
| file | target path in image | semantic change |
|---|---|---|
kda.py |
vllm/models/glm5next/nvidia/kda.py |
remove the save/None/restore strip of vllm_config.quant_config around KDA base-class construction (lines 168–175 upstream) — the layer now sees the real quant config; the checkpoint's ignore list still keeps anything BF16 that should stay BF16 |
model.py |
vllm/models/glm5next/nvidia/model.py |
Glm5NextDecoderLayer.__init__: MLA attention gets quant_config=quant_config instead of quant_config=None (line 329 upstream); one site covers the 11 MLA layers and the MTP draft layer. The vision-tower quant_config=None site is deliberately untouched (upstream warns quantizing the tower NaNs image features; this checkpoint ships the tower BF16) |
modelopt.py |
vllm/model_executor/layers/quantization/modelopt.py |
ModelOptMixedPrecisionConfig: FP8_BLOCK128 / FP8_BLOCK64 / FP8_BLOCK32 per-layer dispatch to dynamic-activation block Fp8LinearMethod, fused-module name resolution, MTP draft-prefix aliases, and a ParallelLMHead method whose weight_scale_inv loader shards scale rows by vocab shard / 128 (exact: 154880 and the 77440-row TP2 shard divide by 128) |
configs/N=12576,K=4096,…block_shape=[32,32].json |
vllm/model_executor/layers/quantization/utils/configs/ |
autotuned triton tile configs for the KDA fused in_proj GEMM on RTX PRO 6000 Blackwell (SM120) — vLLM ships zero SM120 block-FP8 configs; the default config costs 313 µs/call vs 22.3 µs tuned (43.8% of decode GPU time → ~3%) |
Safety property (measured): this patched image is bitwise identical to the vendor image for
BF16-attention checkpoints — the parent NVFP4 checkpoint produces 100.00% teacher-forced
agreement with |Δlogprob| = 0.00000 under it, because its quantization_config.ignore already
excludes every attention module. The patch only allows quantized attention when a checkpoint's
manifest asks for it.
All three .py files are Apache-2.0 vLLM-tree files (SPDX headers retained) carrying local
modifications; they are not MIT-licensed model weights. Credit: the vLLM project; the
official glm5_next per-model image as packaged by
chriswritescode-dev/glm-5.3-flash-sm120
(our base-image lineage). The base image is pulled from Docker Hub at build time and is not
redistributed in this repo.
Parallel-work credit: local-inference-lab
independently landed the equivalent MLA quant-config passthrough in their public vLLM fork
(dev/jovian-judgement@8590bf9c, 2026-08-29), paired with MXFP8 attention via the b12x SM120
kernels. This patch set does the same for the official per-model-image code path and adds the
block-[32,32] FP8 dispatch for the fused KDA in_proj plus the block-FP8 lm_head loader.
Numerics note for the [32,32] block: the tuned configs keep BLOCK_SIZE_K=32, the same k-split
as the default config, so tuned vs untuned kernels are bit-identical in reduction order.