|
Download serving/README.md from awai-network/basho: direct link, hf CLI and curl.
- Browser
- Download file 3.78 kB
-
https://huggingface.co/awai-network/basho/resolve/main/serving/README.md
- Command line
-
hf download hf://awai-network/basho/serving/README.md
-
curl -L -o README.md https://huggingface.co/awai-network/basho/resolve/main/serving/README.md
3.78 kB
| # serving/ — the matched-pair runtime | |
| This checkpoint stores attention, shared/dense-MLP, and `lm_head` weights as block-FP8 | |
| (`F8_E4M3` + FP32 `weight_scale_inv`). The public `glm5_next` vLLM images cannot load that: | |
| they build every attention layer **unquantized** (the official checkpoints keep attention BF16, | |
| so the model code strips `quant_config` at layer construction), and they have no FP8 | |
| `ParallelLMHead` loader. The result on an unpatched image is a `KeyError` on the first attention | |
| scale tensor, or a vocab-embedding shape assert. | |
| The fix is four files on top of `cstechdev/vllm:glm53-flash-nope-sm120-cu130-20260826-r1` | |
| (vLLM `0.1.dev20051+g487ecf187`), applied by the Dockerfile here: | |
| | file | target path in image | semantic change | | |
| |---|---|---| | |
| | `kda.py` | `vllm/models/glm5next/nvidia/kda.py` | remove the save/`None`/restore strip of `vllm_config.quant_config` around KDA base-class construction (lines 168–175 upstream) — the layer now sees the real quant config; the checkpoint's `ignore` list still keeps anything BF16 that should stay BF16 | | |
| | `model.py` | `vllm/models/glm5next/nvidia/model.py` | `Glm5NextDecoderLayer.__init__`: MLA attention gets `quant_config=quant_config` instead of `quant_config=None` (line 329 upstream); one site covers the 11 MLA layers and the MTP draft layer. The vision-tower `quant_config=None` site is deliberately untouched (upstream warns quantizing the tower NaNs image features; this checkpoint ships the tower BF16) | | |
| | `modelopt.py` | `vllm/model_executor/layers/quantization/modelopt.py` | `ModelOptMixedPrecisionConfig`: `FP8_BLOCK128` / `FP8_BLOCK64` / `FP8_BLOCK32` per-layer dispatch to dynamic-activation block `Fp8LinearMethod`, fused-module name resolution, MTP draft-prefix aliases, and a `ParallelLMHead` method whose `weight_scale_inv` loader shards scale rows by vocab shard / 128 (exact: 154880 and the 77440-row TP2 shard divide by 128) | | |
| | `configs/N=12576,K=4096,…block_shape=[32,32].json` | `vllm/model_executor/layers/quantization/utils/configs/` | autotuned triton tile configs for the KDA fused in_proj GEMM on RTX PRO 6000 Blackwell (SM120) — vLLM ships zero SM120 block-FP8 configs; the default config costs 313 µs/call vs 22.3 µs tuned (43.8% of decode GPU time → ~3%) | | |
| Safety property (measured): this patched image is **bitwise identical** to the vendor image for | |
| BF16-attention checkpoints — the parent NVFP4 checkpoint produces 100.00% teacher-forced | |
| agreement with |Δlogprob| = 0.00000 under it, because its `quantization_config.ignore` already | |
| excludes every attention module. The patch only *allows* quantized attention when a checkpoint's | |
| manifest asks for it. | |
| All three `.py` files are Apache-2.0 vLLM-tree files (SPDX headers retained) carrying local | |
| modifications; they are **not** MIT-licensed model weights. Credit: the vLLM project; the | |
| official `glm5_next` per-model image as packaged by | |
| [chriswritescode-dev/glm-5.3-flash-sm120](https://github.com/chriswritescode-dev/glm-5.3-flash-sm120) | |
| (our base-image lineage). The base image is pulled from Docker Hub at build time and is not | |
| redistributed in this repo. | |
| Parallel-work credit: [local-inference-lab](https://github.com/local-inference-lab/vllm) | |
| independently landed the equivalent MLA quant-config passthrough in their public vLLM fork | |
| (`dev/jovian-judgement@8590bf9c`, 2026-08-29), paired with MXFP8 attention via the b12x SM120 | |
| kernels. This patch set does the same for the official per-model-image code path and adds the | |
| block-[32,32] FP8 dispatch for the fused KDA in_proj plus the block-FP8 lm_head loader. | |
| Numerics note for the [32,32] block: the tuned configs keep `BLOCK_SIZE_K=32`, the same k-split | |
| as the default config, so tuned vs untuned kernels are bit-identical in reduction order. | |