ignore / exclude_modules omits the MTP layer (45): checkpoint fails to load with speculative decoding

#6
by jon1012 - opened

Thanks for publishing this — the eval table (lossless on all six, including MMMU-Pro and AA-LCR) is much appreciated.

One packaging issue: the quantization ignore / exclude_modules lists omit the MTP layer (45), so the checkpoint cannot be loaded with speculative decoding enabled.

Symptom

vLLM 0.27.1 (with glm5_next support), TP=4, --speculative-config '{"method":"mtp","num_speculative_tokens":3}', --moe-backend auto on 4× RTX PRO 6000 Blackwell (sm_120):

RuntimeError: The size of tensor a (256) must match the size of tensor b (512) at non-singleton dimension 1
  File "vllm/models/glm5next/nvidia/mtp.py", line 395, in load_weights
  File "vllm/model_executor/layers/fused_moe/routed_experts.py", line 805, in weight_loader
  File "vllm/model_executor/layers/fused_moe/routed_experts.py", line 560, in _load_w2
    expert_data.copy_(loaded_weight)

256 is the packed-NVFP4 width (moe_intermediate_size 2048 / TP 4 = 512, / 2 for uint8 packing); 512 is the unpacked BF16 width. The engine allocated quantized expert parameters for a layer whose tensors are BF16.

Cause

Layer 45 is shipped entirely unquantized — reading the safetensors headers, its 889 tensors are 888 BF16 + 1 F32, with zero weight_scale, weight_scale_2 or input_scale (a genuinely quantized layer, e.g. model.language_model.layers.10.mlp.experts.0.down_proj, carries weight + weight_scale + weight_scale_2 + input_scale).

But no entry in either exclusion list matches layer 45. Both config.json → quantization_config.ignore and hf_quant_config.json → quantization.exclude_modules have 132 entries, and they break down as:

pattern class entries layer ids covered
*.self_attn* 45 0–44
*.mlp.gate 42 3–44
*.mlp.shared_experts* 42 3–44
model.visual* 1 vision tower (correctly handled)
lm_head, model.language_model.embed_tokens 2 —

With num_hidden_layers: 45 and num_nextn_predict_layers: 1, layer 45 is the MTP head. The lists look like they were generated over range(num_hidden_layers) (0–44), so the MTP layer got no entries — while config_groups.group_0 targets all Linear. The vision tower was given an explicit wildcard and is fine; only the MTP layer is affected.

This is also why it can ship unnoticed: layer 45 is only materialised as the draft model, so without speculative decoding the checkpoint loads and serves normally.

Fix

Adding the MTP layer to both lists resolves it completely. Note that two patterns are needed:

"model.language_model.layers.45*",
"model.layers.45*"

The second is the one that actually matters, and the first alone is a silent no-op. vLLM builds the MTP head as a standalone text-only model and strips the multimodal prefix, so the quant-config matcher never sees the language_model spelling for the drafter. From vllm/models/glm5next/nvidia/mtp.py:

Multimodal (Glm5NextForConditionalGeneration) checkpoints prefix the text-tower weights with model.language_model.; the MTP head is built as a text-only model (model.layers.*), so strip the prefix to match.

A wildcard over the whole layer is appropriate here because every layer-45 tensor is BF16 — its self_attn, mlp.gate and shared_experts are mis-declared too, not just the routed experts.

With both patterns added locally, the model loads in ~230 s and selects the expected backends on each side, which confirms the exclusion is being honoured:

Using 'FLASHINFER_CUTLASS' NvFp4 MoE backend        <- the W4A4 routed experts
Using FlashInfer CUTLASS Unquantized MoE backend    <- the BF16 MTP drafter

Happy to provide any further detail.

Confirming this on a second, independent setup (community Jovian Judgement beta 20260917 image, glm53-flash profile, TP4 on 4× RTX PRO 6000).

Same failure with the official export (RuntimeError: ... 256 vs 512 in _load_w2), same resolution: adding both wildcard patterns to the ignore lists fixes it. I additionally verified the fix end-to-end in serving — with the metadata-only patch, MTP3 serves and benches cleanly (full 8000/1000-token concurrency matrix C1–C32, zero failed requests; server-logged MTP acceptance rate ~50–56%, acceptance length ~2.69).

One extra data point on scope: the community image's qualified MTP configuration does not use this checkpoint at all — it binds a quantization-aware-distilled repack (local-inference-lab/GLM-5.3-Flash-NVFP4, MXFP8 MTP experts) and pins the draft MoE to Marlin. So among the artifacts in circulation, MTP on this official export only works via the ignore-list fix described here.

Sign up or log in to comment