MTP layer is unquantized but missing from exclude_modules, so vLLM cannot load it for speculative decoding

#13
by brokenlander - opened

Thanks for publishing this build, it has been serving us well.

One thing looks like an oversight in the quantization metadata rather than the weights. Layer 45 is
the MTP layer here (num_hidden_layers: 45, num_nextn_predict_layers: 1), and it is unquantized
on disk: 888 BF16 tensors and no weight scales at all, where layer 44 has 864 packed U8 plus 864
F8_E4M3 scales. For example:

layers.44.mlp.experts.0.gate_proj.weight   U8    (2048, 2048)   + weight_scale F8_E4M3 (2048, 256)
layers.45.mlp.experts.0.gate_proj.weight   BF16  (2048, 4096)   no weight_scale

The .quant_summary.txt in this repo agrees, it covers layer indices 0 to 44 and stops there, so I
assume ModelOpt simply never visited the MTP layer.

The problem is that quantization_config.ignore in config.json (and exclude_modules in
hf_quant_config.json) does not mention layer 45 either. Both lists have 132 entries covering
layers 0 to 44. vLLM therefore builds those experts as NVFP4 and fails while loading them:

RuntimeError: The size of tensor a (256) must match the size of tensor b (512) at non-singleton dimension 1

That is in _load_w2, copying BF16 of the logical width into a packed uint8 buffer. The factor of
two is the FP4 packing factor, so it reproduces at every tensor parallel size. Serving without
--speculative-config is unaffected, because nothing touches layer 45.

Adding an entry that covers the MTP layer fixes it, in the same spirit as
nvidia/DeepSeek-R1-0528-FP4, which excludes model.layers.61* wholesale and so never hits this.
Once that layer is excluded the model serves with MTP k=3 on 4x B300 at roughly 2.2x single-stream
decode, with about 2.6 accepted tokens per step on real coding traffic, so the head itself is good.

One wrinkle worth flagging if you do add it. vLLM matches these patterns against its own module
path, not the checkpoint path, and it builds the MTP layer as a separate text-only model rooted at
model. So its experts are model.layers.45.mlp.experts, while everything else in this checkpoint
is spelled model.language_model.layers.*. Glm5NextForConditionalGeneration defines no weights
mapper, so a pattern written in this repo's own spelling will not match the MTP module in vLLM,
whereas the DeepSeek-style model.layers.45* does. You may want both forms, since the right
spelling for the checkpoint and the spelling that takes effect in vLLM are not the same here.

I have also opened vllm-project/vllm#58883 on the vLLM side, so that the failure at least names the
layer and says what to add instead of reporting a bare shape mismatch.

Would you consider adding the MTP layer to the exclude list in this repo? Happy to provide any more
detail that helps.

Thanks @brokenlander , we have accepted community fix for this and the root cause in the ModelOpt library is also clear and have fixed. We'll look into vLLM side issue later, but the model.layers.45* is added for now.

Sign up or log in to comment