NextN (MTP) layer stored in BF16 but missing from quantization_config.ignore — speculative decoding fails to load
Summary
All weights of the NextN layer (model.language_model.layers.45, the single MTP layer per text_config.num_nextn_predict_layers=1) are stored in BF16, but the quantization metadata does not declare this. As a result, runtimes that allocate parameters per the declared quantization config fail to load the model's MTP draft for speculative decoding.
Details
config.json → quantization_config:
quant_algo: NVFP4,quant_method: modeloptconfig_groups.group_0targets["Linear"]with 4-bit float weights, group_size 16 — i.e. every Linear module defaults to NVFP4ignorehas 132 entries, none covering layer 45 (...layers.45*/ NextN / MTP patterns are absent)
The legacy hf_quant_config.json (quantization.exclude_modules) has the same omission.
But the stored tensors for layer 45 are BF16 throughout, e.g.:
model.language_model.layers.45.mlp.experts.0.down_proj.weight→BF16 [4096, 2048]inmodel-00002-of-00033.safetensors- the weight index contains no
weight_scaleentries for any layer-45 expert (unlike the main MoE layers, which carry NVFP4 scales)
Effect
An engine that follows the declared config allocates NVFP4-packed expert parameters for the MTP draft model, then cannot copy the BF16 checkpoint tensors. On TP4 the per-rank shard mismatch is 512 (BF16) vs 256 (NVFP4-packed):
RuntimeError: The size of tensor a (256) must match the size of tensor b (512) at non-singleton dimension 1
... vllm/models/glm5next/nvidia/mtp.py:493 load_weights → routed_experts._load_w2
Repro: serve with --speculative-config '{"method":"mtp","num_speculative_tokens":3}' on a glm5_next-capable vLLM build. Plain (non-speculative) serving is unaffected, since the NextN layer is only instantiated for speculation.
Workaround (metadata-only, verified)
Extending the ignore list to cover the whole NextN layer — model.language_model.layers.45* and the draft-tree form model.layers.45* — in both config.json's quantization_config.ignore and hf_quant_config.json's exclude_modules makes MTP load and serve correctly with the unmodified official weights (measured: MTP3 acceptance length ~2.69 in serving). The weights themselves are fine; the declaration is incomplete.
Ask
Could you re-export (or patch the metadata) so the NextN layer is properly declared — either listed in ignore/exclude_modules as BF16, or stored quantized with scales like the main MoE experts? And is MTP/speculative decoding intended to be supported on this artifact?
Tested against revision 09b04e5e (2026-09-11).