fix: declare NextN (MTP) layer 45 as BF16 in quantization ignore lists

#10

Credit: the diagnosis and the fix come from @jon1012 in discussion #6. This PR applies that fix verbatim; we only add an independent serving-level verification.

Problem

The NextN/MTP layer (model.language_model.layers.45, the single num_nextn_predict_layers=1 layer) ships entirely in BF16 — 889 tensors, 888 BF16 + 1 F32, with no weight_scale/weight_scale_2/input_scale entries — but config_groups.group_0 covers all Linear modules with NVFP4, and neither quantization_config.ignore (config.json) nor quantization.exclude_modules (hf_quant_config.json) mentions layer 45.

Engines that allocate parameters per the declared config therefore allocate NVFP4-packed expert parameters for the MTP drafter and fail to load with speculative decoding enabled:

RuntimeError: The size of tensor a (256) must match the size of tensor b (512) at non-singleton dimension 1

(TP4: 512-wide BF16 shard vs 256-wide packed-NVFP4 parameter.) Non-speculative serving is unaffected — the layer is only materialized as the draft model.

Fix (metadata only, weights untouched)

Append two wildcard patterns to both ignore lists:

  • model.language_model.layers.45*
  • model.layers.45*

Both spellings are required: vLLM builds the MTP drafter as a standalone text-only model and strips the model.language_model. prefix, so the first pattern alone is a silent no-op for the drafter.

Verification

With exactly this patch, MTP3 loads and serves correctly on a glm5_next-capable vLLM build (4× RTX PRO 6000 Blackwell, TP4): a full 8000-input/1000-output-token concurrency matrix (C1–C32, 4 prompts per level, cold cache per level) completes with zero failed requests; the server logs MTP acceptance rate ~50–56% and acceptance length ~2.69.

lucifer1004 changed pull request status to open

Closing in favor of a fresh PR: the initial commit here reformatted the JSON files wholesale (indent change), burying the actual two-line fix. Superseded by a PR with a minimal, formatting-preserving diff.

lucifer1004 changed pull request status to closed

Sign up or log in to comment