fix: declare NextN (MTP) layer 45 as BF16 in quantization ignore lists
Credit: the diagnosis and the fix come from @jon1012 in discussion #6. This PR applies that fix verbatim; we only add an independent serving-level verification.
Problem
The NextN/MTP layer (model.language_model.layers.45, the single num_nextn_predict_layers=1 layer) ships entirely in BF16 — 889 tensors, 888 BF16 + 1 F32, with no weight_scale/weight_scale_2/input_scale entries — but config_groups.group_0 covers all Linear modules with NVFP4, and neither quantization_config.ignore (config.json) nor quantization.exclude_modules (hf_quant_config.json) mentions layer 45.
Engines that allocate parameters per the declared config therefore allocate NVFP4-packed expert parameters for the MTP drafter and fail to load with speculative decoding enabled:
RuntimeError: The size of tensor a (256) must match the size of tensor b (512) at non-singleton dimension 1
(TP4: 512-wide BF16 shard vs 256-wide packed-NVFP4 parameter.) Non-speculative serving is unaffected — the layer is only materialized as the draft model.
Fix (metadata only, weights untouched)
Append two wildcard patterns to both ignore lists:
model.language_model.layers.45*model.layers.45*
Both spellings are required: vLLM builds the MTP drafter as a standalone text-only model and strips the model.language_model. prefix, so the first pattern alone is a silent no-op for the drafter.
Verification
With exactly this patch, MTP3 loads and serves correctly on a glm5_next-capable vLLM build (4× RTX PRO 6000 Blackwell, TP4): a full 8000-input/1000-output-token concurrency matrix (C1–C32, 4 prompts per level, cold cache per level) completes with zero failed requests; the server logs MTP acceptance rate ~50–56% and acceptance length ~2.69.
Closing in favor of a fresh PR: the initial commit here reformatted the JSON files wholesale (indent change), burying the actual two-line fix. Superseded by a PR with a minimal, formatting-preserving diff.