MTP works after adding layer 45 to the quantization exclusion lists
1
#15 opened 5 days ago
by
soonh1618
Why don't the commits in this repo have the "verified" badge?
1
#14 opened 5 days ago
by
jesang
MTP layer is unquantized but missing from exclude_modules, so vLLM cannot load it for speculative decoding
1
#13 opened 5 days ago
by
brokenlander
Error in function 'TllmGenFmhaRunner' at .sglang/lib/python3.12/site-packages/flashinfer/data/include/flashinfer/trtllm/fmha/fmhaRunner.cuh:37: Unsupported architecture
#12 opened 11 days ago
by
shakhizat
Serving Recipe: GLM-5.3-Flash-NVFP4 with vLLM at 1M Context on Dual DGX Spark
👍 5
#8 opened 20 days ago
by
Pilcothink
[FIX] Reasoning Loop on Trivial Task same as FP8
🔥 1
#7 opened 20 days ago
by
voves
ignore / exclude_modules omits the MTP layer (45): checkpoint fails to load with speculative decoding
❤️ 4
1
#6 opened 21 days ago
by
jon1012
Where's sglang? Do you have a plan for adding Sglang support?
4
#2 opened 23 days ago
by
IlyaTers