fix: declare NextN (MTP) layer 45 as BF16 in quantization ignore lists (#11) da920bb shengliangx lucifer1004 commited on 9 days ago
GLM-5.3-Flash-NVFP4 (NVFP4 routed experts + dense MLP, FP8 KV cache) 9377ce9 shengliangx commited on Sep 4