Serving Recipe: GLM-5.3-Flash-NVFP4 with vLLM at 1M Context on Dual DGX Spark
#8
by Pilcothink - opened
I got GLM-5.3-Flash-NVFP4 running with vLLM on 2Γ DGX Spark with up to a 1M-token context.
I put the Dockerimage and dual-node recipe here in case it helps anyone else:
https://github.com/gpdev-Pilcothink/DGX_Spark_vllm_Dockerfile/tree/main/0.28/GLM53-flash