GLM-5.3-Vision-FP8

GLM-5.3 with vision: the Kimi K2.6 MoonViT-3d vision tower and the PatchMerger projector trained by Baseten for GLM-5.2-Vision, bolted onto the GLM-5.3 text backbone from zai-org/GLM-5.3. GLM-5.3 shares GLM-5.2's base weights, architecture, and tokenizer, so the projector transfers.

Nothing in the text backbone or the vision tower was modified.

component source status
text backbone zai-org/GLM-5.3 byte-identical
vision tower (MoonViT-3d, 27 layers) moonshotai/Kimi-K2.6 via baseten/GLM-5.2-Vision-NVFP4 byte-identical
projector (PatchMerger MLP, 49.5M) baseten/GLM-5.2-Vision-NVFP4 as trained against GLM-5.2

Results

Not yet evaluated in this precision. The NVFP4 sibling (lunaroute/GLM-5.3-Vision-NVFP4), which uses the identical tower and projector, scores 60.0% on MMMU-Pro standard (4 options), n=300, zero-shot. FP8 should be at least as good; numbers will be added once measured.

Files

If this repo does not (yet) contain model-*.safetensors, download them from the base repo into the same directory:

hf download zai-org/GLM-5.3 --include 'model-*.safetensors' --local-dir ./GLM-5.3-Vision-FP8
hf download lunaroute/GLM-5.3-Vision-FP8 --local-dir ./GLM-5.3-Vision-FP8

Serving (vLLM 0.27.1, tested)

Needs the out-of-tree vllm-glm5v plugin (registers Glm5vForConditionalGeneration via the vllm.general_plugins entry point) and three one-line fixes in vLLM so the sparse-MLA backends read index_topk from hf_text_config instead of hf_config (model_executor/layers/attention/sparse_mla_attention.py x2, v1/attention/backends/mla/flashattn_mla_sparse.py). On B200 also set NCCL_NVLS_ENABLE=0 and VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm in a .pth file (vLLM re-execs workers, so shell exports do not reach them). All of that is plugins/vllm-glm5v/apply_vllm_fixes.sh in this repo; run it inside the vllm/vllm-openai:v0.27.1-x86_64 image.

vllm serve ./GLM-5.3-Vision-FP8 --served-model-name glm-5.3-vision \
  --tensor-parallel-size 4 --quantization modelopt_fp4 --moe-backend flashinfer_trtllm \
  --kv-cache-dtype fp8 --gpu-memory-utilization 0.90 --max-model-len 262144 --trust-remote-code \
  --enable-auto-tool-choice --tool-call-parser glm47 --reasoning-parser glm45 \
  --enable-chunked-prefill --disable-custom-all-reduce --enable-prefix-caching

TP4 on B200 leaves ~46 GiB/GPU for KV (about 1M tokens). TP8 for 1M context. Thinking is controlled by the reasoning_effort chat-template kwarg (low/high/max); enable_thinking is not honoured by the GLM-5.3 template.

Serving (SGLang 0.5.13, Baseten's path, tested)

Same out-of-tree plugin and flags as Baseten's GLM-5.2-Vision (plugins/sglang_glm5v, image lmsysorg/sglang:v0.5.13-cu130). patch.py applies cleanly on 0.5.13. If you run from a pip install instead of that image, FlashInfer JIT needs nvrtc.h and libnvrtc on the CUDA include/lib path (symlink them from the pip nvidia/cu13 package). 4x B200 gives 974k KV tokens at 256k context:

uv pip install ./GLM-5.3-Vision-FP8/plugins
export SGLANG_EXTERNAL_MODEL_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_PROCESSOR_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_MODEL_ARCH=Glm5vForConditionalGeneration
python -m sglang_glm5v.patch

python -m sglang.launch_server \
  --model-path ./GLM-5.3-Vision-FP8 --trust-remote-code \
  --tp-size 8 \
  --disable-shared-experts-fusion --disable-flashinfer-autotune \
  --attention-backend dsa --mm-attention-backend sdpa \
  --kv-cache-dtype fp8_e4m3 --page-size 64 \
  --mem-fraction-static 0.85 \
  --context-length 1048576 \
  --reasoning-parser glm45 --tool-call-parser glm47 \
  --served-model-name glm-5.3-vision --port 30000

License

MIT. GLM-5.3 is MIT (Z.ai). The projector and assembly are MIT (Baseten). The MoonViT tower is from Kimi K2.6 under Moonshot's Modified MIT license, which adds a display-attribution requirement above 100M monthly active users or USD 20M monthly revenue; see LICENSE.

Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lunaroute/GLM-5.3-Vision-FP8

Base model

zai-org/GLM-5.2
Quantized
(8)
this model