Support for NVFP4 kv cache

#3
by g-a-b-y - opened

Is it possible to use NVFP4 KV cache to further reduce VRAM usage?

With GLM-5.2 it always failed on vLLM saying Sparse MLA was not supported.

Sign up or log in to comment