RedHatAI/Kimi-K3-NVFP4

This is a quantized version of moonshotai/Kimi-K3 with MoE layers quantized to NVFP4 for accelerated inference.

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Kimi-K3-NVFP4 \
  --tensor-parallel-size 8 \
  --trust_remote_code \
  --load-format instanttensor \
  --reasoning-parser kimi_k3 \
  --language-model-only  # optional

May require https://github.com/vllm-project/vllm/pull/50500 to run

Creation Process

This model was compressed using LLM Compressor. For more information, please contact Kyle Sayers via the vLLM Developers Slack or ksayers@redhat.com.

Evaluation

Benchmark moonshotai/Kimi-K3 RedHatAI/Kimi-K3-NVFP4
GPQA 93.5 91.0
inspect eval hf/Idavidrein/gpqa/diamond \
  --model vllm/RedHatAI/Kimi-K3-NVFP4 \
  --reasoning-effort high \
  --model-base-url http://localhost:8000/v1
Downloads last month
196
Safetensors
Model size
1.6T params
Tensor type
BF16
·
U8
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Kimi-K3-NVFP4

Quantized
(25)
this model