Instructions to use PatronusAI/kimi-k3-nvfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PatronusAI/kimi-k3-nvfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PatronusAI/kimi-k3-nvfp4", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("PatronusAI/kimi-k3-nvfp4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PatronusAI/kimi-k3-nvfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PatronusAI/kimi-k3-nvfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/PatronusAI/kimi-k3-nvfp4
- SGLang
How to use PatronusAI/kimi-k3-nvfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PatronusAI/kimi-k3-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PatronusAI/kimi-k3-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use PatronusAI/kimi-k3-nvfp4 with Docker Model Runner:
docker model run hf.co/PatronusAI/kimi-k3-nvfp4
Kimi-K3-NVFP4
NVFP4 conversion of moonshotai/Kimi-K3,
produced by a closed-form, bit-exact MXFP4 → NVFP4 cast with
NVIDIA TensorRT Model Optimizer.
Tooling: patronus-ai/kimi-k3-nvfp4
| Cast fidelity | 85,085,650,944 / 85,085,650,944 blocks bit-exact (100.0000%), max_abs_err 0.0 |
| Layers cast | 247,296 (routed MoE experts) |
| GPQA Diamond | 92.96% (pass@1 avg-of-16) |
| Verified on | SGLang kimi-k3 branch, 2 nodes × 8×B200, TP16 |
| Size | ~1.6 TB, 96 shards |
What this is
| Source (Kimi-K3) | This checkpoint | |
|---|---|---|
| Weight format | MXFP4 (E2M1, block 32, E8M0 scale) | NVFP4 (E2M1, block 16, E4M3 scale + FP32 per-tensor scale) |
| Weight values | — | bit-identical to the source |
| Scope | routed MoE experts (w1/w2/w3) |
routed MoE experts (w1/w2/w3) |
Attention, shared experts, dense MLP projections, lm_head and vision modules are
carried through unquantized.
Bit-exactness
The cast is data-free: NVFP4 scales are derived from the source E8M0 exponents
rather than re-estimated from data. E8M0 has no mantissa, so every source scale is
already an exact power of two. With m = k_max - 8:
weight_scale_2 = 2^m (per tensor, FP32)
per_block_scale = 2^(k_j - m) (per 16 elements, E4M3)
=> product = 2^k_j exactly
E4M3 represents 2^k exactly for k ∈ [-9, 8], so any block within 17 binades of
the tensor max reconstructs bit-for-bit. On Kimi-K3 that covered every block.
No calibration data was used, and none is needed.
Note on precision. Because the per-block scale is always pinned to an exact power of two, this cast deliberately uses none of NVFP4's precision advantages at the weight level — neither the finer block-16 granularity nor E4M3's 3 mantissa bits. That is the cost of bit-exactness, and it is the right trade: quantization is irreversible, so a higher-resolution container cannot recover information MXFP4 already discarded. NVFP4's accuracy edge is only reachable by re-quantizing from BF16, which for 2.8T params means materializing ~5.6 TB. What the conversion buys is access to the NVFP4 kernel and tooling ecosystem.
Deployment (SGLang)
Verified on the SGLang kimi-k3 branch (not main), 2 nodes × 8×B200, TP16.
Run the identical command on every node, varying only --node-rank.
# --- per-node environment ---
export SGLANG_HOST_IP=$(ip -o -4 addr show ens1 | awk '{print $4}' | cut -d/ -f1)
export GLOO_SOCKET_IFNAME=ens1 NCCL_SOCKET_IFNAME=ens1
export NCCL_IB_HCA=mlx5_0,mlx5_1,mlx5_2,mlx5_3,mlx5_4,mlx5_9,mlx5_12,mlx5_13
export NCCL_IB_GID_INDEX=7 NCCL_IB_DISABLE=0 NCCL_NVLS_ENABLE=0 NCCL_CROSS_NIC=0
export PATH=/path/to/venv/bin:$PATH # ninja must resolve from SHARED storage
python -m sglang.launch_server \
--model-path /path/to/kimi-k3-nvfp4 \
--trust-remote-code \
--tp-size 16 --nnodes 2 --node-rank ${NODE_RANK} \
--dist-init-addr ${HEAD_IP}:20000 \
--moe-runner-backend marlin \
--mem-fraction-static 0.85 \
--mamba-full-memory-ratio 7.21 \
--disable-flashinfer-autotune \
--watchdog-timeout 3600 --dist-timeout 3600 \
--reasoning-parser kimi_k3 --tool-call-parser kimi_k3 \
--model-loader-extra-config '{"enable_multithread_load": true}' \
--host 0.0.0.0 --port 30000
SGLANG_HOST_IP is mandatory for multi-node — without it the ranks never rendezvous
and startup hangs with no error.
Five non-obvious requirements
Each of these independently prevents the model from serving:
--moe-runner-backend marlin. Kimi-K3 uses thesitu(SiTuGlu) activation, which is not implemented in SGLang's NVFP4 FlashInfer runner (_SUPPORTED_FP4_ACTIVATIONS = {silu, relu2, gelu}). Marlin supports gated{silu, situ}and forwards both SiTu parameters (gemm1_alpha=β=4.0,clamp_limit=linear_β=25.0).- Patch
modelopt_quant.py:_SUPPORTED_ACT_STRS += ("situ",). The assert at:2552runs before the per-backend dispatch at:2556, so it rejectssitueven when the selected backend supports it. Every backend already applies its own stricter check afterwards. - Raise
UNBALANCED_MODEL_LOADING_TIMEOUT_S(load_model_utils.py, default 480 s, no CLI flag). 1.6 TB over network storage means per-node load skew routinely exceeds 8 minutes and aborts an otherwise healthy run. --dist-timeout 3600. Marlin'sprepare_moe_nvfp4_layer_for_marlinrepacks 247k expert weights, exceeding NCCL's 600 s default; the first rank to finish then times out alone atSeqNum=1 _ALLGATHER_BASE.- Omit
--dcp-sizeunless FlashInfer is new enough —cutedsl_mla_backendpassesenable_dcp/cp_worldtotrtllm_batch_decode_with_kv_cache_mla, which flashinfer 0.6.15 does not accept.
Patches 2–3 are scripted in
serving/patch_situ_nvfp4.py.
Querying
K3 is a reasoning model — it emits thinking tokens before the answer. Use a generous
max_tokens (≥512) and read choices[0].message.content; the thinking is in
reasoning_content. reasoning_effort accepts low / high / max.
curl $URL/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "k3",
"messages": [{"role":"user","content":"What is 17 * 23? Reply with only the number."}],
"max_tokens": 1200, "temperature": 0, "reasoning_effort": "low"}'
With too small a budget you get content: "" and finish_reason: "length" — that is
truncated thinking, not a broken model.
This is a W4A16 deployment
Marlin dequantizes the 4-bit weights to BF16 inside the GEMM, so weights stay 4-bit in HBM (the memory and bandwidth win is kept) but the math runs at BF16 tensor-core rate. Activations are never quantized.
Consequences: no activation calibration is needed or used and this deployment setup does not deliver FP4 compute throughput. Accuracy should be at least that of native MXFP4, since the weights are bit-identical and the activations are higher precision than the source's MXFP8.
Evaluation
GPQA Diamond: 92.96% (pass@1 avg-of-16, sd 1.24).
Per-seed: 96.0 / 93.4 / 91.9 / 91.4 / 92.9 / 93.4 / 90.9 / 93.9 / 92.4 / 91.4 / 93.9 / 92.4 / 92.9 / 93.9 / 93.4 / 92.9.
Harness: nemo-skills with prompt eval/aai/mcq-4choices and
++eval_type=multichoice, matching the
Artificial Analysis methodology;
198 Diamond questions, greedy for seed 0 and temperature=0.6, top_p=0.95 thereafter.
tokens_to_generate=65536 is load-bearing: at 16384 the model exhausted its budget
inside reasoning on ~10% of questions and emitted an empty answer, scoring as wrong
for a purely harness reason — worth ~8.6 points on the greedy pass (87.4% → 96.0%).
Truncation fell from 9.6% to 0.58%.
Layout
Per quantized module <base>:
<base>.weight uint8 (out, in/2) NVFP4-packed E2M1 nibbles
<base>.weight_scale E4M3 (out, in/16) per-block scale
<base>.weight_scale_2 fp32 scalar per-tensor scale
config.json reports format: nvfp4-pack-quantized, quant_method: modelopt,
quant_algo: NVFP4, group_size: 16, and an exclude_modules list of 55 globs
derived from the modules that actually kept an unpacked 2-D float .weight.
Branches.
mainis weight-only (input_activations: null).w4a4addsinput_scaleplaceholders of 1.0; they are unused by the marlin path.
Bundled fix
modeling_kimi_linear.py here includes a fix to KimiDeltaAttention: A_log is
stored at head_dim (128) but the upstream release builds it at num_heads (96), so
from_pretrained fails to load its own weights.
Note: The checkpoint holds 96 trained decay values zero-padded to 128 and verified across
all 69 A_log tensors here: elements [96:128] are exactly zero in every one. The
FLA kernel indexes by num_heads, so the padding is inert. Future readers should refer to
discussion #150 for fixing it.
Known limitation: the FP4 fast path
The trtllm-gen cubin pool ships SiTu kernels at NVFP4's 16-element scale block
(57 × BF16-activation, 45 × E2M1-activation), but they are unreachable from SGLang's
NVFP4 runner, which resolves activations via FlashInfer's ActivationType (no SiTu
member) and reads a different cubin store. Wiring the dispatch to
sglang.kernels.ops.moe.trtllm_gen_moe reaches the JIT build stage and then fails on
CUDA 12.8 with namespace "cuda" has no member "maximum" — cuda::maximum requires
CUDA 13's CCCL. Untested beyond that point.
Reproducing
python cast_mxfp4_to_nvfp4_offline.py \
--source_ckpt /path/to/Kimi-K3 \
--output_ckpt /path/to/Kimi-K3-NVFP4 --verify
Runs on CPU without instantiating the model (~7 h, bounded by shard I/O).
License
Derivative of moonshotai/Kimi-K3; the upstream model's license and terms apply.
- Downloads last month
- 15
Model tree for PatronusAI/kimi-k3-nvfp4
Base model
moonshotai/Kimi-K3