Kimi-K3-INT4

Model Overview

  • Model Architecture: KimiK3ForConditionalGeneration
  • Input: Text / Image
  • Output: Text
  • Model Optimizations:
    • Weight quantization: INT4
    • Activation quantization: None (original precision)
  • Release Date: 2026-09-28
  • Version: 1.0
  • Model Developers: RedHatAI

This model is a quantized version of moonshotai/Kimi-K3.

Model Optimizations

This model was obtained by quantizing the weights of moonshotai/Kimi-K3 to signed INT4 with group size 128 while keeping activations in their original precision. The weights are stored in the packed compressed-tensors format for vLLM inference.

The quantized weight representation reduces the weight precision from 16 to 4 bits per parameter, reducing quantized weight storage and associated GPU memory requirements by approximately 75%.

The quantization was performed with LLM Compressor. The final checkpoint uses the compressed-tensors pack-quantized representation and is served by vLLM with its WNA16/MARLIN backend.

Deployment

Use with vLLM

Kimi K3 requires substantial hardware. The example below follows the upstream Kimi K3 serving guidance and the configuration used for validation of this checkpoint. Increase --max-model-len if the available memory and workload permit it.

vllm serve RedHatAI/Kimi-K3-INT4 \
  --tensor-parallel-size 8 \
  --enforce-eager \
  --trust-remote-code \
  --reasoning-parser kimi_k3 \
  --tool-call-parser kimi_k3 \
  --enable-auto-tool-choice

The model always has reasoning enabled. For multi-turn conversations, preserve the complete assistant message, including reasoning_content and any tool calls, when passing the response back to the model.

from openai import OpenAI

client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")

response = client.chat.completions.create(
    model="RedHatAI/Kimi-K3-INT4",
    messages=[{"role": "user", "content": "Explain quantum mechanics clearly."}],
)

print(response.choices[0].message.content)

Creation

The checkpoint was produced with layerwise compression and decompression.

The core quantization configuration was:

from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier

recipe = QuantizationModifier(
    targets="Linear",
    scheme="W4A16",
    ignore=[
        "lm_head",
        r"re:.*block_sparse_moe\.gate",
        r"re:.*vision_tower.*",
    ],
)

oneshot(
    model=model,
    tokenizer=processor.tokenizer,
    dataset="perfectblend",
    splits="train[:512]",
    recipe=recipe,
    max_seq_length=2048,
    num_calibration_samples=512,
    trust_remote_code_model=True,
    pipeline="sequential",
    sequential_targets=["KimiDecoderLayer"],
    sequential_targets_per_subgraph=4,
    layerwise_decompression=True,
    layerwise_compression=True,
)

Evaluation

This model was evaluated on MATH-500 and GPQA Diamond using lighteval, with the model served through vLLM's OpenAI-compatible API. The baseline was evaluated with the same protocol. These are preliminary smoke-test results, not full benchmark scores.

Accuracy

Category Benchmark moonshotai/Kimi-K3 RedHatAI/Kimi-K3-INT4 Recovery
Reasoning MATH-500 (0-shot, pass@1 90% 87% 97%
GPQA Diamond (0-shot, pass@1) 91% 92% 101%

License

The model is provided under the Kimi K3 license. Review the original model license and terms before use or redistribution.

Downloads last month
-
Safetensors
Model size
2.8T params
Tensor type
F32
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Kimi-K3-INT4

Quantized
(53)
this model