Instructions to use RedHatAI/Kimi-K3-INT4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RedHatAI/Kimi-K3-INT4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="RedHatAI/Kimi-K3-INT4", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("RedHatAI/Kimi-K3-INT4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use RedHatAI/Kimi-K3-INT4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RedHatAI/Kimi-K3-INT4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/Kimi-K3-INT4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/RedHatAI/Kimi-K3-INT4
- SGLang
How to use RedHatAI/Kimi-K3-INT4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "RedHatAI/Kimi-K3-INT4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/Kimi-K3-INT4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "RedHatAI/Kimi-K3-INT4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/Kimi-K3-INT4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use RedHatAI/Kimi-K3-INT4 with Docker Model Runner:
docker model run hf.co/RedHatAI/Kimi-K3-INT4
Kimi-K3-INT4
Model Overview
- Model Architecture: KimiK3ForConditionalGeneration
- Input: Text / Image
- Output: Text
- Model Optimizations:
- Weight quantization: INT4
- Activation quantization: None (original precision)
- Release Date: 2026-09-28
- Version: 1.0
- Model Developers: RedHatAI
This model is a quantized version of moonshotai/Kimi-K3.
Model Optimizations
This model was obtained by quantizing the weights of moonshotai/Kimi-K3 to signed INT4 with group size 128 while keeping activations in their original precision. The weights are stored in the packed compressed-tensors format for vLLM inference.
The quantized weight representation reduces the weight precision from 16 to 4 bits per parameter, reducing quantized weight storage and associated GPU memory requirements by approximately 75%.
The quantization was performed with LLM Compressor. The final checkpoint uses the compressed-tensors pack-quantized representation and is served by vLLM with its WNA16/MARLIN backend.
Deployment
Use with vLLM
Kimi K3 requires substantial hardware. The example below follows the upstream Kimi K3 serving guidance and the configuration used for validation of this checkpoint. Increase --max-model-len if the available memory and workload permit it.
vllm serve RedHatAI/Kimi-K3-INT4 \
--tensor-parallel-size 8 \
--enforce-eager \
--trust-remote-code \
--reasoning-parser kimi_k3 \
--tool-call-parser kimi_k3 \
--enable-auto-tool-choice
The model always has reasoning enabled. For multi-turn conversations, preserve the complete assistant message, including reasoning_content and any tool calls, when passing the response back to the model.
from openai import OpenAI
client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")
response = client.chat.completions.create(
model="RedHatAI/Kimi-K3-INT4",
messages=[{"role": "user", "content": "Explain quantum mechanics clearly."}],
)
print(response.choices[0].message.content)
Creation
The checkpoint was produced with layerwise compression and decompression.
The core quantization configuration was:
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
recipe = QuantizationModifier(
targets="Linear",
scheme="W4A16",
ignore=[
"lm_head",
r"re:.*block_sparse_moe\.gate",
r"re:.*vision_tower.*",
],
)
oneshot(
model=model,
tokenizer=processor.tokenizer,
dataset="perfectblend",
splits="train[:512]",
recipe=recipe,
max_seq_length=2048,
num_calibration_samples=512,
trust_remote_code_model=True,
pipeline="sequential",
sequential_targets=["KimiDecoderLayer"],
sequential_targets_per_subgraph=4,
layerwise_decompression=True,
layerwise_compression=True,
)
Evaluation
This model was evaluated on MATH-500 and GPQA Diamond using lighteval, with the model served through vLLM's OpenAI-compatible API. The baseline was evaluated with the same protocol. These are preliminary smoke-test results, not full benchmark scores.
Accuracy
| Category | Benchmark | moonshotai/Kimi-K3 | RedHatAI/Kimi-K3-INT4 | Recovery |
|---|---|---|---|---|
| Reasoning | MATH-500 (0-shot, pass@1 | 90% | 87% | 97% |
| GPQA Diamond (0-shot, pass@1) | 91% | 92% | 101% |
License
The model is provided under the Kimi K3 license. Review the original model license and terms before use or redistribution.
- Downloads last month
- -
Model tree for RedHatAI/Kimi-K3-INT4
Base model
moonshotai/Kimi-K3