Kimi-K3-8L-dummy

An 8-layer, config-only variant of moonshotai/Kimi-K3 for testing and debugging inference engines. It contains no weights.

Usage

Load it with random weights, e.g. in vLLM:

vllm serve riverclouds/Kimi-K3-8L-dummy \
  --trust-remote-code \
  --load-format dummy \
  --tensor-parallel-size 4

Outputs are meaningless. It is meant for exercising the real Kimi-K3 architecture (KDA linear attention + MLA, MoE, KV-cache layout, scheduling, prefix caching) at a fraction of the size.

Changes from the official config

Only three fields in text_config differ from moonshotai/Kimi-K3's config.json:

Field Kimi-K3 This repo
num_hidden_layers 93 8
linear_attn_config.kda_layers [1, 2, 3, 5, ..., 91] [1, 2, 3, 5, 6, 7]
linear_attn_config.full_attn_layers [4, 8, ..., 92, 93] [4, 8]

This keeps the 3:1 KDA-to-MLA ratio, with a full-attention last layer as in the original. All other fields (including max_position_embeddings, auto_map, and vision_config) are unchanged. Tokenizer, processor, generation config and remote code files are copied verbatim from the official repo.

License

Distributed under the original Kimi K3 License from Moonshot AI.

Downloads last month
1,498
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for riverclouds/Kimi-K3-8L-dummy

Finetuned
(52)
this model