Kimi-K3-8L-dummy
An 8-layer, config-only variant of moonshotai/Kimi-K3 for testing and debugging inference engines. It contains no weights.
Usage
Load it with random weights, e.g. in vLLM:
vllm serve riverclouds/Kimi-K3-8L-dummy \
--trust-remote-code \
--load-format dummy \
--tensor-parallel-size 4
Outputs are meaningless. It is meant for exercising the real Kimi-K3 architecture (KDA linear attention + MLA, MoE, KV-cache layout, scheduling, prefix caching) at a fraction of the size.
Changes from the official config
Only three fields in text_config differ from moonshotai/Kimi-K3's
config.json:
| Field | Kimi-K3 | This repo |
|---|---|---|
num_hidden_layers |
93 | 8 |
linear_attn_config.kda_layers |
[1, 2, 3, 5, ..., 91] |
[1, 2, 3, 5, 6, 7] |
linear_attn_config.full_attn_layers |
[4, 8, ..., 92, 93] |
[4, 8] |
This keeps the 3:1 KDA-to-MLA ratio, with a full-attention last layer as in the
original. All other fields (including max_position_embeddings, auto_map,
and vision_config) are unchanged. Tokenizer, processor, generation config and
remote code files are copied verbatim from the official repo.
License
Distributed under the original Kimi K3 License from Moonshot AI.
- Downloads last month
- 1,498
Model tree for riverclouds/Kimi-K3-8L-dummy
Base model
moonshotai/Kimi-K3