Instructions to use slippedJim/Kimi-K3-Draft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use slippedJim/Kimi-K3-Draft with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("slippedJim/Kimi-K3-Draft", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Kimi-K3 DSpark Draft (ATOM / ROCm)
An MLA-native DSpark draft model for Kimi-K3, served through ATOM's
dspark speculative method on 8 x MI355X. 5 dense layers with non-causal
attention drafting 7 tokens per pass, a low-rank sequential Markov head, and a
confidence head. 68 tensors, 3,562,312,961 parameters, bf16. Serving window
32768.
Performance
tok/fwd = tokens per target forward step = 1 + accepted draft tokens / forward steps. It includes the bonus token the target emits every step, so it
is 1.0 at zero acceptance and capped at 8 when drafting 7 tokens. Same scale as
the acceptance length on
Inferact/Kimi-K3-DSpark's card.
/!\ The three readings are not all on the same stack
The reference card's numbers were measured on vLLM + GB300. Everything in the two middle columns was measured here, on ATOM + 8 x MI355X, under the protocol below.
To show what that difference is worth, we downloaded the reference draft and re-measured it here too: it averages 3.6398 against 3.9023 on its own card, ranging per benchmark from 76.2% (HumanEval) to 116.3% (SPEED-Bench low-entropy). There is no single constant between the stacks, which is why the reference draft's own same-stack reading is included rather than left out.
| benchmark | prompts | this draft (ATOM) | reference draft (ATOM, same machine) | vs reference (ATOM) | reference card (vLLM) | vs card |
|---|---|---|---|---|---|---|
| GSM8K | 1319 | 4.8176 | 4.9022 | 98.3% | 5.64 | 85.4% |
| MATH-500 | 500 | 3.7630 | 3.4388 | 109.4% | 3.82 | 98.5% |
| AIME 2026 | 30 | 3.0481 | 2.7620 | 110.4% | 2.72 | 112.1% |
| HumanEval | 164 | 4.4131 | 4.0701 | 108.4% | 5.34 | 82.6% |
| MBPP | 256 | 3.9237 | 3.7474 | 104.7% | 4.44 | 88.4% |
| MT-Bench | 80 | 3.3564 | 3.1022 | 108.2% | 3.14 | 106.9% |
| SWE-bench Pro | 128 | 3.5499 | 3.4707 | 102.3% | 3.35 | 106.0% |
| SPEED-Bench coding | 80 | 4.9226 | 4.5745 | 107.6% | 4.38 | 112.4% |
| SPEED-Bench multilingual | 80 | 3.7174 | 3.4759 | 106.9% | 4.21 | 88.3% |
| SPEED-Bench rag | 80 | 3.8543 | 3.5558 | 108.4% | 4.11 | 93.8% |
| SPEED-Bench qa | 80 | 3.4747 | 3.0705 | 113.2% | 3.07 | 113.2% |
| SPEED-Bench writing | 80 | 3.0419 | 2.8200 | 107.9% | 2.79 | 109.0% |
| SPEED-Bench low-entropy (16k) | 512 | 4.5993 | 4.3274 | 106.3% | 3.72 | 123.6% |
| AA-LCR (~95k input) | 100 | not measured | not measured | โ | 3.19 | โ |
| mean (13 measured) | 3.8832 | 3.6398 | 106.7% | 3.9023 | 99.5% |
Mean over the thirteen measured benchmarks: 3.8832, against 3.6398 for the reference draft on the same stack (106.7%) and 3.9023 on its card (99.5%). At num_speculative_tokens=7 that mean corresponds to an acceptance rate of 41.19%, against 37.71% for the reference draft on the same stack.
vs reference (ATOM) is the like-for-like number -- same node, image, protocol and prompt counts. vs card crosses stacks, and that difference is not small: the reference draft reads 93.3% of its own card here, ranging per benchmark from 76.2% (HumanEval) to 116.3% (SPEED-Bench low-entropy), with no constant between them.
The card reports a 14-benchmark mean of 3.85. AA-LCR is the one we cannot reproduce: it is measured on 71k-115k-token multi-document prompts and the published dataset exposes only the question text (mean 206 characters), so it is excluded rather than substituted with something shorter.
Raw per-benchmark results, including the full accepted-length histograms for
both drafts, are in benchmarks/ in this repo.
Protocol
Identical for both drafts, every run:
| engine | ATOM, TP=8, 8 x MI355X |
| image | rocm/atom-dev@sha256:2f8bd4206ad15d014ae48115eae1ee9f1db83781848a8542de7177cfbd4ac914 |
--num-speculative-tokens |
7 |
| KV cache | fp8 |
| sampling | greedy (temperature=0, top_p=1.0) |
| concurrency | 1 |
max_tokens |
12288 (8192 for the low-entropy 16k split) |
window (--max-model-len) |
16384 (32768 for the low-entropy 16k split) |
max_num_seqs |
8 |
| prefix caching | off |
| prompts | official count for every benchmark, no sampling |
| counters | read from /debug/mtp_stats, differenced before/after each set |
Acceptance is read from the engine's own counters, not inferred from wall-clock time. The counters are cumulative since server start, so each benchmark is a before/after snapshot difference.
/!\ Why the benchmark window is 16384 and not 32768
The model serves a 32768 window -- that is what the Quick Start above uses. The benchmark runs twelve of the thirteen sets at 16384 for one reason: that is the window the reference draft was measured at, and changing only our side would break the comparison.
It costs nothing, because on those twelve sets the window never binds. The
longest prompt + completion across all of them is 14,994 tokens
(SPEED-Bench writing: 2,699 + 12,295), under the 16384 ceiling. Re-running them
at 32768 returns the same numbers.
max_tokens does bind: eight of the twelve sets have generations that stop at
the 12288 ceiling. Acceptance length is close to insensitive to it -- raising
max_tokens from 2048 to 15360 on an earlier draft lengthened generations by
3.9x and quadrupled forward steps while moving tok/fwd by +0.2% -- but it
is a real ceiling and it is applied identically to both drafts.
The thirteenth set (SPEED-Bench low-entropy) is the only one with long prompts:
inputs average 10.4k tokens and reach 22.5k, so it gets its own server at 32768
with max_num_batched_tokens raised to match.
Quick Start
ATOM pinned by digest: atom 0.1.6rc1.dev275, torch 2.13.0+rocm7.14.0,
HIP 7.14.60850.
docker run -d --name atom-dspark \
--device=/dev/kfd --device=/dev/dri --group-add video \
--security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
--ipc=host --shm-size 128g --network host \
-v /path/to/Kimi-K3:/target:ro -v /path/to/this/repo:/draft:ro \
rocm/atom-dev@sha256:2f8bd4206ad15d014ae48115eae1ee9f1db83781848a8542de7177cfbd4ac914 \
python -m atom.entrypoints.openai_server \
--model /target --served-model-name Kimi-K3 \
--method dspark --draft-model /draft --num-speculative-tokens 7 \
--kv_cache_dtype fp8 -tp 8 --trust-remote-code \
--max-model-len 32768 --max-num-seqs 8 --max-num-batched-tokens 32768 \
--gpu-memory-utilization 0.93 --block-size 128 \
--no-enable_prefix_caching --server-port 8000
The server log should show Detected MLA DSpark drafter and
DSparkProposer aux capture on target layers: (2, 23, 47, 71, 89). Acceptance
counters are at /debug/mtp_stats.
/!\ max_num_batched_tokens must be at least max_model_len: chunked prefill
is off, and a prompt larger than the batched-token budget cannot be scheduled
at all.
/!\ max_num_seqs stays small. K3's KDA recurrent state is allocated per slot;
at 64 the state pool alone asks for 28.9 GB against a ~19 GB KV budget and the
server refuses to start.
Architecture
{
"architectures": ["K3DSparkModel"],
"model_type": "k3_dspark",
"hidden_size": 7168,
"intermediate_size": 14336,
"num_hidden_layers": 5,
"num_attention_heads": 64,
"num_key_value_heads": 64,
"q_lora_rank": 1536,
"kv_lora_rank": 512,
"qk_nope_head_dim": 128,
"qk_rope_head_dim": 64,
"v_head_dim": 128,
"vocab_size": 163840,
"target_num_hidden_layers": 93,
"target_layer_ids": [2, 23, 47, 71, 89],
"markov_rank": 256,
"enable_confidence_head": true,
"max_position_embeddings": 1048576,
"rope_parameters": {
"rope_type": "yarn", "factor": 32.0,
"original_max_position_embeddings": 32768,
"rope_theta": 50000.0, "beta_fast": 32, "beta_slow": 1,
"mscale": 1.0, "mscale_all_dim": 1.0
}
}
/!\ rope_interleave is deliberately absent. ATOM maps that key onto aiter's
is_neox_style, so setting it would select half-split rotation; this draft uses
interleaved pairs. Emitting the key silently changes the attention definition
and collapses acceptance.
Limitations
- AA-LCR is not measured (see above).
- Measured at a 32768 window. The config declares YaRN out to 1048576, but nothing above 32768 has been measured.
- GSM8K is the one benchmark below the reference draft on this stack.
- No training checkpoint is published.
- Downloads last month
- 3
Model tree for slippedJim/Kimi-K3-Draft
Base model
moonshotai/Kimi-K3