You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Kimi-K3 DSpark Draft (ATOM / ROCm)

An MLA-native DSpark draft model for Kimi-K3, served through ATOM's dspark speculative method on 8 x MI355X. 5 dense layers with non-causal attention drafting 7 tokens per pass, a low-rank sequential Markov head, and a confidence head. 68 tensors, 3,562,312,961 parameters, bf16. Serving window 32768.


Performance

tok/fwd = tokens per target forward step = 1 + accepted draft tokens / forward steps. It includes the bonus token the target emits every step, so it is 1.0 at zero acceptance and capped at 8 when drafting 7 tokens. Same scale as the acceptance length on Inferact/Kimi-K3-DSpark's card.

/!\ The three readings are not all on the same stack

The reference card's numbers were measured on vLLM + GB300. Everything in the two middle columns was measured here, on ATOM + 8 x MI355X, under the protocol below.

To show what that difference is worth, we downloaded the reference draft and re-measured it here too: it averages 3.6398 against 3.9023 on its own card, ranging per benchmark from 76.2% (HumanEval) to 116.3% (SPEED-Bench low-entropy). There is no single constant between the stacks, which is why the reference draft's own same-stack reading is included rather than left out.

benchmark prompts this draft (ATOM) reference draft (ATOM, same machine) vs reference (ATOM) reference card (vLLM) vs card
GSM8K 1319 4.8176 4.9022 98.3% 5.64 85.4%
MATH-500 500 3.7630 3.4388 109.4% 3.82 98.5%
AIME 2026 30 3.0481 2.7620 110.4% 2.72 112.1%
HumanEval 164 4.4131 4.0701 108.4% 5.34 82.6%
MBPP 256 3.9237 3.7474 104.7% 4.44 88.4%
MT-Bench 80 3.3564 3.1022 108.2% 3.14 106.9%
SWE-bench Pro 128 3.5499 3.4707 102.3% 3.35 106.0%
SPEED-Bench coding 80 4.9226 4.5745 107.6% 4.38 112.4%
SPEED-Bench multilingual 80 3.7174 3.4759 106.9% 4.21 88.3%
SPEED-Bench rag 80 3.8543 3.5558 108.4% 4.11 93.8%
SPEED-Bench qa 80 3.4747 3.0705 113.2% 3.07 113.2%
SPEED-Bench writing 80 3.0419 2.8200 107.9% 2.79 109.0%
SPEED-Bench low-entropy (16k) 512 4.5993 4.3274 106.3% 3.72 123.6%
AA-LCR (~95k input) 100 not measured not measured โ€” 3.19 โ€”
mean (13 measured) 3.8832 3.6398 106.7% 3.9023 99.5%

Mean over the thirteen measured benchmarks: 3.8832, against 3.6398 for the reference draft on the same stack (106.7%) and 3.9023 on its card (99.5%). At num_speculative_tokens=7 that mean corresponds to an acceptance rate of 41.19%, against 37.71% for the reference draft on the same stack.

vs reference (ATOM) is the like-for-like number -- same node, image, protocol and prompt counts. vs card crosses stacks, and that difference is not small: the reference draft reads 93.3% of its own card here, ranging per benchmark from 76.2% (HumanEval) to 116.3% (SPEED-Bench low-entropy), with no constant between them.

The card reports a 14-benchmark mean of 3.85. AA-LCR is the one we cannot reproduce: it is measured on 71k-115k-token multi-document prompts and the published dataset exposes only the question text (mean 206 characters), so it is excluded rather than substituted with something shorter.

Raw per-benchmark results, including the full accepted-length histograms for both drafts, are in benchmarks/ in this repo.

Protocol

Identical for both drafts, every run:

engine ATOM, TP=8, 8 x MI355X
image rocm/atom-dev@sha256:2f8bd4206ad15d014ae48115eae1ee9f1db83781848a8542de7177cfbd4ac914
--num-speculative-tokens 7
KV cache fp8
sampling greedy (temperature=0, top_p=1.0)
concurrency 1
max_tokens 12288 (8192 for the low-entropy 16k split)
window (--max-model-len) 16384 (32768 for the low-entropy 16k split)
max_num_seqs 8
prefix caching off
prompts official count for every benchmark, no sampling
counters read from /debug/mtp_stats, differenced before/after each set

Acceptance is read from the engine's own counters, not inferred from wall-clock time. The counters are cumulative since server start, so each benchmark is a before/after snapshot difference.

/!\ Why the benchmark window is 16384 and not 32768

The model serves a 32768 window -- that is what the Quick Start above uses. The benchmark runs twelve of the thirteen sets at 16384 for one reason: that is the window the reference draft was measured at, and changing only our side would break the comparison.

It costs nothing, because on those twelve sets the window never binds. The longest prompt + completion across all of them is 14,994 tokens (SPEED-Bench writing: 2,699 + 12,295), under the 16384 ceiling. Re-running them at 32768 returns the same numbers.

max_tokens does bind: eight of the twelve sets have generations that stop at the 12288 ceiling. Acceptance length is close to insensitive to it -- raising max_tokens from 2048 to 15360 on an earlier draft lengthened generations by 3.9x and quadrupled forward steps while moving tok/fwd by +0.2% -- but it is a real ceiling and it is applied identically to both drafts.

The thirteenth set (SPEED-Bench low-entropy) is the only one with long prompts: inputs average 10.4k tokens and reach 22.5k, so it gets its own server at 32768 with max_num_batched_tokens raised to match.


Quick Start

ATOM pinned by digest: atom 0.1.6rc1.dev275, torch 2.13.0+rocm7.14.0, HIP 7.14.60850.

docker run -d --name atom-dspark \
  --device=/dev/kfd --device=/dev/dri --group-add video \
  --security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
  --ipc=host --shm-size 128g --network host \
  -v /path/to/Kimi-K3:/target:ro -v /path/to/this/repo:/draft:ro \
  rocm/atom-dev@sha256:2f8bd4206ad15d014ae48115eae1ee9f1db83781848a8542de7177cfbd4ac914 \
  python -m atom.entrypoints.openai_server \
    --model /target --served-model-name Kimi-K3 \
    --method dspark --draft-model /draft --num-speculative-tokens 7 \
    --kv_cache_dtype fp8 -tp 8 --trust-remote-code \
    --max-model-len 32768 --max-num-seqs 8 --max-num-batched-tokens 32768 \
    --gpu-memory-utilization 0.93 --block-size 128 \
    --no-enable_prefix_caching --server-port 8000

The server log should show Detected MLA DSpark drafter and DSparkProposer aux capture on target layers: (2, 23, 47, 71, 89). Acceptance counters are at /debug/mtp_stats.

/!\ max_num_batched_tokens must be at least max_model_len: chunked prefill is off, and a prompt larger than the batched-token budget cannot be scheduled at all.

/!\ max_num_seqs stays small. K3's KDA recurrent state is allocated per slot; at 64 the state pool alone asks for 28.9 GB against a ~19 GB KV budget and the server refuses to start.


Architecture

{
  "architectures": ["K3DSparkModel"],
  "model_type": "k3_dspark",
  "hidden_size": 7168,
  "intermediate_size": 14336,
  "num_hidden_layers": 5,
  "num_attention_heads": 64,
  "num_key_value_heads": 64,
  "q_lora_rank": 1536,
  "kv_lora_rank": 512,
  "qk_nope_head_dim": 128,
  "qk_rope_head_dim": 64,
  "v_head_dim": 128,
  "vocab_size": 163840,
  "target_num_hidden_layers": 93,
  "target_layer_ids": [2, 23, 47, 71, 89],
  "markov_rank": 256,
  "enable_confidence_head": true,
  "max_position_embeddings": 1048576,
  "rope_parameters": {
    "rope_type": "yarn", "factor": 32.0,
    "original_max_position_embeddings": 32768,
    "rope_theta": 50000.0, "beta_fast": 32, "beta_slow": 1,
    "mscale": 1.0, "mscale_all_dim": 1.0
  }
}

/!\ rope_interleave is deliberately absent. ATOM maps that key onto aiter's is_neox_style, so setting it would select half-split rotation; this draft uses interleaved pairs. Emitting the key silently changes the attention definition and collapses acceptance.


Limitations

  • AA-LCR is not measured (see above).
  • Measured at a 32768 window. The config declares YaRN out to 1048576, but nothing above 32768 has been measured.
  • GSM8K is the one benchmark below the reference draft on this stack.
  • No training checkpoint is published.
Downloads last month
3
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for slippedJim/Kimi-K3-Draft

Finetuned
(52)
this model