k3-a40-bootstrap / FINDINGS.md
patdev's picture
Upload FINDINGS.md with huggingface_hub
4448aee verified
|
Raw History Blame Contribute Delete
31.7 kB

Kimi-K3 on 2Γ— A40 β€” measured findings

Everything here was measured on RunPod pod cumv30dk07dfdy: 2Γ— NVIDIA A40 (2 Γ— 46068 MiB), 93 GB container RAM, 16 container vCPU, 240 GB ephemeral disk, $0.92/h. Model: mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf UD-IQ1_S, 181 GiB. Runtime: mmnga/llama.cpp branch kimi-k3-width-support, CUDA 12.8, sm_86.

Numbers without a measurement behind them are marked as estimates.

Throughput

Configuration decode prefill static VRAM
autofit, 16 experts, greedy 2.94 tok/s avg 1.2–4.1 tok/s 87.7 GB
--cpu-moe --moe-cache on (unconfigured) 2.81 tok/s 0.9–2.4 tok/s 60.2 GB
autofit, 8 experts 3.15 tok/s 1.88 tok/s 87.4 GB

Prefill never exceeding decode by much is the tell: this deployment is not compute-bound on the GPU, it is dominated by whatever runs on the host.

Result: 7.3 tok/s

Step decode prefill cost
starting point (previous effort, L40S) 2.32 β€” β€”
2Γ— A40 after the fixes below 2.94 1.2–4.1 $0.92/h
5Γ— A40, identical placement 5.44 6.8–8.0 $2.20/h
5Γ— A40, everything VRAM-resident 7.35 27.3 $2.20/h

3.2Γ— on decode, 27Γ— on prefill. Two independent factors, both of which this investigation initially got wrong: host vCPU (which RunPod scales with GPU count) and full VRAM residency. Neither is a tuning parameter.

Output quality at 7.3 tok/s: English prose and reasoning are correct ("The trip duration is 3.5 hours"), code is correct (fibonacci(n-1) + fibonacci(n-2)), and the model continues fluently in Chinese β€” Moonshot's other strong language. French generation stays broken, which is the Width50 build, not the hardware.

DSpark speculative decoding gives nothing. Once the draft finally loaded β€” after clearing four separate blockers β€” throughput was 7.32/7.33 against 7.33/7.35 without it, and llama-server exposed no draft_n / draft_n_accepted telemetry to explain why. The predicted floor of 1.2Γ— was still too optimistic; the measured figure is 1.0Γ—.

The answer: host vCPU, not VRAM

The same configuration, on the same model, using the same two GPUs β€” three further GPUs sitting idle at 267 MiB β€” runs at 1.9Γ— the speed on a bigger host:

Pod GPUs vCPU RAM decode
cumv30dk07dfdy 2Γ— A40 16 93 GB 2.31 / 2.98 / 2.88
okinf3ln9k8tsp 5Γ— A40 41 234 GB 4.86 / 5.28 / 5.44 / 5.64

Prefill moves even harder: 0.99 β†’ 8.01 tok/s.

RunPod scales vCPU with GPU count, so renting more GPUs buys host parallelism as much as it buys VRAM. Every placement experiment below kept 70–90 GB of weights being processed by 16 saturated threads, which is why moving that work between CPU and GPU changed nothing: the CPU side was the wall in all of them.

This is the mechanism the section below says was missing. It was found by accident β€” a 5-GPU pod picked up the 2-GPU placement script before the guard was published, and ran the identical configuration on a larger host.

Why the placement experiments all looked the same

~3 tok/s is the ceiling for this model on 2Γ— A40, and the reason is not the one that seems obvious.

Throughput is invariant across configurations that differ enormously:

Configuration VRAM used attention CPU-side bytes/token decode
autofit, 16 experts 87.7 GB 48 layers on host ~33 GB 2.46 / 3.00
--cpu-moe + moe-cache 60.2 GB all in VRAM ~4 GB 2.21 / 2.68
--n-cpu-moe 86 69.9 GB all in VRAM ~4 GB 2.51 / 3.06
manual balanced -ot 87.8 GB all in VRAM, 44.3/43.5 split ~4 GB 2.31 / 2.98
autofit, 8 experts 87.4 GB 48 layers on host ~17 GB 3.15
-ngl 0 (no GPU) 1.5 GB all on host all 0.65

The manual placement is the cleanest of these: explicit -ot ...=CUDA0/CUDA1 layer ranges fill both cards to 96%/94% with every layer's attention resident and only 72 layers' routed experts on the host. It took three calibration rounds, each guided by the shortfall rather than a guess β€” 45256 MiB over, then 1038 MiB over, then fitting. It performs exactly like everything else.

Cutting host traffic by 8Γ— moved nothing. Halving expert compute moved +28%.

No validated mechanism. Three hypotheses were proposed and each was refuted by measurement:

  1. Routed-expert CPU compute dominates β€” refuted: halving experts per token gave +28%, not the ~100% that would follow.
  2. Attention re-read from host RAM dominates β€” refuted: --n-cpu-moe 86 puts every layer's attention in VRAM and cuts host bytes ~8Γ—, and changed nothing.
  3. CPU↔GPU boundary crossings dominate β€” refuted by comparing the two: autofit splits into contiguous GPU-then-CPU blocks (1–2 crossings per token) while --n-cpu-moe 86 crosses at 86 layers (~172 per token). If crossings dominated, autofit would be far faster. The two are within 2% of each other.

What stands is empirical, not theoretical: no rebalancing of CPU against GPU within 92 GB of VRAM improves anything, across placements spanning 60–88 GB of VRAM and 4–33 GB of host traffic per token.

The GPUs are not decorative, though β€” verified. Running with -ngl 0, so that essentially nothing sits in VRAM (1265 MiB and 269 MiB), gives 0.65 tok/s against 2.51 for the same prompt with autofit. The two A40s are worth 3.9Γ—. That check was run specifically because, if the GPUs had contributed nothing, "buy more VRAM" would have been the wrong advice.

So the direction β€” fit more of the model in VRAM β€” is sound, and a ggml microbenchmark says the magnitude is large. test-backend-ops timing MUL_MAT_ID on the K3 expert shapes (MXFP4, 3584β†’3072) on an L40S:

 8 experts, C=1    29.25 Β΅s
16 experts, C=1    59.83 Β΅s
 8 experts, C=7   159.53 Β΅s
16 experts, C=7   318.97 Β΅s      ~7.73 TFLOPS

At 59.83 Β΅s per MoE tensor and 92 MoE layers, the routed-expert work costs 5.5 ms/token when the experts are VRAM-resident β€” a ~180 tok/s ceiling from that term alone. We measure 333 ms/token, so GPU MoE compute is about 1.6% of the time; the other 98% is the host path.

That makes the naive linear extrapolation (5.4 tok/s at full residency) almost certainly far too pessimistic: full residency does not scale the existing regime, it deletes the term that dominates it.

What did NOT work, and why

Halving experts per token (16 β†’ 8). Bought only +28% (2.46 β†’ 3.15 on the same prompt) and destroyed the output β€” # # # # after six tokens. REAP had already cut 896 experts to 448 with the router calibrated for 16 active. The small gain also disproved the hypothesis that routed-expert CPU compute was the bottleneck: halving it would otherwise have nearly doubled throughput.

The CUDA MoE expert cache (llama-k3-moe-cache.patch, from the Code145 build). Applies cleanly to mmnga's branch once tests/ is excluded, compiles for sm_86, and --moe-cache appears in --help. Measured 14% slower on decode and roughly half the prefill. Caveat: that run set none of the seven GGML_CUDA_MOE_CACHE_* variables its own production launcher uses, so the verdict is against an unconfigured cache, not the design.

DSpark speculative decoding. Three separate blockers, each found only after clearing the previous one:

  1. --draft-max was removed; the flag is now --spec-draft-n-max.
  2. The published GGUF declares general.architecture = "dflash-draft"; llama.cpp registers LLM_ARCH_DFLASH as plain "dflash". Fixed by fix_dspark_arch.py (renames the architecture and its 30 namespaced KV keys, verified by reopening and comparing every tensor's shape, dtype and byte size).
  3. The rewritten file still fails: "DFlash model requires 'target_layers' in GGUF metadata". The third-party converter also names tensors dflash.dspark.markov.w1 where llama.cpp expects markov_w1. The clean fix is to convert from RadixArk/Kimi-K3-DSpark with llama.cpp's own converter.

Expected ceiling if it ever loads: 1.2–2Γ—, not the ~3Γ— vLLM reports. On a MoE each drafted token routes to its own experts, so batch verification amortises only attention and shared experts, not the routed reads.

GGML_OP_OFFLOAD_MIN_BATCH=1. ggml can copy only the selected experts of an offloaded MUL_MAT_ID to the GPU instead of computing them on the host, but the path is gated on batch size and the default of 32 means it never fires during decode. Forcing it on halved throughput: 2.46 β†’ 1.13 and 3.00 β†’ 1.38 tok/s. With 16 experts across 68 offloaded layers that is 1088 separate small transfers per token, so it is latency-bound, not bandwidth-bound, and loses to host compute. The upstream default is right.

--n-cpu-moe at any rung (62/68/74/80). Never loaded. Tensor overrides set model_params::tensor_buft_overrides, which aborts autofit exactly like -ngl and --tensor-split do β€” and once autofit is off, neither --split-mode layer nor --tensor-split distributes anything: every allocation lands on device 1 while device 0 stays empty. Rung 80 missed by 135 MiB (46203 requested against 46068 available) with all attention plus twelve expert layers on one card. Autofit is the only mechanism that actually uses both GPUs; balancing manually requires explicit -ot ...=CUDA0 / =CUDA1 layer ranges.

TurboQuant / KV-cache quantisation. Not applicable: 69 of 93 layers are KDA recurrent and carry no growing KV, so the whole cache is about 0.8 GB.

ONNX custom ops from earlier work target Qwen3.5 on SM75, not Kimi-K3 on sm_86, and ONNX Runtime has no kimi_k3 architecture.

Output quality

English and code are genuinely usable; French is not.

"The capital of Japan is"  -> Tokyo, Washington, Stockholm, Berlin   correct
"train 14:00 -> 17:30"     -> "The trip takes 3 hours 30 minutes."   correct
"def invert(d):"           -> return {v: k for k, v in d.items()}    correct
French free-form           -> "cunce", "s'fotrenci", "zocneine"      non-words

French knowledge survives β€” "La Revolution francaise a commence en" continues "1789" β€” but French generation is corrupted at the sub-word level. The prime suspect is Width50: this build halves moe_intermediate_size from 3072 to 1536.

Sampler, six configurations on one prompt

Setting Result
greedy, rp 1.0 facts correct, then hard repetition loop
greedy, rp 1.05 / last_n 64 no loop, but "Germany is PLN" β€” the penalty pushes it off correct answers while enumerating
greedy, rp 1.15 / last_n 256 drifts into SQL-flavoured nonsense
temp 0.3, min_p 0.05 fluent, hallucinates ("Tokyo Island", "Sumaru")
temp 0.6, top_k 20 degrades
temp 0.3, min_p 0.1, rp 1.05 fluent, no loop, nothing glaringly wrong β€” chosen default

Repetition penalty measurably reduces factual accuracy on this model. Clients wanting maximum accuracy should still ask for temperature: 0.

Deployment mechanics

common_fit_params sizes GPU placement against actually-free VRAM across all devices, but aborts the moment any placement parameter is pinned:

-ngl 99        -> "n_gpu_layers already set by user, abort"
--tensor-split -> "model_params::tensor_split already set by user, abort"

Both aborts produced the same failure: 84790 MiB shoved onto device 1 alone, cudaMalloc OOM, device 0 untouched. With no placement flags, autofit fills both cards to 95%. --cpu-moe and --n-cpu-moe do not abort it.

Boot time went from 14m28s to ~4 minutes by publishing the compiled binaries to the Hub (46 MB tarball, 2 s to extract vs 8m34s to compile). The artifact name pins base image, CUDA version and GPU arch so a mismatched tarball can never be picked up silently, and it is smoke-tested by running llama-server --version before it is trusted.

RunPod costs. A stopped pod's Volume disk bills at $0.20/GB/month, double the running rate β€” a 1600 GB volume cost $0.44/h doing nothing, which was 99% of one day's spend. Container disk is erased on stop but still billed, so for re-downloadable data the rule is: always Terminate, never Stop. A stopped pod is also pinned to its host and may refuse to restart while the same GPU type shows High stock elsewhere.

Kimi-K3-256K-REAP-Code145

Finalised during this session. All 96 shards were already on the Hub (347.5 GB); only config.json, the index and the tokenizer were missing, and finalize_code145.py β€” which downloads only metadata, never the weights β€” had never been run. The repo is now loadable.

num_experts             145
num_experts_per_token    16
moe_intermediate_size  3072      full width, unlike the 1536 we serve
source_index_keys    497 220
dropped_expert_keys  414 552
output_index_keys     82 668     = 497220 - 414552
renamed_expert_keys   80 040     = 92 layers x 145 experts x 6 tensors

Full width makes it the serious candidate for working French. It does not fit 2Γ— A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4Γ— A40 at $1.80/h, or an imatrix-guided quantisation down to ~1.7 bit.


Session 2 β€” what actually bounds decode, and what long context costs

Decode is serialisation-bound, not bandwidth-bound

Six placements had moved 60–88 GB of weights between CPU and GPU without changing throughput, and no mechanism had survived measurement. The question none of them asked was whether the hardware is idle between tokens. It is.

Aggregate throughput against concurrent streams, 5Γ— A40, 64 tokens each, distinct prefixes so no stream rides another's prompt cache:

1 stream     5.56 tok/s aggregate    5.56 per stream
2 streams    9.85 tok/s aggregate    4.93 per stream   1.77x
4 streams   14.53 tok/s aggregate    3.63 per stream   2.61x

Concurrency scales, so single-stream decode leaves silicon idle. The bandwidth arithmetic agrees independently: 7 GB of weights are touched per token, and at 138 ms/token that is **51 GB/s against 696 GB/s available** on the active card β€” 7% of the roof. llama.cpp splits layers across GPUs, so at batch 1 exactly one card computes while the other four wait.

This retires the last open question from session 1 and explains why every kernel-level optimisation attempt returned nothing: the expert GEMM was never the constraint.

Sustained decode on real output: 7.22 tok/s

The 7.3 tok/s headline was measured on selftests, three of whose five prompts produce degenerate repetition (# [CS_ah], et d'et d') β€” and a repetition loop decodes cheap. Re-measured on 600 tokens of coherent prose: 7.22 tok/s. The number holds; it is now measured on output worth having.

Prefill: 167 tok/s cold, ~free warm

Every selftest prompt is 6–22 tokens, so its "prefill tok/s" is per-request overhead and says nothing. Measured properly, 3408-token prompt:

cold   3408 tok in 20.4 s = 167 tok/s     (23x the decode rate β€” prefill batches)
warm      4 tok in  0.53 s                (slot prompt cache reprocessed only the delta)

For a Claude Code client sending a 32k system prompt: ~3 min once, then only the changed tokens on every later turn.

Long context is nearly free on this architecture

From the config, not from assumption: kv_lora_rank 512, qk_rope_head_dim 64, full_attn_layers every 4th layer, max_position_embeddings 1048576.

24 of 93 layers cache at all (MLA); the other 69 are KDA β€” recurrent state,
constant regardless of context length.

KV per token = 24 x (512 + 64) x 2 = 27.6 KB

  65536 ctx ->  1.8 GB      262144 ->  7.2 GB      1048576 -> 29 GB

Against ~38 GB left free by the weights on a 5-GPU host. 16384 was never a hardware limit, just an untested default β€” and it could not accept a single Claude Code request. What constrains context here is the weights sharing the cards, never the cache.

Cost reality check

Runpod hosts an official Kimi-K3 endpoint (moonshot-kimi). Measured: 300 tokens in 10.281 s, billed $0.004797.

                    speed         $/1M output    quality
official          29.2 tok/s          ~16        full precision, French works
ours (1 stream)    7.2 tok/s          ~85        1.14 bpw, French broken
ours (4 streams)  14.5 tok/s          ~42

The official endpoint is 4x faster and 5.7x cheaper. Self-hosting buys privacy, no per-token billing, and context control β€” not price or quality. Worth restating whenever the effort seems to justify itself on cost.

No smaller GGUF exists

Ours is the most compact Kimi-K3 published, because it is the Width50 variant (moe_intermediate_size 1536 rather than 3072) β€” which is also why French is broken.

ours   REAP448 Width50 IQ1_S       181 GiB
       0xTank REAP568 UD-IQ1_S     246 GiB
       mmnga-o REAP50 UD-IQ1_S     305 GiB
       prometheusAIR REAP55 IQ1_M  319 GiB
       hellohazime REAP640ja       411 GiB

A "light" build for a single A40 (46 GB, ~37 GB of weights after a 256k KV) does not exist and would have to be produced β€” roughly a 64-expert REAP.

Under $1 per million tokens is not reachable with K3

At $0.44/h per A40, $1/1M output requires 122 tok/s aggregate per card. The 5Γ— pod at $2.20/h would need 611 tok/s; it delivers 14.5. That is a 42x gap, not a tuning problem.

What does reach it, from HF configs (throughput figures are bandwidth-derived estimates, not measured):

Qwen3-Coder-30B-A3B  30B/3B active  ~30 GB FP8  KV 12 GiB@256k  1x A40  ~$0.10-0.25/1M
Qwen3-Coder-Next     80B/3B active  ~74 GB FP8  KV 12 GiB@256k  2x A40  ~$0.20-0.80/1M
Qwen3.8-27B          dense          ~26 GB FP8  KV 32 GiB@256k  256k does not fit on 1 card

Served by vLLM, not llama.cpp β€” tensor parallelism and continuous batching are exactly what the concurrency measurement above shows llama.cpp lacking.

Operational trap: bare wait kills the hot-reload watcher

llama-server is started with &, so a bare wait anywhere later in the script waits for it too, forever. The server keeps serving, the selftest never finishes, the watcher never starts, and the pod silently ignores every published update. Symptom: /props reports the old config long after a push. Always wait "$pid" with an explicit pid.


Session 3 β€” 256k validated, and a measured alternative that dominates it

Kimi-K3 at 262144 context: it works, and it costs nothing in decode

--ctx-size 262144, all-VRAM on 5x A40, 39.5-42.5 GB used per card:

n_ctx        262144        (was 16384 -- never a hardware limit, just a default)
decode          7.25 tok/s (unchanged from 16k: 7.2-7.3)

16x the context for no throughput cost, exactly as the KDA/MLA geometry predicted. The DSpark rungs OOM'd at this context and the ladder correctly fell through to all-VRAM without a draft β€” no loss, DSpark was measured at zero gain.

Prefill degrades gracefully, and the proxy times out before it finishes

A cold 38k prefill takes ~230 s. Runpod's HTTP proxy sits behind Cloudflare, which cuts the connection at ~125 s (error code: 524). The workaround is to build the prefix in chunks β€” each request extends the cached prefix β€” which also measures the degradation curve a single call would have hidden:

chunk 1  +8400 tok @  8k ctx   228.6 tok/s
chunk 2  +8404 tok @ 17k ctx   203.6 tok/s
chunk 3  +8404 tok @ 25k ctx   182.4 tok/s
chunk 4  +8605 tok @ 34k ctx   164.1 tok/s
chunk 5  +8704 tok @ 43k ctx   149.8 tok/s

-34% across 5x the context, far better than quadratic β€” 69 of 93 layers are KDA (linear), only 24 are MLA. 42,517 tokens prefilled in 3 min 55 total. Re-submitting a cached 32k prompt returns prompt_n=4 in 0.6 s.

vLLM's Kimi-K3 support is real and unreachable

vLLM and SGLang both ship day-0 K3 support, including the two things that failed here: a hybrid prefix cache over recurrent KDA state, and working DSpark block-diffusion speculative decoding. None of it is usable on A40s, because every vLLM-loadable K3 checkpoint is the full unpruned model:

nvidia/Kimi-K3-NVFP4        1499 GiB   + requires Blackwell (sm_100)
RedHatAI/Kimi-K3-NVFP4      1533 GiB   + requires Blackwell
RedHatAI/Kimi-K3-FP8-BLOCK  2626 GiB   + requires sm_89+

No AWQ, no GPTQ, no W4A16. The 181 GiB IQ1_S we run is llama.cpp-only, and llama.cpp is exactly what caps us. The unlock would be quantising a REAP'd K3 to W4A16 for vLLM β€” Code145 at 4-bit lands near 260 GB (6x A40 / 4x A100-80). That build, not any config change, is the only path to the good engine.

The measured alternative

Qwen3-Coder-30B-A3B-Instruct, AWQ 4-bit (16.9 GiB), vLLM 0.27.1, one A40 at $0.44/h, 262144 context with fp8 KV. Everything below is measured, not derived:

concurrency   tokens   duration   aggregate      per stream   $/1M output
     1          128      1.6 s      77.6 tok/s      77.6        $1.575
     4          512      1.9 s     265.3 tok/s      66.3        $0.461
     8         1024      2.1 s     496.8 tok/s      62.1        $0.246
    16         2048      2.6 s     782.4 tok/s      48.9        $0.156
    32         4096      3.5 s    1165.6 tok/s      36.4        $0.105

prefill  38,494 tok cold in 10.42 s = 3693 tok/s;  warm 1.07 s

Head to head:

                        Kimi-K3 5x A40   Qwen3-Coder 1x A40   K3 official API
hourly                      $2.20              $0.44           per-token
decode, 1 stream          7.2 tok/s         77.6 tok/s        29.2 tok/s
prefill @38k              ~165 tok/s        3693 tok/s             --
32k prompt, cold           3 min 55            10 s                --
context 262144            5 cards            1 card                --
$/1M output                  ~85         $1.58 .. $0.105          ~16
French                      broken            correct            correct

Qwen is 10.8x on decode, 22x on prefill, on one fifth of the hardware, and it answers in French. The under-$1/1M target is met from two concurrent streams and reaches $0.105 at 32. Kimi-K3 self-hosted on A40s is not competitive on any axis measured here; it remains interesting only where the specific K3 model is the requirement.

Two mistakes worth not repeating

  • Creating a pod from create-pod without a start command: the schema has no args field, so the container starts with the image default. Use create-template with dockerStartCmd, then deploy with templateId.
  • --disable-log-requests was removed in vLLM 0.27; it crash-loops the server. A pod's args cannot be edited after creation β€” only the template can β€” so a bad flag costs a full recreate.

The context ceiling on 5x A40, and why q8 KV is not the way out

ctx      16384    7.2-7.3 tok/s
ctx     262144    7.25 tok/s
ctx     524288    7.29 tok/s     <- ceiling, all-VRAM, no draft
ctx    1048576    OOM: 29 GB of f16 MLA cache does not fit alongside the weights
KV q8_0            rejected outright, at every context

Decode is flat across a 32x range of context. That is the KDA/MLA hybrid doing exactly what it is for: only 24 of 93 layers cache at all, and the other 69 hold a recurrent state whose size is independent of context.

q8_0 KV cannot be used here, and the failure signature says why. Every placement rung failed in 1.6-2.5 s with failed to create llama_context from model and no cudaMalloc line β€” a parameter rejection. Contrast the genuine OOM at 512k on the draft rung: 29.9 s, with cudaMalloc failed: out of memory printed. Fast failure with no allocator message means llama.cpp refused the configuration; slow failure with one means it tried and ran out. The 69 KDA layers carry a recurrent state llama.cpp will not quantise, so f16 is forced and the ceiling is set by memory: 512k fits, 1M does not.

TurboQuant (Google, ICLR 2026) would take the cache to ~3 bits and put 1M within reach, but as of this run it is a paper with an open SGLang feature request (#21618) and no merged implementation in any engine.

Qwen decode against context length

short context   77.6 tok/s
119k context    34.3 tok/s      (prompt served warm from prefix cache)

Halves across 119k, and still 4.7x Kimi's short-context rate. Tool calling was verified on the OpenAI endpoint: two parallel calls, correct arguments, finish_reason: tool_calls.

Runpod's HTTP proxy caps a single request at ~125 s

Cloudflare sits in front of the pod proxy and returns error code: 524. A cold prefill longer than that cannot complete in one call β€” measured at ~40k on Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached prefix) or expose a TCP port instead.


Session 4 β€” speculative decoding, and why the regime decides everything

n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40, --speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'. Same model, same hardware, same quant, two pods running side by side so the comparison is simultaneous rather than sequential:

A. pure generation (no prompt/output overlap)
     baseline    102.21 tok/s
     ngram        58.49 tok/s      -43%

B. code editing (output re-emits the input)
     baseline    130.16 tok/s
     ngram       242.39 tok/s      +86%

Speculation is not a free win β€” it is a bet on repetition. With nothing to guess, every draft is rejected and the verification cost is pure loss; that is the -43%. Claude Code lives almost entirely in regime B (read a file, emit a modified version, repeat identifiers), so it is the right default there and the wrong default for prose.

vLLM's telemetry, aggregated over both regimes:

drafts 207   draft tokens 1035   accepted 706  = 68.2%
accepted per position: 179 / 156 / 131 / 120 / 120
mean accepted length 3.41  ->  ~4.4 tokens emitted per forward pass

This is the measurement llama.cpp never produced: DSpark there exposed no draft_n at all and moved throughput by 0.0%. The mechanism was never broken β€” the engine was.

The number that matters for the original goal

Single stream, code editing, one A40 at $0.44/h:

242.39 tok/s  ->  $0.50 per 1M output tokens

The under-$1/1M target is met on a single stream, without needing concurrency to amortise anything. For reference the same target on Kimi-K3 required 611 tok/s aggregate against 14.5 measured.

Note also that the baseline itself reads higher here (102-130 tok/s) than the 77.6 measured earlier: that earlier figure was a 128-token request whose wall clock was dominated by per-request overhead. Longer outputs amortise it. Quote 77.6 for short replies and ~130 for sustained generation.


Session 5 β€” Kimi-Linear-48B-A3B: genuine Kimi, 1M context, one A40

cyankiwi/Kimi-Linear-48B-A3B-Instruct-AWQ-4bit, 28.4 GiB, vLLM 0.27.1, one A40 at $0.44/h. KimiLinearForCausalLM, model_type: kimi_linear β€” the same architecture family as K3, with the same MLA geometry (kv_lora_rank 512, qk_rope_head_dim 64).

Why it fits where K3 could not

27 layers = 7 MLA (full_attn_layers [4,8,12,16,20,24,27]) + 20 KDA
KV/token  = 7 x (512 + 64) x 2 = 7.88 KB      (K3: 24 layers -> 27.6 KB)

weights AWQ4   28.4 GiB
1,048,576 ctx   7.88 GiB KV   =>  36.3 GB on a 46 GB card

vLLM confirmed it at boot: GPU KV cache size: 1,416,566 tokens on a single card. K3 needed five A40s at $2.20/h to reach 524,288.

Measured

flux    aggregate   per stream   $/1M
  1        67.4        67.4      1.814
  4       206.6        51.6      0.592
  8       352.4        44.0      0.347
 16       556.1        34.8      0.220
 32       779.9        24.4      0.157

prefill cold  6162 tok/s     (better than Qwen3-Coder-30B's 3957)
prefill warm  6430 tok/s     ratio 1.04 -> NO prefix caching
code editing  89.4 tok/s
tool calling  2 parallel calls, correct arguments
code          correct;  French fluent

n-gram speculation is BROKEN on this backend

Rejection sampling is supposed to be lossless, so any output difference between speculative and non-speculative decoding is a bug. Here it corrupts code, reproducibly and identically across runs:

                with speculation                    without
add()     return a -b):\n    return a -      return a - b\n\ndef multiply(a,
fibonacci fibonacci(n-1(n-1) + fibonacci(n-2) fibonacci(n-1) + fibonacci(n-2)

The signature is a just-emitted fragment re-spliced at the wrong place β€” a bad draft accept. Plain prose and verbatim copying stayed perfect, which is why it took targeted probes to find: code is dense in the short repeats (parentheses, operators) that n-gram matching latches onto.

It is also 3-4x slower, because vLLM disables CUDA graphs under speculation on TritonMLABackend:

flux      with spec   without    ratio
  1          23.1       71.2      3.1x
  4          58.7      249.6      4.3x
 32         203.9      781.0      3.8x
$/1M @32     0.600      0.157

Turn speculation off for Kimi-Linear. Worth reporting upstream.

Two fixes that made it work at all

  • Tokenizer. tokenization_kimi.py imports bytes_to_unicode from transformers.convert_slow_tokenizer, removed in transformers >= 5.5.3 which every vLLM >= 0.24 requires. No Kimi-Linear repo ships a fast tokenizer.json, so the Python tokenizer is mandatory. Fixed by publishing patdev/kimi-linear-tokenizer-fix: same files with a local fallback definition, verified byte-identical to the transformers reference before deployment. Point vLLM at it with --tokenizer.
  • Tool parser. hermes yields zero calls. The chat template uses <|tool_calls_section_begin|> / <|tool_call_begin|>, the K2 token scheme β€” so the parser is kimi_k2. (kimi_k3 is for <|open|>/<|close|>/<|sep|> and does not apply.) It also sets skip_special_tokens=False, which those markers need.
  • Headroom. --gpu-memory-utilization 0.93 OOMs at the very last step, in the speculative rejection sampler asking for 160 MiB with 143 MiB left. 0.90 leaves room and still yields >1M tokens of KV.

Head to head

                    Kimi-Linear 48B    Qwen3-Coder-30B
hourly                   $0.44              $0.44
decode 1 stream        67-71 tok/s        77.6-79.2 tok/s
aggregate @32            780               1166
prefill cold            6162               3957
prefill warm            6430              63315
context              1,048,576            262,144
$/1M @32                0.157              0.105
tool calling            yes                 yes
French                  fluent             correct
brand                GENUINE KIMI          Qwen

Kimi-Linear wins on context (4x) and cold prefill (1.6x). It loses on decode and, decisively for an agent client, on prefix caching: 6430 vs 63315, ten times worse, because vLLM does not yet support prefix caching over recurrent hybrid state. A 32k system prompt costs ~5 s every turn instead of ~0.5 s.

Qwen3-Coder-Next 80B: no data

Two attempts, both infrastructure failures, zero requests served.

attempt 1   hung 22 min in NCCL peer-to-peer init on 2x A40
attempt 2   NCCL_P2P_DISABLE=1 cleared that, then looped ~10 min on
            "No available shared memory broadcast block found in 60 seconds"

NCCL_P2P_DISABLE=1 is required for tensor parallelism on Runpod A40 pairs. The second hang is unproven β€” the hypothesis is torch.compile on a 512-expert MoE, testable with --enforce-eager. Its boot log did confirm one thing: enable_prefix_caching=False, the same hybrid limitation as Kimi-Linear.