# Kimi-K3 on 2× A40 — measured findings Everything here was measured on RunPod pod `cumv30dk07dfdy`: 2× NVIDIA A40 (2 × 46068 MiB), 93 GB container RAM, 16 container vCPU, 240 GB ephemeral disk, $0.92/h. Model: `mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf` UD-IQ1_S, 181 GiB. Runtime: `mmnga/llama.cpp` branch `kimi-k3-width-support`, CUDA 12.8, sm_86. Numbers without a measurement behind them are marked as estimates. ## Throughput | Configuration | decode | prefill | static VRAM | |---|---|---|---| | autofit, 16 experts, greedy | **2.94 tok/s** avg | 1.2–4.1 tok/s | 87.7 GB | | `--cpu-moe --moe-cache on` (unconfigured) | 2.81 tok/s | 0.9–2.4 tok/s | 60.2 GB | | autofit, 8 experts | 3.15 tok/s | 1.88 tok/s | 87.4 GB | Prefill never exceeding decode by much is the tell: this deployment is not compute-bound on the GPU, it is dominated by whatever runs on the host. ## Result: 7.3 tok/s | Step | decode | prefill | cost | |---|---|---|---| | starting point (previous effort, L40S) | 2.32 | — | — | | 2× A40 after the fixes below | 2.94 | 1.2–4.1 | $0.92/h | | 5× A40, *identical* placement | 5.44 | 6.8–8.0 | $2.20/h | | 5× A40, everything VRAM-resident | **7.35** | **27.3** | $2.20/h | 3.2× on decode, 27× on prefill. Two independent factors, both of which this investigation initially got wrong: **host vCPU** (which RunPod scales with GPU count) and **full VRAM residency**. Neither is a tuning parameter. Output quality at 7.3 tok/s: English prose and reasoning are correct ("The trip duration is 3.5 hours"), code is correct (`fibonacci(n-1) + fibonacci(n-2)`), and the model continues fluently in Chinese — Moonshot's other strong language. French generation stays broken, which is the Width50 build, not the hardware. **DSpark speculative decoding gives nothing.** Once the draft finally loaded — after clearing four separate blockers — throughput was 7.32/7.33 against 7.33/7.35 without it, and llama-server exposed no `draft_n` / `draft_n_accepted` telemetry to explain why. The predicted floor of 1.2× was still too optimistic; the measured figure is 1.0×. ## The answer: host vCPU, not VRAM The same configuration, on the same model, using the same two GPUs — three further GPUs sitting idle at 267 MiB — runs at **1.9× the speed** on a bigger host: | Pod | GPUs | vCPU | RAM | decode | |---|---|---|---|---| | `cumv30dk07dfdy` | 2× A40 | **16** | 93 GB | 2.31 / 2.98 / 2.88 | | `okinf3ln9k8tsp` | 5× A40 | **41** | 234 GB | 4.86 / 5.28 / 5.44 / 5.64 | Prefill moves even harder: 0.99 → 8.01 tok/s. RunPod scales vCPU with GPU count, so renting more GPUs buys host parallelism as much as it buys VRAM. Every placement experiment below kept 70–90 GB of weights being processed by 16 saturated threads, which is why moving that work between CPU and GPU changed nothing: the CPU side was the wall in all of them. This is the mechanism the section below says was missing. It was found by accident — a 5-GPU pod picked up the 2-GPU placement script before the guard was published, and ran the *identical* configuration on a larger host. ## Why the placement experiments all looked the same ~3 tok/s is the ceiling for this model on 2× A40, and the reason is not the one that seems obvious. Throughput is **invariant** across configurations that differ enormously: | Configuration | VRAM used | attention | CPU-side bytes/token | decode | |---|---|---|---|---| | autofit, 16 experts | 87.7 GB | 48 layers on host | ~33 GB | 2.46 / 3.00 | | `--cpu-moe` + moe-cache | 60.2 GB | all in VRAM | ~4 GB | 2.21 / 2.68 | | `--n-cpu-moe 86` | 69.9 GB | all in VRAM | ~4 GB | 2.51 / 3.06 | | manual balanced `-ot` | 87.8 GB | all in VRAM, 44.3/43.5 split | ~4 GB | 2.31 / 2.98 | | autofit, 8 experts | 87.4 GB | 48 layers on host | ~17 GB | 3.15 | | `-ngl 0` (no GPU) | 1.5 GB | all on host | all | **0.65** | The manual placement is the cleanest of these: explicit `-ot ...=CUDA0/CUDA1` layer ranges fill both cards to 96%/94% with every layer's attention resident and only 72 layers' routed experts on the host. It took three calibration rounds, each guided by the shortfall rather than a guess — 45256 MiB over, then 1038 MiB over, then fitting. It performs exactly like everything else. Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%. **No validated mechanism.** Three hypotheses were proposed and each was refuted by measurement: 1. *Routed-expert CPU compute dominates* — refuted: halving experts per token gave +28%, not the ~100% that would follow. 2. *Attention re-read from host RAM dominates* — refuted: `--n-cpu-moe 86` puts every layer's attention in VRAM and cuts host bytes ~8×, and changed nothing. 3. *CPU↔GPU boundary crossings dominate* — refuted by comparing the two: autofit splits into contiguous GPU-then-CPU blocks (1–2 crossings per token) while `--n-cpu-moe 86` crosses at 86 layers (~172 per token). If crossings dominated, autofit would be far faster. The two are within 2% of each other. What stands is empirical, not theoretical: **no rebalancing of CPU against GPU within 92 GB of VRAM improves anything**, across placements spanning 60–88 GB of VRAM and 4–33 GB of host traffic per token. **The GPUs are not decorative, though — verified.** Running with `-ngl 0`, so that essentially nothing sits in VRAM (1265 MiB and 269 MiB), gives **0.65 tok/s** against 2.51 for the same prompt with autofit. The two A40s are worth **3.9×**. That check was run specifically because, if the GPUs had contributed nothing, "buy more VRAM" would have been the wrong advice. So the direction — fit more of the model in VRAM — is sound, and a ggml microbenchmark says the magnitude is large. `test-backend-ops` timing `MUL_MAT_ID` on the K3 expert shapes (MXFP4, 3584→3072) on an L40S: ``` 8 experts, C=1 29.25 µs 16 experts, C=1 59.83 µs 8 experts, C=7 159.53 µs 16 experts, C=7 318.97 µs ~7.73 TFLOPS ``` At 59.83 µs per MoE tensor and 92 MoE layers, the routed-expert work costs **5.5 ms/token when the experts are VRAM-resident** — a ~180 tok/s ceiling from that term alone. We measure 333 ms/token, so GPU MoE compute is about **1.6% of the time**; the other 98% is the host path. That makes the naive linear extrapolation (5.4 tok/s at full residency) almost certainly far too pessimistic: full residency does not scale the existing regime, it deletes the term that dominates it. ## What did NOT work, and why **Halving experts per token (16 → 8).** Bought only +28% (2.46 → 3.15 on the same prompt) and destroyed the output — `# # # #` after six tokens. REAP had already cut 896 experts to 448 with the router calibrated for 16 active. The small gain also *disproved* the hypothesis that routed-expert CPU compute was the bottleneck: halving it would otherwise have nearly doubled throughput. **The CUDA MoE expert cache** (`llama-k3-moe-cache.patch`, from the Code145 build). Applies cleanly to mmnga's branch once `tests/` is excluded, compiles for sm_86, and `--moe-cache` appears in `--help`. Measured 14% *slower* on decode and roughly half the prefill. Caveat: that run set none of the seven `GGML_CUDA_MOE_CACHE_*` variables its own production launcher uses, so the verdict is against an unconfigured cache, not the design. **DSpark speculative decoding.** Three separate blockers, each found only after clearing the previous one: 1. `--draft-max` was removed; the flag is now `--spec-draft-n-max`. 2. The published GGUF declares `general.architecture = "dflash-draft"`; llama.cpp registers `LLM_ARCH_DFLASH` as plain `"dflash"`. Fixed by `fix_dspark_arch.py` (renames the architecture and its 30 namespaced KV keys, verified by reopening and comparing every tensor's shape, dtype and byte size). 3. The rewritten file still fails: *"DFlash model requires 'target_layers' in GGUF metadata"*. The third-party converter also names tensors `dflash.dspark.markov.w1` where llama.cpp expects `markov_w1`. The clean fix is to convert from `RadixArk/Kimi-K3-DSpark` with llama.cpp's own converter. Expected ceiling if it ever loads: **1.2–2×**, not the ~3× vLLM reports. On a MoE each drafted token routes to its own experts, so batch verification amortises only attention and shared experts, not the routed reads. **`GGML_OP_OFFLOAD_MIN_BATCH=1`.** ggml can copy only the *selected* experts of an offloaded `MUL_MAT_ID` to the GPU instead of computing them on the host, but the path is gated on batch size and the default of 32 means it never fires during decode. Forcing it on **halved throughput**: 2.46 → 1.13 and 3.00 → 1.38 tok/s. With 16 experts across 68 offloaded layers that is 1088 separate small transfers per token, so it is latency-bound, not bandwidth-bound, and loses to host compute. The upstream default is right. **`--n-cpu-moe` at any rung (62/68/74/80).** Never loaded. Tensor overrides set `model_params::tensor_buft_overrides`, which aborts autofit exactly like `-ngl` and `--tensor-split` do — and once autofit is off, **neither `--split-mode layer` nor `--tensor-split` distributes anything**: every allocation lands on device 1 while device 0 stays empty. Rung 80 missed by 135 MiB (46203 requested against 46068 available) with all attention plus twelve expert layers on one card. Autofit is the only mechanism that actually uses both GPUs; balancing manually requires explicit `-ot ...=CUDA0` / `=CUDA1` layer ranges. **TurboQuant / KV-cache quantisation.** Not applicable: 69 of 93 layers are KDA recurrent and carry no growing KV, so the whole cache is about 0.8 GB. **ONNX custom ops** from earlier work target Qwen3.5 on SM75, not Kimi-K3 on sm_86, and ONNX Runtime has no `kimi_k3` architecture. ## Output quality English and code are genuinely usable; French is not. ``` "The capital of Japan is" -> Tokyo, Washington, Stockholm, Berlin correct "train 14:00 -> 17:30" -> "The trip takes 3 hours 30 minutes." correct "def invert(d):" -> return {v: k for k, v in d.items()} correct French free-form -> "cunce", "s'fotrenci", "zocneine" non-words ``` French *knowledge* survives — "La Revolution francaise a commence en" continues "1789" — but French *generation* is corrupted at the sub-word level. The prime suspect is Width50: this build halves `moe_intermediate_size` from 3072 to 1536. ### Sampler, six configurations on one prompt | Setting | Result | |---|---| | greedy, rp 1.0 | facts correct, then hard repetition loop | | greedy, rp 1.05 / last_n 64 | no loop, but "Germany is PLN" — the penalty pushes it off correct answers while enumerating | | greedy, rp 1.15 / last_n 256 | drifts into SQL-flavoured nonsense | | temp 0.3, min_p 0.05 | fluent, hallucinates ("Tokyo Island", "Sumaru") | | temp 0.6, top_k 20 | degrades | | **temp 0.3, min_p 0.1, rp 1.05** | fluent, no loop, nothing glaringly wrong — chosen default | Repetition penalty measurably *reduces factual accuracy* on this model. Clients wanting maximum accuracy should still ask for `temperature: 0`. ## Deployment mechanics `common_fit_params` sizes GPU placement against actually-free VRAM across all devices, but **aborts the moment any placement parameter is pinned**: ``` -ngl 99 -> "n_gpu_layers already set by user, abort" --tensor-split -> "model_params::tensor_split already set by user, abort" ``` Both aborts produced the same failure: 84790 MiB shoved onto device 1 alone, cudaMalloc OOM, device 0 untouched. With no placement flags, autofit fills both cards to 95%. `--cpu-moe` and `--n-cpu-moe` do *not* abort it. **Boot time** went from 14m28s to ~4 minutes by publishing the compiled binaries to the Hub (46 MB tarball, 2 s to extract vs 8m34s to compile). The artifact name pins base image, CUDA version and GPU arch so a mismatched tarball can never be picked up silently, and it is smoke-tested by running `llama-server --version` before it is trusted. **RunPod costs.** A stopped pod's Volume disk bills at $0.20/GB/month, double the running rate — a 1600 GB volume cost $0.44/h doing nothing, which was 99% of one day's spend. Container disk is erased on stop but still billed, so for re-downloadable data the rule is: always Terminate, never Stop. A stopped pod is also pinned to its host and may refuse to restart while the same GPU type shows High stock elsewhere. ## Kimi-K3-256K-REAP-Code145 Finalised during this session. All 96 shards were already on the Hub (347.5 GB); only `config.json`, the index and the tokenizer were missing, and `finalize_code145.py` — which downloads only metadata, never the weights — had never been run. The repo is now loadable. ``` num_experts 145 num_experts_per_token 16 moe_intermediate_size 3072 full width, unlike the 1536 we serve source_index_keys 497 220 dropped_expert_keys 414 552 output_index_keys 82 668 = 497220 - 414552 renamed_expert_keys 80 040 = 92 layers x 145 experts x 6 tensors ``` Full width makes it the serious candidate for working French. It does not fit 2× A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4× A40 at $1.80/h, or an imatrix-guided quantisation down to ~1.7 bit. --- # Session 2 — what actually bounds decode, and what long context costs ## Decode is serialisation-bound, not bandwidth-bound Six placements had moved 60–88 GB of weights between CPU and GPU without changing throughput, and no mechanism had survived measurement. The question none of them asked was whether the hardware is *idle* between tokens. It is. Aggregate throughput against concurrent streams, 5× A40, 64 tokens each, distinct prefixes so no stream rides another's prompt cache: ``` 1 stream 5.56 tok/s aggregate 5.56 per stream 2 streams 9.85 tok/s aggregate 4.93 per stream 1.77x 4 streams 14.53 tok/s aggregate 3.63 per stream 2.61x ``` Concurrency scales, so single-stream decode leaves silicon idle. The bandwidth arithmetic agrees independently: ~7 GB of weights are touched per token, and at 138 ms/token that is **~51 GB/s against 696 GB/s available** on the active card — 7% of the roof. llama.cpp splits layers across GPUs, so at batch 1 exactly one card computes while the other four wait. This retires the last open question from session 1 and explains why every kernel-level optimisation attempt returned nothing: the expert GEMM was never the constraint. ## Sustained decode on real output: 7.22 tok/s The 7.3 tok/s headline was measured on selftests, three of whose five prompts produce degenerate repetition (`# [CS_ah]`, `et d'et d'`) — and a repetition loop decodes cheap. Re-measured on 600 tokens of coherent prose: **7.22 tok/s**. The number holds; it is now measured on output worth having. ## Prefill: 167 tok/s cold, ~free warm Every selftest prompt is 6–22 tokens, so its "prefill tok/s" is per-request overhead and says nothing. Measured properly, 3408-token prompt: ``` cold 3408 tok in 20.4 s = 167 tok/s (23x the decode rate — prefill batches) warm 4 tok in 0.53 s (slot prompt cache reprocessed only the delta) ``` For a Claude Code client sending a 32k system prompt: **~3 min once**, then only the changed tokens on every later turn. ## Long context is nearly free on this architecture From the config, not from assumption: `kv_lora_rank 512`, `qk_rope_head_dim 64`, `full_attn_layers` every 4th layer, `max_position_embeddings 1048576`. ``` 24 of 93 layers cache at all (MLA); the other 69 are KDA — recurrent state, constant regardless of context length. KV per token = 24 x (512 + 64) x 2 = 27.6 KB 65536 ctx -> 1.8 GB 262144 -> 7.2 GB 1048576 -> 29 GB ``` Against ~38 GB left free by the weights on a 5-GPU host. **16384 was never a hardware limit, just an untested default** — and it could not accept a single Claude Code request. What constrains context here is the weights sharing the cards, never the cache. ## Cost reality check Runpod hosts an official Kimi-K3 endpoint (`moonshot-kimi`). Measured: 300 tokens in 10.281 s, billed $0.004797. ``` speed $/1M output quality official 29.2 tok/s ~16 full precision, French works ours (1 stream) 7.2 tok/s ~85 1.14 bpw, French broken ours (4 streams) 14.5 tok/s ~42 ``` The official endpoint is 4x faster and 5.7x cheaper. Self-hosting buys privacy, no per-token billing, and context control — not price or quality. Worth restating whenever the effort seems to justify itself on cost. ## No smaller GGUF exists Ours is the most compact Kimi-K3 published, because it is the Width50 variant (`moe_intermediate_size` 1536 rather than 3072) — which is also why French is broken. ``` ours REAP448 Width50 IQ1_S 181 GiB 0xTank REAP568 UD-IQ1_S 246 GiB mmnga-o REAP50 UD-IQ1_S 305 GiB prometheusAIR REAP55 IQ1_M 319 GiB hellohazime REAP640ja 411 GiB ``` A "light" build for a single A40 (46 GB, ~37 GB of weights after a 256k KV) does not exist and would have to be produced — roughly a 64-expert REAP. ## Under $1 per million tokens is not reachable with K3 At $0.44/h per A40, $1/1M output requires **122 tok/s aggregate per card**. The 5× pod at $2.20/h would need 611 tok/s; it delivers 14.5. That is a 42x gap, not a tuning problem. What does reach it, from HF configs (throughput figures are bandwidth-derived estimates, not measured): ``` Qwen3-Coder-30B-A3B 30B/3B active ~30 GB FP8 KV 12 GiB@256k 1x A40 ~$0.10-0.25/1M Qwen3-Coder-Next 80B/3B active ~74 GB FP8 KV 12 GiB@256k 2x A40 ~$0.20-0.80/1M Qwen3.8-27B dense ~26 GB FP8 KV 32 GiB@256k 256k does not fit on 1 card ``` Served by **vLLM, not llama.cpp** — tensor parallelism and continuous batching are exactly what the concurrency measurement above shows llama.cpp lacking. ## Operational trap: bare `wait` kills the hot-reload watcher `llama-server` is started with `&`, so a bare `wait` anywhere later in the script waits for *it* too, forever. The server keeps serving, the selftest never finishes, the watcher never starts, and the pod silently ignores every published update. Symptom: `/props` reports the old config long after a push. Always `wait "$pid"` with an explicit pid. --- # Session 3 — 256k validated, and a measured alternative that dominates it ## Kimi-K3 at 262144 context: it works, and it costs nothing in decode `--ctx-size 262144`, all-VRAM on 5x A40, 39.5-42.5 GB used per card: ``` n_ctx 262144 (was 16384 -- never a hardware limit, just a default) decode 7.25 tok/s (unchanged from 16k: 7.2-7.3) ``` 16x the context for no throughput cost, exactly as the KDA/MLA geometry predicted. The DSpark rungs OOM'd at this context and the ladder correctly fell through to all-VRAM without a draft — no loss, DSpark was measured at zero gain. ## Prefill degrades gracefully, and the proxy times out before it finishes A cold 38k prefill takes ~230 s. Runpod's HTTP proxy sits behind Cloudflare, which cuts the connection at ~125 s (`error code: 524`). The workaround is to build the prefix in chunks — each request extends the cached prefix — which also measures the degradation curve a single call would have hidden: ``` chunk 1 +8400 tok @ 8k ctx 228.6 tok/s chunk 2 +8404 tok @ 17k ctx 203.6 tok/s chunk 3 +8404 tok @ 25k ctx 182.4 tok/s chunk 4 +8605 tok @ 34k ctx 164.1 tok/s chunk 5 +8704 tok @ 43k ctx 149.8 tok/s ``` -34% across 5x the context, far better than quadratic — 69 of 93 layers are KDA (linear), only 24 are MLA. 42,517 tokens prefilled in 3 min 55 total. Re-submitting a cached 32k prompt returns `prompt_n=4` in 0.6 s. ## vLLM's Kimi-K3 support is real and unreachable vLLM and SGLang both ship day-0 K3 support, including the two things that failed here: a hybrid prefix cache over recurrent KDA state, and working DSpark block-diffusion speculative decoding. None of it is usable on A40s, because every vLLM-loadable K3 checkpoint is the full unpruned model: ``` nvidia/Kimi-K3-NVFP4 1499 GiB + requires Blackwell (sm_100) RedHatAI/Kimi-K3-NVFP4 1533 GiB + requires Blackwell RedHatAI/Kimi-K3-FP8-BLOCK 2626 GiB + requires sm_89+ ``` No AWQ, no GPTQ, no W4A16. The 181 GiB IQ1_S we run is llama.cpp-only, and llama.cpp is exactly what caps us. **The unlock would be quantising a REAP'd K3 to W4A16 for vLLM** — Code145 at 4-bit lands near 260 GB (6x A40 / 4x A100-80). That build, not any config change, is the only path to the good engine. ## The measured alternative Qwen3-Coder-30B-A3B-Instruct, AWQ 4-bit (16.9 GiB), vLLM 0.27.1, **one** A40 at $0.44/h, 262144 context with fp8 KV. Everything below is measured, not derived: ``` concurrency tokens duration aggregate per stream $/1M output 1 128 1.6 s 77.6 tok/s 77.6 $1.575 4 512 1.9 s 265.3 tok/s 66.3 $0.461 8 1024 2.1 s 496.8 tok/s 62.1 $0.246 16 2048 2.6 s 782.4 tok/s 48.9 $0.156 32 4096 3.5 s 1165.6 tok/s 36.4 $0.105 prefill 38,494 tok cold in 10.42 s = 3693 tok/s; warm 1.07 s ``` Head to head: ``` Kimi-K3 5x A40 Qwen3-Coder 1x A40 K3 official API hourly $2.20 $0.44 per-token decode, 1 stream 7.2 tok/s 77.6 tok/s 29.2 tok/s prefill @38k ~165 tok/s 3693 tok/s -- 32k prompt, cold 3 min 55 10 s -- context 262144 5 cards 1 card -- $/1M output ~85 $1.58 .. $0.105 ~16 French broken correct correct ``` Qwen is 10.8x on decode, 22x on prefill, on one fifth of the hardware, and it answers in French. The under-$1/1M target is met from two concurrent streams and reaches $0.105 at 32. Kimi-K3 self-hosted on A40s is not competitive on any axis measured here; it remains interesting only where the specific K3 model is the requirement. ## Two mistakes worth not repeating - Creating a pod from `create-pod` without a start command: the schema has no `args` field, so the container starts with the image default. Use `create-template` with `dockerStartCmd`, then deploy with `templateId`. - `--disable-log-requests` was removed in vLLM 0.27; it crash-loops the server. A pod's `args` cannot be edited after creation — only the template can — so a bad flag costs a full recreate. ## The context ceiling on 5x A40, and why q8 KV is not the way out ``` ctx 16384 7.2-7.3 tok/s ctx 262144 7.25 tok/s ctx 524288 7.29 tok/s <- ceiling, all-VRAM, no draft ctx 1048576 OOM: 29 GB of f16 MLA cache does not fit alongside the weights KV q8_0 rejected outright, at every context ``` Decode is flat across a 32x range of context. That is the KDA/MLA hybrid doing exactly what it is for: only 24 of 93 layers cache at all, and the other 69 hold a recurrent state whose size is independent of context. **q8_0 KV cannot be used here**, and the failure signature says why. Every placement rung failed in 1.6-2.5 s with `failed to create llama_context from model` and *no* `cudaMalloc` line — a parameter rejection. Contrast the genuine OOM at 512k on the draft rung: 29.9 s, with `cudaMalloc failed: out of memory` printed. Fast failure with no allocator message means llama.cpp refused the configuration; slow failure with one means it tried and ran out. The 69 KDA layers carry a recurrent state llama.cpp will not quantise, so f16 is forced and the ceiling is set by memory: 512k fits, 1M does not. TurboQuant (Google, ICLR 2026) would take the cache to ~3 bits and put 1M within reach, but as of this run it is a paper with an open SGLang feature request (#21618) and no merged implementation in any engine. ## Qwen decode against context length ``` short context 77.6 tok/s 119k context 34.3 tok/s (prompt served warm from prefix cache) ``` Halves across 119k, and still 4.7x Kimi's short-context rate. Tool calling was verified on the OpenAI endpoint: two parallel calls, correct arguments, `finish_reason: tool_calls`. ## Runpod's HTTP proxy caps a single request at ~125 s Cloudflare sits in front of the pod proxy and returns `error code: 524`. A cold prefill longer than that cannot complete in one call — measured at ~40k on Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached prefix) or expose a TCP port instead. --- # Session 4 — speculative decoding, and why the regime decides everything n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40, `--speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'`. Same model, same hardware, same quant, two pods running side by side so the comparison is simultaneous rather than sequential: ``` A. pure generation (no prompt/output overlap) baseline 102.21 tok/s ngram 58.49 tok/s -43% B. code editing (output re-emits the input) baseline 130.16 tok/s ngram 242.39 tok/s +86% ``` **Speculation is not a free win — it is a bet on repetition.** With nothing to guess, every draft is rejected and the verification cost is pure loss; that is the -43%. Claude Code lives almost entirely in regime B (read a file, emit a modified version, repeat identifiers), so it is the right default *there* and the wrong default for prose. vLLM's telemetry, aggregated over both regimes: ``` drafts 207 draft tokens 1035 accepted 706 = 68.2% accepted per position: 179 / 156 / 131 / 120 / 120 mean accepted length 3.41 -> ~4.4 tokens emitted per forward pass ``` This is the measurement llama.cpp never produced: DSpark there exposed no `draft_n` at all and moved throughput by 0.0%. The mechanism was never broken — the engine was. ## The number that matters for the original goal Single stream, code editing, one A40 at $0.44/h: ``` 242.39 tok/s -> $0.50 per 1M output tokens ``` The under-$1/1M target is met **on a single stream**, without needing concurrency to amortise anything. For reference the same target on Kimi-K3 required 611 tok/s aggregate against 14.5 measured. Note also that the baseline itself reads higher here (102-130 tok/s) than the 77.6 measured earlier: that earlier figure was a 128-token request whose wall clock was dominated by per-request overhead. Longer outputs amortise it. Quote 77.6 for short replies and ~130 for sustained generation. --- # Session 5 — Kimi-Linear-48B-A3B: genuine Kimi, 1M context, one A40 `cyankiwi/Kimi-Linear-48B-A3B-Instruct-AWQ-4bit`, 28.4 GiB, vLLM 0.27.1, one A40 at $0.44/h. `KimiLinearForCausalLM`, `model_type: kimi_linear` — the same architecture family as K3, with the same MLA geometry (`kv_lora_rank 512`, `qk_rope_head_dim 64`). ## Why it fits where K3 could not ``` 27 layers = 7 MLA (full_attn_layers [4,8,12,16,20,24,27]) + 20 KDA KV/token = 7 x (512 + 64) x 2 = 7.88 KB (K3: 24 layers -> 27.6 KB) weights AWQ4 28.4 GiB 1,048,576 ctx 7.88 GiB KV => 36.3 GB on a 46 GB card ``` vLLM confirmed it at boot: **`GPU KV cache size: 1,416,566 tokens`** on a single card. K3 needed five A40s at $2.20/h to reach 524,288. ## Measured ``` flux aggregate per stream $/1M 1 67.4 67.4 1.814 4 206.6 51.6 0.592 8 352.4 44.0 0.347 16 556.1 34.8 0.220 32 779.9 24.4 0.157 prefill cold 6162 tok/s (better than Qwen3-Coder-30B's 3957) prefill warm 6430 tok/s ratio 1.04 -> NO prefix caching code editing 89.4 tok/s tool calling 2 parallel calls, correct arguments code correct; French fluent ``` ## n-gram speculation is BROKEN on this backend Rejection sampling is supposed to be lossless, so any output difference between speculative and non-speculative decoding is a bug. Here it corrupts code, reproducibly and identically across runs: ``` with speculation without add() return a -b):\n return a - return a - b\n\ndef multiply(a, fibonacci fibonacci(n-1(n-1) + fibonacci(n-2) fibonacci(n-1) + fibonacci(n-2) ``` The signature is a just-emitted fragment re-spliced at the wrong place — a bad draft accept. Plain prose and verbatim copying stayed perfect, which is why it took targeted probes to find: code is dense in the short repeats (parentheses, operators) that n-gram matching latches onto. It is also 3-4x *slower*, because vLLM disables CUDA graphs under speculation on `TritonMLABackend`: ``` flux with spec without ratio 1 23.1 71.2 3.1x 4 58.7 249.6 4.3x 32 203.9 781.0 3.8x $/1M @32 0.600 0.157 ``` **Turn speculation off for Kimi-Linear.** Worth reporting upstream. ## Two fixes that made it work at all - **Tokenizer.** `tokenization_kimi.py` imports `bytes_to_unicode` from `transformers.convert_slow_tokenizer`, removed in transformers >= 5.5.3 which every vLLM >= 0.24 requires. No Kimi-Linear repo ships a fast `tokenizer.json`, so the Python tokenizer is mandatory. Fixed by publishing `patdev/kimi-linear-tokenizer-fix`: same files with a local fallback definition, verified byte-identical to the transformers reference before deployment. Point vLLM at it with `--tokenizer`. - **Tool parser.** `hermes` yields zero calls. The chat template uses `<|tool_calls_section_begin|>` / `<|tool_call_begin|>`, the K2 token scheme — so the parser is `kimi_k2`. (`kimi_k3` is for `<|open|>`/`<|close|>`/`<|sep|>` and does not apply.) It also sets `skip_special_tokens=False`, which those markers need. - **Headroom.** `--gpu-memory-utilization 0.93` OOMs at the very last step, in the speculative rejection sampler asking for 160 MiB with 143 MiB left. 0.90 leaves room and still yields >1M tokens of KV. ## Head to head ``` Kimi-Linear 48B Qwen3-Coder-30B hourly $0.44 $0.44 decode 1 stream 67-71 tok/s 77.6-79.2 tok/s aggregate @32 780 1166 prefill cold 6162 3957 prefill warm 6430 63315 context 1,048,576 262,144 $/1M @32 0.157 0.105 tool calling yes yes French fluent correct brand GENUINE KIMI Qwen ``` Kimi-Linear wins on context (4x) and cold prefill (1.6x). It loses on decode and, decisively for an agent client, on prefix caching: **6430 vs 63315, ten times worse**, because vLLM does not yet support prefix caching over recurrent hybrid state. A 32k system prompt costs ~5 s every turn instead of ~0.5 s. ## Qwen3-Coder-Next 80B: no data Two attempts, both infrastructure failures, zero requests served. ``` attempt 1 hung 22 min in NCCL peer-to-peer init on 2x A40 attempt 2 NCCL_P2P_DISABLE=1 cleared that, then looped ~10 min on "No available shared memory broadcast block found in 60 seconds" ``` `NCCL_P2P_DISABLE=1` is required for tensor parallelism on Runpod A40 pairs. The second hang is unproven — the hypothesis is torch.compile on a 512-expert MoE, testable with `--enforce-eager`. Its boot log did confirm one thing: `enable_prefix_caching=False`, the same hybrid limitation as Kimi-Linear.