k3-a40-bootstrap / FINDINGS.md
patdev's picture
Upload FINDINGS.md with huggingface_hub
4448aee verified
|
Raw History Blame Contribute Delete
31.7 kB
# Kimi-K3 on 2× A40 — measured findings
Everything here was measured on RunPod pod `cumv30dk07dfdy`: 2× NVIDIA A40
(2 × 46068 MiB), 93 GB container RAM, 16 container vCPU, 240 GB ephemeral disk,
$0.92/h. Model: `mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf` UD-IQ1_S, 181 GiB.
Runtime: `mmnga/llama.cpp` branch `kimi-k3-width-support`, CUDA 12.8, sm_86.
Numbers without a measurement behind them are marked as estimates.
## Throughput
| Configuration | decode | prefill | static VRAM |
|---|---|---|---|
| autofit, 16 experts, greedy | **2.94 tok/s** avg | 1.2–4.1 tok/s | 87.7 GB |
| `--cpu-moe --moe-cache on` (unconfigured) | 2.81 tok/s | 0.9–2.4 tok/s | 60.2 GB |
| autofit, 8 experts | 3.15 tok/s | 1.88 tok/s | 87.4 GB |
Prefill never exceeding decode by much is the tell: this deployment is not
compute-bound on the GPU, it is dominated by whatever runs on the host.
## Result: 7.3 tok/s
| Step | decode | prefill | cost |
|---|---|---|---|
| starting point (previous effort, L40S) | 2.32 | — | — |
| 2× A40 after the fixes below | 2.94 | 1.2–4.1 | $0.92/h |
| 5× A40, *identical* placement | 5.44 | 6.8–8.0 | $2.20/h |
| 5× A40, everything VRAM-resident | **7.35** | **27.3** | $2.20/h |
3.2× on decode, 27× on prefill. Two independent factors, both of which this
investigation initially got wrong: **host vCPU** (which RunPod scales with GPU
count) and **full VRAM residency**. Neither is a tuning parameter.
Output quality at 7.3 tok/s: English prose and reasoning are correct ("The trip
duration is 3.5 hours"), code is correct (`fibonacci(n-1) + fibonacci(n-2)`),
and the model continues fluently in Chinese — Moonshot's other strong language.
French generation stays broken, which is the Width50 build, not the hardware.
**DSpark speculative decoding gives nothing.** Once the draft finally loaded —
after clearing four separate blockers — throughput was 7.32/7.33 against
7.33/7.35 without it, and llama-server exposed no `draft_n` / `draft_n_accepted`
telemetry to explain why. The predicted floor of 1.2× was still too optimistic;
the measured figure is 1.0×.
## The answer: host vCPU, not VRAM
The same configuration, on the same model, using the same two GPUs — three
further GPUs sitting idle at 267 MiB — runs at **1.9× the speed** on a bigger
host:
| Pod | GPUs | vCPU | RAM | decode |
|---|---|---|---|---|
| `cumv30dk07dfdy` | 2× A40 | **16** | 93 GB | 2.31 / 2.98 / 2.88 |
| `okinf3ln9k8tsp` | 5× A40 | **41** | 234 GB | 4.86 / 5.28 / 5.44 / 5.64 |
Prefill moves even harder: 0.99 → 8.01 tok/s.
RunPod scales vCPU with GPU count, so renting more GPUs buys host parallelism as
much as it buys VRAM. Every placement experiment below kept 70–90 GB of weights
being processed by 16 saturated threads, which is why moving that work between
CPU and GPU changed nothing: the CPU side was the wall in all of them.
This is the mechanism the section below says was missing. It was found by
accident — a 5-GPU pod picked up the 2-GPU placement script before the guard
was published, and ran the *identical* configuration on a larger host.
## Why the placement experiments all looked the same
~3 tok/s is the ceiling for this model on 2× A40, and the reason is not the one
that seems obvious.
Throughput is **invariant** across configurations that differ enormously:
| Configuration | VRAM used | attention | CPU-side bytes/token | decode |
|---|---|---|---|---|
| autofit, 16 experts | 87.7 GB | 48 layers on host | ~33 GB | 2.46 / 3.00 |
| `--cpu-moe` + moe-cache | 60.2 GB | all in VRAM | ~4 GB | 2.21 / 2.68 |
| `--n-cpu-moe 86` | 69.9 GB | all in VRAM | ~4 GB | 2.51 / 3.06 |
| manual balanced `-ot` | 87.8 GB | all in VRAM, 44.3/43.5 split | ~4 GB | 2.31 / 2.98 |
| autofit, 8 experts | 87.4 GB | 48 layers on host | ~17 GB | 3.15 |
| `-ngl 0` (no GPU) | 1.5 GB | all on host | all | **0.65** |
The manual placement is the cleanest of these: explicit `-ot ...=CUDA0/CUDA1`
layer ranges fill both cards to 96%/94% with every layer's attention resident
and only 72 layers' routed experts on the host. It took three calibration
rounds, each guided by the shortfall rather than a guess — 45256 MiB over, then
1038 MiB over, then fitting. It performs exactly like everything else.
Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%.
**No validated mechanism.** Three hypotheses were proposed and each was refuted
by measurement:
1. *Routed-expert CPU compute dominates* — refuted: halving experts per token
gave +28%, not the ~100% that would follow.
2. *Attention re-read from host RAM dominates* — refuted: `--n-cpu-moe 86` puts
every layer's attention in VRAM and cuts host bytes ~8×, and changed nothing.
3. *CPU↔GPU boundary crossings dominate* — refuted by comparing the two: autofit
splits into contiguous GPU-then-CPU blocks (1–2 crossings per token) while
`--n-cpu-moe 86` crosses at 86 layers (~172 per token). If crossings
dominated, autofit would be far faster. The two are within 2% of each other.
What stands is empirical, not theoretical: **no rebalancing of CPU against GPU
within 92 GB of VRAM improves anything**, across placements spanning 60–88 GB of
VRAM and 4–33 GB of host traffic per token.
**The GPUs are not decorative, though — verified.** Running with `-ngl 0`, so
that essentially nothing sits in VRAM (1265 MiB and 269 MiB), gives **0.65
tok/s** against 2.51 for the same prompt with autofit. The two A40s are worth
**3.9×**. That check was run specifically because, if the GPUs had contributed
nothing, "buy more VRAM" would have been the wrong advice.
So the direction — fit more of the model in VRAM — is sound, and a ggml
microbenchmark says the magnitude is large. `test-backend-ops` timing
`MUL_MAT_ID` on the K3 expert shapes (MXFP4, 3584→3072) on an L40S:
```
8 experts, C=1 29.25 µs
16 experts, C=1 59.83 µs
8 experts, C=7 159.53 µs
16 experts, C=7 318.97 µs ~7.73 TFLOPS
```
At 59.83 µs per MoE tensor and 92 MoE layers, the routed-expert work costs
**5.5 ms/token when the experts are VRAM-resident** — a ~180 tok/s ceiling from
that term alone. We measure 333 ms/token, so GPU MoE compute is about **1.6% of
the time**; the other 98% is the host path.
That makes the naive linear extrapolation (5.4 tok/s at full residency) almost
certainly far too pessimistic: full residency does not scale the existing
regime, it deletes the term that dominates it.
## What did NOT work, and why
**Halving experts per token (16 → 8).** Bought only +28% (2.46 → 3.15 on the
same prompt) and destroyed the output — `# # # #` after six tokens. REAP had
already cut 896 experts to 448 with the router calibrated for 16 active. The
small gain also *disproved* the hypothesis that routed-expert CPU compute was
the bottleneck: halving it would otherwise have nearly doubled throughput.
**The CUDA MoE expert cache** (`llama-k3-moe-cache.patch`, from the Code145
build). Applies cleanly to mmnga's branch once `tests/` is excluded, compiles
for sm_86, and `--moe-cache` appears in `--help`. Measured 14% *slower* on
decode and roughly half the prefill. Caveat: that run set none of the seven
`GGML_CUDA_MOE_CACHE_*` variables its own production launcher uses, so the
verdict is against an unconfigured cache, not the design.
**DSpark speculative decoding.** Three separate blockers, each found only after
clearing the previous one:
1. `--draft-max` was removed; the flag is now `--spec-draft-n-max`.
2. The published GGUF declares `general.architecture = "dflash-draft"`;
llama.cpp registers `LLM_ARCH_DFLASH` as plain `"dflash"`. Fixed by
`fix_dspark_arch.py` (renames the architecture and its 30 namespaced KV
keys, verified by reopening and comparing every tensor's shape, dtype and
byte size).
3. The rewritten file still fails: *"DFlash model requires 'target_layers' in
GGUF metadata"*. The third-party converter also names tensors
`dflash.dspark.markov.w1` where llama.cpp expects `markov_w1`. The clean fix
is to convert from `RadixArk/Kimi-K3-DSpark` with llama.cpp's own converter.
Expected ceiling if it ever loads: **1.2–2×**, not the ~3× vLLM reports. On a
MoE each drafted token routes to its own experts, so batch verification
amortises only attention and shared experts, not the routed reads.
**`GGML_OP_OFFLOAD_MIN_BATCH=1`.** ggml can copy only the *selected* experts of
an offloaded `MUL_MAT_ID` to the GPU instead of computing them on the host, but
the path is gated on batch size and the default of 32 means it never fires
during decode. Forcing it on **halved throughput**: 2.46 → 1.13 and 3.00 → 1.38
tok/s. With 16 experts across 68 offloaded layers that is 1088 separate small
transfers per token, so it is latency-bound, not bandwidth-bound, and loses to
host compute. The upstream default is right.
**`--n-cpu-moe` at any rung (62/68/74/80).** Never loaded. Tensor overrides set
`model_params::tensor_buft_overrides`, which aborts autofit exactly like `-ngl`
and `--tensor-split` do — and once autofit is off, **neither `--split-mode
layer` nor `--tensor-split` distributes anything**: every allocation lands on
device 1 while device 0 stays empty. Rung 80 missed by 135 MiB (46203 requested
against 46068 available) with all attention plus twelve expert layers on one
card. Autofit is the only mechanism that actually uses both GPUs; balancing
manually requires explicit `-ot ...=CUDA0` / `=CUDA1` layer ranges.
**TurboQuant / KV-cache quantisation.** Not applicable: 69 of 93 layers are KDA
recurrent and carry no growing KV, so the whole cache is about 0.8 GB.
**ONNX custom ops** from earlier work target Qwen3.5 on SM75, not Kimi-K3 on
sm_86, and ONNX Runtime has no `kimi_k3` architecture.
## Output quality
English and code are genuinely usable; French is not.
```
"The capital of Japan is" -> Tokyo, Washington, Stockholm, Berlin correct
"train 14:00 -> 17:30" -> "The trip takes 3 hours 30 minutes." correct
"def invert(d):" -> return {v: k for k, v in d.items()} correct
French free-form -> "cunce", "s'fotrenci", "zocneine" non-words
```
French *knowledge* survives — "La Revolution francaise a commence en" continues
"1789" — but French *generation* is corrupted at the sub-word level. The prime
suspect is Width50: this build halves `moe_intermediate_size` from 3072 to 1536.
### Sampler, six configurations on one prompt
| Setting | Result |
|---|---|
| greedy, rp 1.0 | facts correct, then hard repetition loop |
| greedy, rp 1.05 / last_n 64 | no loop, but "Germany is PLN" — the penalty pushes it off correct answers while enumerating |
| greedy, rp 1.15 / last_n 256 | drifts into SQL-flavoured nonsense |
| temp 0.3, min_p 0.05 | fluent, hallucinates ("Tokyo Island", "Sumaru") |
| temp 0.6, top_k 20 | degrades |
| **temp 0.3, min_p 0.1, rp 1.05** | fluent, no loop, nothing glaringly wrong — chosen default |
Repetition penalty measurably *reduces factual accuracy* on this model. Clients
wanting maximum accuracy should still ask for `temperature: 0`.
## Deployment mechanics
`common_fit_params` sizes GPU placement against actually-free VRAM across all
devices, but **aborts the moment any placement parameter is pinned**:
```
-ngl 99 -> "n_gpu_layers already set by user, abort"
--tensor-split -> "model_params::tensor_split already set by user, abort"
```
Both aborts produced the same failure: 84790 MiB shoved onto device 1 alone,
cudaMalloc OOM, device 0 untouched. With no placement flags, autofit fills both
cards to 95%. `--cpu-moe` and `--n-cpu-moe` do *not* abort it.
**Boot time** went from 14m28s to ~4 minutes by publishing the compiled
binaries to the Hub (46 MB tarball, 2 s to extract vs 8m34s to compile). The
artifact name pins base image, CUDA version and GPU arch so a mismatched
tarball can never be picked up silently, and it is smoke-tested by running
`llama-server --version` before it is trusted.
**RunPod costs.** A stopped pod's Volume disk bills at $0.20/GB/month, double
the running rate — a 1600 GB volume cost $0.44/h doing nothing, which was 99%
of one day's spend. Container disk is erased on stop but still billed, so for
re-downloadable data the rule is: always Terminate, never Stop. A stopped pod
is also pinned to its host and may refuse to restart while the same GPU type
shows High stock elsewhere.
## Kimi-K3-256K-REAP-Code145
Finalised during this session. All 96 shards were already on the Hub (347.5 GB);
only `config.json`, the index and the tokenizer were missing, and
`finalize_code145.py` — which downloads only metadata, never the weights — had
never been run. The repo is now loadable.
```
num_experts 145
num_experts_per_token 16
moe_intermediate_size 3072 full width, unlike the 1536 we serve
source_index_keys 497 220
dropped_expert_keys 414 552
output_index_keys 82 668 = 497220 - 414552
renamed_expert_keys 80 040 = 92 layers x 145 experts x 6 tensors
```
Full width makes it the serious candidate for working French. It does not fit
2× A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4× A40 at
$1.80/h, or an imatrix-guided quantisation down to ~1.7 bit.
---
# Session 2 — what actually bounds decode, and what long context costs
## Decode is serialisation-bound, not bandwidth-bound
Six placements had moved 60–88 GB of weights between CPU and GPU without
changing throughput, and no mechanism had survived measurement. The question
none of them asked was whether the hardware is *idle* between tokens. It is.
Aggregate throughput against concurrent streams, 5× A40, 64 tokens each,
distinct prefixes so no stream rides another's prompt cache:
```
1 stream 5.56 tok/s aggregate 5.56 per stream
2 streams 9.85 tok/s aggregate 4.93 per stream 1.77x
4 streams 14.53 tok/s aggregate 3.63 per stream 2.61x
```
Concurrency scales, so single-stream decode leaves silicon idle. The
bandwidth arithmetic agrees independently: ~7 GB of weights are touched per
token, and at 138 ms/token that is **~51 GB/s against 696 GB/s available** on
the active card — 7% of the roof. llama.cpp splits layers across GPUs, so at
batch 1 exactly one card computes while the other four wait.
This retires the last open question from session 1 and explains why every
kernel-level optimisation attempt returned nothing: the expert GEMM was never
the constraint.
## Sustained decode on real output: 7.22 tok/s
The 7.3 tok/s headline was measured on selftests, three of whose five prompts
produce degenerate repetition (`# [CS_ah]`, `et d'et d'`) — and a repetition
loop decodes cheap. Re-measured on 600 tokens of coherent prose: **7.22 tok/s**.
The number holds; it is now measured on output worth having.
## Prefill: 167 tok/s cold, ~free warm
Every selftest prompt is 6–22 tokens, so its "prefill tok/s" is per-request
overhead and says nothing. Measured properly, 3408-token prompt:
```
cold 3408 tok in 20.4 s = 167 tok/s (23x the decode rate — prefill batches)
warm 4 tok in 0.53 s (slot prompt cache reprocessed only the delta)
```
For a Claude Code client sending a 32k system prompt: **~3 min once**, then
only the changed tokens on every later turn.
## Long context is nearly free on this architecture
From the config, not from assumption: `kv_lora_rank 512`, `qk_rope_head_dim
64`, `full_attn_layers` every 4th layer, `max_position_embeddings 1048576`.
```
24 of 93 layers cache at all (MLA); the other 69 are KDA — recurrent state,
constant regardless of context length.
KV per token = 24 x (512 + 64) x 2 = 27.6 KB
65536 ctx -> 1.8 GB 262144 -> 7.2 GB 1048576 -> 29 GB
```
Against ~38 GB left free by the weights on a 5-GPU host. **16384 was never a
hardware limit, just an untested default** — and it could not accept a single
Claude Code request. What constrains context here is the weights sharing the
cards, never the cache.
## Cost reality check
Runpod hosts an official Kimi-K3 endpoint (`moonshot-kimi`). Measured:
300 tokens in 10.281 s, billed $0.004797.
```
speed $/1M output quality
official 29.2 tok/s ~16 full precision, French works
ours (1 stream) 7.2 tok/s ~85 1.14 bpw, French broken
ours (4 streams) 14.5 tok/s ~42
```
The official endpoint is 4x faster and 5.7x cheaper. Self-hosting buys
privacy, no per-token billing, and context control — not price or quality.
Worth restating whenever the effort seems to justify itself on cost.
## No smaller GGUF exists
Ours is the most compact Kimi-K3 published, because it is the Width50 variant
(`moe_intermediate_size` 1536 rather than 3072) — which is also why French is
broken.
```
ours REAP448 Width50 IQ1_S 181 GiB
0xTank REAP568 UD-IQ1_S 246 GiB
mmnga-o REAP50 UD-IQ1_S 305 GiB
prometheusAIR REAP55 IQ1_M 319 GiB
hellohazime REAP640ja 411 GiB
```
A "light" build for a single A40 (46 GB, ~37 GB of weights after a 256k KV)
does not exist and would have to be produced — roughly a 64-expert REAP.
## Under $1 per million tokens is not reachable with K3
At $0.44/h per A40, $1/1M output requires **122 tok/s aggregate per card**. The
5× pod at $2.20/h would need 611 tok/s; it delivers 14.5. That is a 42x gap,
not a tuning problem.
What does reach it, from HF configs (throughput figures are bandwidth-derived
estimates, not measured):
```
Qwen3-Coder-30B-A3B 30B/3B active ~30 GB FP8 KV 12 GiB@256k 1x A40 ~$0.10-0.25/1M
Qwen3-Coder-Next 80B/3B active ~74 GB FP8 KV 12 GiB@256k 2x A40 ~$0.20-0.80/1M
Qwen3.8-27B dense ~26 GB FP8 KV 32 GiB@256k 256k does not fit on 1 card
```
Served by **vLLM, not llama.cpp** — tensor parallelism and continuous batching
are exactly what the concurrency measurement above shows llama.cpp lacking.
## Operational trap: bare `wait` kills the hot-reload watcher
`llama-server` is started with `&`, so a bare `wait` anywhere later in the
script waits for *it* too, forever. The server keeps serving, the selftest
never finishes, the watcher never starts, and the pod silently ignores every
published update. Symptom: `/props` reports the old config long after a push.
Always `wait "$pid"` with an explicit pid.
---
# Session 3 — 256k validated, and a measured alternative that dominates it
## Kimi-K3 at 262144 context: it works, and it costs nothing in decode
`--ctx-size 262144`, all-VRAM on 5x A40, 39.5-42.5 GB used per card:
```
n_ctx 262144 (was 16384 -- never a hardware limit, just a default)
decode 7.25 tok/s (unchanged from 16k: 7.2-7.3)
```
16x the context for no throughput cost, exactly as the KDA/MLA geometry
predicted. The DSpark rungs OOM'd at this context and the ladder correctly fell
through to all-VRAM without a draft — no loss, DSpark was measured at zero gain.
## Prefill degrades gracefully, and the proxy times out before it finishes
A cold 38k prefill takes ~230 s. Runpod's HTTP proxy sits behind Cloudflare,
which cuts the connection at ~125 s (`error code: 524`). The workaround is to
build the prefix in chunks — each request extends the cached prefix — which
also measures the degradation curve a single call would have hidden:
```
chunk 1 +8400 tok @ 8k ctx 228.6 tok/s
chunk 2 +8404 tok @ 17k ctx 203.6 tok/s
chunk 3 +8404 tok @ 25k ctx 182.4 tok/s
chunk 4 +8605 tok @ 34k ctx 164.1 tok/s
chunk 5 +8704 tok @ 43k ctx 149.8 tok/s
```
-34% across 5x the context, far better than quadratic — 69 of 93 layers are
KDA (linear), only 24 are MLA. 42,517 tokens prefilled in 3 min 55 total.
Re-submitting a cached 32k prompt returns `prompt_n=4` in 0.6 s.
## vLLM's Kimi-K3 support is real and unreachable
vLLM and SGLang both ship day-0 K3 support, including the two things that
failed here: a hybrid prefix cache over recurrent KDA state, and working DSpark
block-diffusion speculative decoding. None of it is usable on A40s, because
every vLLM-loadable K3 checkpoint is the full unpruned model:
```
nvidia/Kimi-K3-NVFP4 1499 GiB + requires Blackwell (sm_100)
RedHatAI/Kimi-K3-NVFP4 1533 GiB + requires Blackwell
RedHatAI/Kimi-K3-FP8-BLOCK 2626 GiB + requires sm_89+
```
No AWQ, no GPTQ, no W4A16. The 181 GiB IQ1_S we run is llama.cpp-only, and
llama.cpp is exactly what caps us. **The unlock would be quantising a REAP'd K3
to W4A16 for vLLM** — Code145 at 4-bit lands near 260 GB (6x A40 / 4x A100-80).
That build, not any config change, is the only path to the good engine.
## The measured alternative
Qwen3-Coder-30B-A3B-Instruct, AWQ 4-bit (16.9 GiB), vLLM 0.27.1, **one** A40 at
$0.44/h, 262144 context with fp8 KV. Everything below is measured, not derived:
```
concurrency tokens duration aggregate per stream $/1M output
1 128 1.6 s 77.6 tok/s 77.6 $1.575
4 512 1.9 s 265.3 tok/s 66.3 $0.461
8 1024 2.1 s 496.8 tok/s 62.1 $0.246
16 2048 2.6 s 782.4 tok/s 48.9 $0.156
32 4096 3.5 s 1165.6 tok/s 36.4 $0.105
prefill 38,494 tok cold in 10.42 s = 3693 tok/s; warm 1.07 s
```
Head to head:
```
Kimi-K3 5x A40 Qwen3-Coder 1x A40 K3 official API
hourly $2.20 $0.44 per-token
decode, 1 stream 7.2 tok/s 77.6 tok/s 29.2 tok/s
prefill @38k ~165 tok/s 3693 tok/s --
32k prompt, cold 3 min 55 10 s --
context 262144 5 cards 1 card --
$/1M output ~85 $1.58 .. $0.105 ~16
French broken correct correct
```
Qwen is 10.8x on decode, 22x on prefill, on one fifth of the hardware, and it
answers in French. The under-$1/1M target is met from two concurrent streams
and reaches $0.105 at 32. Kimi-K3 self-hosted on A40s is not competitive on any
axis measured here; it remains interesting only where the specific K3 model is
the requirement.
## Two mistakes worth not repeating
- Creating a pod from `create-pod` without a start command: the schema has no
`args` field, so the container starts with the image default. Use
`create-template` with `dockerStartCmd`, then deploy with `templateId`.
- `--disable-log-requests` was removed in vLLM 0.27; it crash-loops the server.
A pod's `args` cannot be edited after creation — only the template can — so a
bad flag costs a full recreate.
## The context ceiling on 5x A40, and why q8 KV is not the way out
```
ctx 16384 7.2-7.3 tok/s
ctx 262144 7.25 tok/s
ctx 524288 7.29 tok/s <- ceiling, all-VRAM, no draft
ctx 1048576 OOM: 29 GB of f16 MLA cache does not fit alongside the weights
KV q8_0 rejected outright, at every context
```
Decode is flat across a 32x range of context. That is the KDA/MLA hybrid doing
exactly what it is for: only 24 of 93 layers cache at all, and the other 69 hold
a recurrent state whose size is independent of context.
**q8_0 KV cannot be used here**, and the failure signature says why. Every
placement rung failed in 1.6-2.5 s with `failed to create llama_context from
model` and *no* `cudaMalloc` line — a parameter rejection. Contrast the genuine
OOM at 512k on the draft rung: 29.9 s, with `cudaMalloc failed: out of memory`
printed. Fast failure with no allocator message means llama.cpp refused the
configuration; slow failure with one means it tried and ran out. The 69 KDA
layers carry a recurrent state llama.cpp will not quantise, so f16 is forced and
the ceiling is set by memory: 512k fits, 1M does not.
TurboQuant (Google, ICLR 2026) would take the cache to ~3 bits and put 1M
within reach, but as of this run it is a paper with an open SGLang feature
request (#21618) and no merged implementation in any engine.
## Qwen decode against context length
```
short context 77.6 tok/s
119k context 34.3 tok/s (prompt served warm from prefix cache)
```
Halves across 119k, and still 4.7x Kimi's short-context rate. Tool calling was
verified on the OpenAI endpoint: two parallel calls, correct arguments,
`finish_reason: tool_calls`.
## Runpod's HTTP proxy caps a single request at ~125 s
Cloudflare sits in front of the pod proxy and returns `error code: 524`. A cold
prefill longer than that cannot complete in one call — measured at ~40k on
Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached
prefix) or expose a TCP port instead.
---
# Session 4 — speculative decoding, and why the regime decides everything
n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40,
`--speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'`.
Same model, same hardware, same quant, two pods running side by side so the
comparison is simultaneous rather than sequential:
```
A. pure generation (no prompt/output overlap)
baseline 102.21 tok/s
ngram 58.49 tok/s -43%
B. code editing (output re-emits the input)
baseline 130.16 tok/s
ngram 242.39 tok/s +86%
```
**Speculation is not a free win — it is a bet on repetition.** With nothing to
guess, every draft is rejected and the verification cost is pure loss; that is
the -43%. Claude Code lives almost entirely in regime B (read a file, emit a
modified version, repeat identifiers), so it is the right default *there* and
the wrong default for prose.
vLLM's telemetry, aggregated over both regimes:
```
drafts 207 draft tokens 1035 accepted 706 = 68.2%
accepted per position: 179 / 156 / 131 / 120 / 120
mean accepted length 3.41 -> ~4.4 tokens emitted per forward pass
```
This is the measurement llama.cpp never produced: DSpark there exposed no
`draft_n` at all and moved throughput by 0.0%. The mechanism was never broken —
the engine was.
## The number that matters for the original goal
Single stream, code editing, one A40 at $0.44/h:
```
242.39 tok/s -> $0.50 per 1M output tokens
```
The under-$1/1M target is met **on a single stream**, without needing
concurrency to amortise anything. For reference the same target on Kimi-K3
required 611 tok/s aggregate against 14.5 measured.
Note also that the baseline itself reads higher here (102-130 tok/s) than the
77.6 measured earlier: that earlier figure was a 128-token request whose wall
clock was dominated by per-request overhead. Longer outputs amortise it. Quote
77.6 for short replies and ~130 for sustained generation.
---
# Session 5 — Kimi-Linear-48B-A3B: genuine Kimi, 1M context, one A40
`cyankiwi/Kimi-Linear-48B-A3B-Instruct-AWQ-4bit`, 28.4 GiB, vLLM 0.27.1, one
A40 at $0.44/h. `KimiLinearForCausalLM`, `model_type: kimi_linear` — the same
architecture family as K3, with the same MLA geometry (`kv_lora_rank 512`,
`qk_rope_head_dim 64`).
## Why it fits where K3 could not
```
27 layers = 7 MLA (full_attn_layers [4,8,12,16,20,24,27]) + 20 KDA
KV/token = 7 x (512 + 64) x 2 = 7.88 KB (K3: 24 layers -> 27.6 KB)
weights AWQ4 28.4 GiB
1,048,576 ctx 7.88 GiB KV => 36.3 GB on a 46 GB card
```
vLLM confirmed it at boot: **`GPU KV cache size: 1,416,566 tokens`** on a single
card. K3 needed five A40s at $2.20/h to reach 524,288.
## Measured
```
flux aggregate per stream $/1M
1 67.4 67.4 1.814
4 206.6 51.6 0.592
8 352.4 44.0 0.347
16 556.1 34.8 0.220
32 779.9 24.4 0.157
prefill cold 6162 tok/s (better than Qwen3-Coder-30B's 3957)
prefill warm 6430 tok/s ratio 1.04 -> NO prefix caching
code editing 89.4 tok/s
tool calling 2 parallel calls, correct arguments
code correct; French fluent
```
## n-gram speculation is BROKEN on this backend
Rejection sampling is supposed to be lossless, so any output difference between
speculative and non-speculative decoding is a bug. Here it corrupts code,
reproducibly and identically across runs:
```
with speculation without
add() return a -b):\n return a - return a - b\n\ndef multiply(a,
fibonacci fibonacci(n-1(n-1) + fibonacci(n-2) fibonacci(n-1) + fibonacci(n-2)
```
The signature is a just-emitted fragment re-spliced at the wrong place — a bad
draft accept. Plain prose and verbatim copying stayed perfect, which is why it
took targeted probes to find: code is dense in the short repeats (parentheses,
operators) that n-gram matching latches onto.
It is also 3-4x *slower*, because vLLM disables CUDA graphs under speculation
on `TritonMLABackend`:
```
flux with spec without ratio
1 23.1 71.2 3.1x
4 58.7 249.6 4.3x
32 203.9 781.0 3.8x
$/1M @32 0.600 0.157
```
**Turn speculation off for Kimi-Linear.** Worth reporting upstream.
## Two fixes that made it work at all
- **Tokenizer.** `tokenization_kimi.py` imports `bytes_to_unicode` from
`transformers.convert_slow_tokenizer`, removed in transformers >= 5.5.3 which
every vLLM >= 0.24 requires. No Kimi-Linear repo ships a fast `tokenizer.json`,
so the Python tokenizer is mandatory. Fixed by publishing
`patdev/kimi-linear-tokenizer-fix`: same files with a local fallback
definition, verified byte-identical to the transformers reference before
deployment. Point vLLM at it with `--tokenizer`.
- **Tool parser.** `hermes` yields zero calls. The chat template uses
`<|tool_calls_section_begin|>` / `<|tool_call_begin|>`, the K2 token scheme —
so the parser is `kimi_k2`. (`kimi_k3` is for `<|open|>`/`<|close|>`/`<|sep|>`
and does not apply.) It also sets `skip_special_tokens=False`, which those
markers need.
- **Headroom.** `--gpu-memory-utilization 0.93` OOMs at the very last step, in
the speculative rejection sampler asking for 160 MiB with 143 MiB left. 0.90
leaves room and still yields >1M tokens of KV.
## Head to head
```
Kimi-Linear 48B Qwen3-Coder-30B
hourly $0.44 $0.44
decode 1 stream 67-71 tok/s 77.6-79.2 tok/s
aggregate @32 780 1166
prefill cold 6162 3957
prefill warm 6430 63315
context 1,048,576 262,144
$/1M @32 0.157 0.105
tool calling yes yes
French fluent correct
brand GENUINE KIMI Qwen
```
Kimi-Linear wins on context (4x) and cold prefill (1.6x). It loses on decode
and, decisively for an agent client, on prefix caching: **6430 vs 63315, ten
times worse**, because vLLM does not yet support prefix caching over recurrent
hybrid state. A 32k system prompt costs ~5 s every turn instead of ~0.5 s.
## Qwen3-Coder-Next 80B: no data
Two attempts, both infrastructure failures, zero requests served.
```
attempt 1 hung 22 min in NCCL peer-to-peer init on 2x A40
attempt 2 NCCL_P2P_DISABLE=1 cleared that, then looped ~10 min on
"No available shared memory broadcast block found in 60 seconds"
```
`NCCL_P2P_DISABLE=1` is required for tensor parallelism on Runpod A40 pairs.
The second hang is unproven — the hypothesis is torch.compile on a 512-expert
MoE, testable with `--enforce-eager`. Its boot log did confirm one thing:
`enable_prefix_caching=False`, the same hybrid limitation as Kimi-Linear.