Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
|
Download FINDINGS.md from patdev/k3-a40-bootstrap: direct link, hf CLI and curl.
- Browser
- Download file 31.7 kB
-
https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/FINDINGS.md
- Command line
-
hf download hf://patdev/k3-a40-bootstrap/FINDINGS.md
-
curl -L -o FINDINGS.md https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/FINDINGS.md
31.7 kB
| # Kimi-K3 on 2× A40 — measured findings | |
| Everything here was measured on RunPod pod `cumv30dk07dfdy`: 2× NVIDIA A40 | |
| (2 × 46068 MiB), 93 GB container RAM, 16 container vCPU, 240 GB ephemeral disk, | |
| $0.92/h. Model: `mmnga-o/Kimi-K3-REAP50-Width50-UD-gguf` UD-IQ1_S, 181 GiB. | |
| Runtime: `mmnga/llama.cpp` branch `kimi-k3-width-support`, CUDA 12.8, sm_86. | |
| Numbers without a measurement behind them are marked as estimates. | |
| ## Throughput | |
| | Configuration | decode | prefill | static VRAM | | |
| |---|---|---|---| | |
| | autofit, 16 experts, greedy | **2.94 tok/s** avg | 1.2–4.1 tok/s | 87.7 GB | | |
| | `--cpu-moe --moe-cache on` (unconfigured) | 2.81 tok/s | 0.9–2.4 tok/s | 60.2 GB | | |
| | autofit, 8 experts | 3.15 tok/s | 1.88 tok/s | 87.4 GB | | |
| Prefill never exceeding decode by much is the tell: this deployment is not | |
| compute-bound on the GPU, it is dominated by whatever runs on the host. | |
| ## Result: 7.3 tok/s | |
| | Step | decode | prefill | cost | | |
| |---|---|---|---| | |
| | starting point (previous effort, L40S) | 2.32 | — | — | | |
| | 2× A40 after the fixes below | 2.94 | 1.2–4.1 | $0.92/h | | |
| | 5× A40, *identical* placement | 5.44 | 6.8–8.0 | $2.20/h | | |
| | 5× A40, everything VRAM-resident | **7.35** | **27.3** | $2.20/h | | |
| 3.2× on decode, 27× on prefill. Two independent factors, both of which this | |
| investigation initially got wrong: **host vCPU** (which RunPod scales with GPU | |
| count) and **full VRAM residency**. Neither is a tuning parameter. | |
| Output quality at 7.3 tok/s: English prose and reasoning are correct ("The trip | |
| duration is 3.5 hours"), code is correct (`fibonacci(n-1) + fibonacci(n-2)`), | |
| and the model continues fluently in Chinese — Moonshot's other strong language. | |
| French generation stays broken, which is the Width50 build, not the hardware. | |
| **DSpark speculative decoding gives nothing.** Once the draft finally loaded — | |
| after clearing four separate blockers — throughput was 7.32/7.33 against | |
| 7.33/7.35 without it, and llama-server exposed no `draft_n` / `draft_n_accepted` | |
| telemetry to explain why. The predicted floor of 1.2× was still too optimistic; | |
| the measured figure is 1.0×. | |
| ## The answer: host vCPU, not VRAM | |
| The same configuration, on the same model, using the same two GPUs — three | |
| further GPUs sitting idle at 267 MiB — runs at **1.9× the speed** on a bigger | |
| host: | |
| | Pod | GPUs | vCPU | RAM | decode | | |
| |---|---|---|---|---| | |
| | `cumv30dk07dfdy` | 2× A40 | **16** | 93 GB | 2.31 / 2.98 / 2.88 | | |
| | `okinf3ln9k8tsp` | 5× A40 | **41** | 234 GB | 4.86 / 5.28 / 5.44 / 5.64 | | |
| Prefill moves even harder: 0.99 → 8.01 tok/s. | |
| RunPod scales vCPU with GPU count, so renting more GPUs buys host parallelism as | |
| much as it buys VRAM. Every placement experiment below kept 70–90 GB of weights | |
| being processed by 16 saturated threads, which is why moving that work between | |
| CPU and GPU changed nothing: the CPU side was the wall in all of them. | |
| This is the mechanism the section below says was missing. It was found by | |
| accident — a 5-GPU pod picked up the 2-GPU placement script before the guard | |
| was published, and ran the *identical* configuration on a larger host. | |
| ## Why the placement experiments all looked the same | |
| ~3 tok/s is the ceiling for this model on 2× A40, and the reason is not the one | |
| that seems obvious. | |
| Throughput is **invariant** across configurations that differ enormously: | |
| | Configuration | VRAM used | attention | CPU-side bytes/token | decode | | |
| |---|---|---|---|---| | |
| | autofit, 16 experts | 87.7 GB | 48 layers on host | ~33 GB | 2.46 / 3.00 | | |
| | `--cpu-moe` + moe-cache | 60.2 GB | all in VRAM | ~4 GB | 2.21 / 2.68 | | |
| | `--n-cpu-moe 86` | 69.9 GB | all in VRAM | ~4 GB | 2.51 / 3.06 | | |
| | manual balanced `-ot` | 87.8 GB | all in VRAM, 44.3/43.5 split | ~4 GB | 2.31 / 2.98 | | |
| | autofit, 8 experts | 87.4 GB | 48 layers on host | ~17 GB | 3.15 | | |
| | `-ngl 0` (no GPU) | 1.5 GB | all on host | all | **0.65** | | |
| The manual placement is the cleanest of these: explicit `-ot ...=CUDA0/CUDA1` | |
| layer ranges fill both cards to 96%/94% with every layer's attention resident | |
| and only 72 layers' routed experts on the host. It took three calibration | |
| rounds, each guided by the shortfall rather than a guess — 45256 MiB over, then | |
| 1038 MiB over, then fitting. It performs exactly like everything else. | |
| Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%. | |
| **No validated mechanism.** Three hypotheses were proposed and each was refuted | |
| by measurement: | |
| 1. *Routed-expert CPU compute dominates* — refuted: halving experts per token | |
| gave +28%, not the ~100% that would follow. | |
| 2. *Attention re-read from host RAM dominates* — refuted: `--n-cpu-moe 86` puts | |
| every layer's attention in VRAM and cuts host bytes ~8×, and changed nothing. | |
| 3. *CPU↔GPU boundary crossings dominate* — refuted by comparing the two: autofit | |
| splits into contiguous GPU-then-CPU blocks (1–2 crossings per token) while | |
| `--n-cpu-moe 86` crosses at 86 layers (~172 per token). If crossings | |
| dominated, autofit would be far faster. The two are within 2% of each other. | |
| What stands is empirical, not theoretical: **no rebalancing of CPU against GPU | |
| within 92 GB of VRAM improves anything**, across placements spanning 60–88 GB of | |
| VRAM and 4–33 GB of host traffic per token. | |
| **The GPUs are not decorative, though — verified.** Running with `-ngl 0`, so | |
| that essentially nothing sits in VRAM (1265 MiB and 269 MiB), gives **0.65 | |
| tok/s** against 2.51 for the same prompt with autofit. The two A40s are worth | |
| **3.9×**. That check was run specifically because, if the GPUs had contributed | |
| nothing, "buy more VRAM" would have been the wrong advice. | |
| So the direction — fit more of the model in VRAM — is sound, and a ggml | |
| microbenchmark says the magnitude is large. `test-backend-ops` timing | |
| `MUL_MAT_ID` on the K3 expert shapes (MXFP4, 3584→3072) on an L40S: | |
| ``` | |
| 8 experts, C=1 29.25 µs | |
| 16 experts, C=1 59.83 µs | |
| 8 experts, C=7 159.53 µs | |
| 16 experts, C=7 318.97 µs ~7.73 TFLOPS | |
| ``` | |
| At 59.83 µs per MoE tensor and 92 MoE layers, the routed-expert work costs | |
| **5.5 ms/token when the experts are VRAM-resident** — a ~180 tok/s ceiling from | |
| that term alone. We measure 333 ms/token, so GPU MoE compute is about **1.6% of | |
| the time**; the other 98% is the host path. | |
| That makes the naive linear extrapolation (5.4 tok/s at full residency) almost | |
| certainly far too pessimistic: full residency does not scale the existing | |
| regime, it deletes the term that dominates it. | |
| ## What did NOT work, and why | |
| **Halving experts per token (16 → 8).** Bought only +28% (2.46 → 3.15 on the | |
| same prompt) and destroyed the output — `# # # #` after six tokens. REAP had | |
| already cut 896 experts to 448 with the router calibrated for 16 active. The | |
| small gain also *disproved* the hypothesis that routed-expert CPU compute was | |
| the bottleneck: halving it would otherwise have nearly doubled throughput. | |
| **The CUDA MoE expert cache** (`llama-k3-moe-cache.patch`, from the Code145 | |
| build). Applies cleanly to mmnga's branch once `tests/` is excluded, compiles | |
| for sm_86, and `--moe-cache` appears in `--help`. Measured 14% *slower* on | |
| decode and roughly half the prefill. Caveat: that run set none of the seven | |
| `GGML_CUDA_MOE_CACHE_*` variables its own production launcher uses, so the | |
| verdict is against an unconfigured cache, not the design. | |
| **DSpark speculative decoding.** Three separate blockers, each found only after | |
| clearing the previous one: | |
| 1. `--draft-max` was removed; the flag is now `--spec-draft-n-max`. | |
| 2. The published GGUF declares `general.architecture = "dflash-draft"`; | |
| llama.cpp registers `LLM_ARCH_DFLASH` as plain `"dflash"`. Fixed by | |
| `fix_dspark_arch.py` (renames the architecture and its 30 namespaced KV | |
| keys, verified by reopening and comparing every tensor's shape, dtype and | |
| byte size). | |
| 3. The rewritten file still fails: *"DFlash model requires 'target_layers' in | |
| GGUF metadata"*. The third-party converter also names tensors | |
| `dflash.dspark.markov.w1` where llama.cpp expects `markov_w1`. The clean fix | |
| is to convert from `RadixArk/Kimi-K3-DSpark` with llama.cpp's own converter. | |
| Expected ceiling if it ever loads: **1.2–2×**, not the ~3× vLLM reports. On a | |
| MoE each drafted token routes to its own experts, so batch verification | |
| amortises only attention and shared experts, not the routed reads. | |
| **`GGML_OP_OFFLOAD_MIN_BATCH=1`.** ggml can copy only the *selected* experts of | |
| an offloaded `MUL_MAT_ID` to the GPU instead of computing them on the host, but | |
| the path is gated on batch size and the default of 32 means it never fires | |
| during decode. Forcing it on **halved throughput**: 2.46 → 1.13 and 3.00 → 1.38 | |
| tok/s. With 16 experts across 68 offloaded layers that is 1088 separate small | |
| transfers per token, so it is latency-bound, not bandwidth-bound, and loses to | |
| host compute. The upstream default is right. | |
| **`--n-cpu-moe` at any rung (62/68/74/80).** Never loaded. Tensor overrides set | |
| `model_params::tensor_buft_overrides`, which aborts autofit exactly like `-ngl` | |
| and `--tensor-split` do — and once autofit is off, **neither `--split-mode | |
| layer` nor `--tensor-split` distributes anything**: every allocation lands on | |
| device 1 while device 0 stays empty. Rung 80 missed by 135 MiB (46203 requested | |
| against 46068 available) with all attention plus twelve expert layers on one | |
| card. Autofit is the only mechanism that actually uses both GPUs; balancing | |
| manually requires explicit `-ot ...=CUDA0` / `=CUDA1` layer ranges. | |
| **TurboQuant / KV-cache quantisation.** Not applicable: 69 of 93 layers are KDA | |
| recurrent and carry no growing KV, so the whole cache is about 0.8 GB. | |
| **ONNX custom ops** from earlier work target Qwen3.5 on SM75, not Kimi-K3 on | |
| sm_86, and ONNX Runtime has no `kimi_k3` architecture. | |
| ## Output quality | |
| English and code are genuinely usable; French is not. | |
| ``` | |
| "The capital of Japan is" -> Tokyo, Washington, Stockholm, Berlin correct | |
| "train 14:00 -> 17:30" -> "The trip takes 3 hours 30 minutes." correct | |
| "def invert(d):" -> return {v: k for k, v in d.items()} correct | |
| French free-form -> "cunce", "s'fotrenci", "zocneine" non-words | |
| ``` | |
| French *knowledge* survives — "La Revolution francaise a commence en" continues | |
| "1789" — but French *generation* is corrupted at the sub-word level. The prime | |
| suspect is Width50: this build halves `moe_intermediate_size` from 3072 to 1536. | |
| ### Sampler, six configurations on one prompt | |
| | Setting | Result | | |
| |---|---| | |
| | greedy, rp 1.0 | facts correct, then hard repetition loop | | |
| | greedy, rp 1.05 / last_n 64 | no loop, but "Germany is PLN" — the penalty pushes it off correct answers while enumerating | | |
| | greedy, rp 1.15 / last_n 256 | drifts into SQL-flavoured nonsense | | |
| | temp 0.3, min_p 0.05 | fluent, hallucinates ("Tokyo Island", "Sumaru") | | |
| | temp 0.6, top_k 20 | degrades | | |
| | **temp 0.3, min_p 0.1, rp 1.05** | fluent, no loop, nothing glaringly wrong — chosen default | | |
| Repetition penalty measurably *reduces factual accuracy* on this model. Clients | |
| wanting maximum accuracy should still ask for `temperature: 0`. | |
| ## Deployment mechanics | |
| `common_fit_params` sizes GPU placement against actually-free VRAM across all | |
| devices, but **aborts the moment any placement parameter is pinned**: | |
| ``` | |
| -ngl 99 -> "n_gpu_layers already set by user, abort" | |
| --tensor-split -> "model_params::tensor_split already set by user, abort" | |
| ``` | |
| Both aborts produced the same failure: 84790 MiB shoved onto device 1 alone, | |
| cudaMalloc OOM, device 0 untouched. With no placement flags, autofit fills both | |
| cards to 95%. `--cpu-moe` and `--n-cpu-moe` do *not* abort it. | |
| **Boot time** went from 14m28s to ~4 minutes by publishing the compiled | |
| binaries to the Hub (46 MB tarball, 2 s to extract vs 8m34s to compile). The | |
| artifact name pins base image, CUDA version and GPU arch so a mismatched | |
| tarball can never be picked up silently, and it is smoke-tested by running | |
| `llama-server --version` before it is trusted. | |
| **RunPod costs.** A stopped pod's Volume disk bills at $0.20/GB/month, double | |
| the running rate — a 1600 GB volume cost $0.44/h doing nothing, which was 99% | |
| of one day's spend. Container disk is erased on stop but still billed, so for | |
| re-downloadable data the rule is: always Terminate, never Stop. A stopped pod | |
| is also pinned to its host and may refuse to restart while the same GPU type | |
| shows High stock elsewhere. | |
| ## Kimi-K3-256K-REAP-Code145 | |
| Finalised during this session. All 96 shards were already on the Hub (347.5 GB); | |
| only `config.json`, the index and the tokenizer were missing, and | |
| `finalize_code145.py` — which downloads only metadata, never the weights — had | |
| never been run. The repo is now loadable. | |
| ``` | |
| num_experts 145 | |
| num_experts_per_token 16 | |
| moe_intermediate_size 3072 full width, unlike the 1536 we serve | |
| source_index_keys 497 220 | |
| dropped_expert_keys 414 552 | |
| output_index_keys 82 668 = 497220 - 414552 | |
| renamed_expert_keys 80 040 = 92 layers x 145 experts x 6 tensors | |
| ``` | |
| Full width makes it the serious candidate for working French. It does not fit | |
| 2× A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4× A40 at | |
| $1.80/h, or an imatrix-guided quantisation down to ~1.7 bit. | |
| --- | |
| # Session 2 — what actually bounds decode, and what long context costs | |
| ## Decode is serialisation-bound, not bandwidth-bound | |
| Six placements had moved 60–88 GB of weights between CPU and GPU without | |
| changing throughput, and no mechanism had survived measurement. The question | |
| none of them asked was whether the hardware is *idle* between tokens. It is. | |
| Aggregate throughput against concurrent streams, 5× A40, 64 tokens each, | |
| distinct prefixes so no stream rides another's prompt cache: | |
| ``` | |
| 1 stream 5.56 tok/s aggregate 5.56 per stream | |
| 2 streams 9.85 tok/s aggregate 4.93 per stream 1.77x | |
| 4 streams 14.53 tok/s aggregate 3.63 per stream 2.61x | |
| ``` | |
| Concurrency scales, so single-stream decode leaves silicon idle. The | |
| bandwidth arithmetic agrees independently: ~7 GB of weights are touched per | |
| token, and at 138 ms/token that is **~51 GB/s against 696 GB/s available** on | |
| the active card — 7% of the roof. llama.cpp splits layers across GPUs, so at | |
| batch 1 exactly one card computes while the other four wait. | |
| This retires the last open question from session 1 and explains why every | |
| kernel-level optimisation attempt returned nothing: the expert GEMM was never | |
| the constraint. | |
| ## Sustained decode on real output: 7.22 tok/s | |
| The 7.3 tok/s headline was measured on selftests, three of whose five prompts | |
| produce degenerate repetition (`# [CS_ah]`, `et d'et d'`) — and a repetition | |
| loop decodes cheap. Re-measured on 600 tokens of coherent prose: **7.22 tok/s**. | |
| The number holds; it is now measured on output worth having. | |
| ## Prefill: 167 tok/s cold, ~free warm | |
| Every selftest prompt is 6–22 tokens, so its "prefill tok/s" is per-request | |
| overhead and says nothing. Measured properly, 3408-token prompt: | |
| ``` | |
| cold 3408 tok in 20.4 s = 167 tok/s (23x the decode rate — prefill batches) | |
| warm 4 tok in 0.53 s (slot prompt cache reprocessed only the delta) | |
| ``` | |
| For a Claude Code client sending a 32k system prompt: **~3 min once**, then | |
| only the changed tokens on every later turn. | |
| ## Long context is nearly free on this architecture | |
| From the config, not from assumption: `kv_lora_rank 512`, `qk_rope_head_dim | |
| 64`, `full_attn_layers` every 4th layer, `max_position_embeddings 1048576`. | |
| ``` | |
| 24 of 93 layers cache at all (MLA); the other 69 are KDA — recurrent state, | |
| constant regardless of context length. | |
| KV per token = 24 x (512 + 64) x 2 = 27.6 KB | |
| 65536 ctx -> 1.8 GB 262144 -> 7.2 GB 1048576 -> 29 GB | |
| ``` | |
| Against ~38 GB left free by the weights on a 5-GPU host. **16384 was never a | |
| hardware limit, just an untested default** — and it could not accept a single | |
| Claude Code request. What constrains context here is the weights sharing the | |
| cards, never the cache. | |
| ## Cost reality check | |
| Runpod hosts an official Kimi-K3 endpoint (`moonshot-kimi`). Measured: | |
| 300 tokens in 10.281 s, billed $0.004797. | |
| ``` | |
| speed $/1M output quality | |
| official 29.2 tok/s ~16 full precision, French works | |
| ours (1 stream) 7.2 tok/s ~85 1.14 bpw, French broken | |
| ours (4 streams) 14.5 tok/s ~42 | |
| ``` | |
| The official endpoint is 4x faster and 5.7x cheaper. Self-hosting buys | |
| privacy, no per-token billing, and context control — not price or quality. | |
| Worth restating whenever the effort seems to justify itself on cost. | |
| ## No smaller GGUF exists | |
| Ours is the most compact Kimi-K3 published, because it is the Width50 variant | |
| (`moe_intermediate_size` 1536 rather than 3072) — which is also why French is | |
| broken. | |
| ``` | |
| ours REAP448 Width50 IQ1_S 181 GiB | |
| 0xTank REAP568 UD-IQ1_S 246 GiB | |
| mmnga-o REAP50 UD-IQ1_S 305 GiB | |
| prometheusAIR REAP55 IQ1_M 319 GiB | |
| hellohazime REAP640ja 411 GiB | |
| ``` | |
| A "light" build for a single A40 (46 GB, ~37 GB of weights after a 256k KV) | |
| does not exist and would have to be produced — roughly a 64-expert REAP. | |
| ## Under $1 per million tokens is not reachable with K3 | |
| At $0.44/h per A40, $1/1M output requires **122 tok/s aggregate per card**. The | |
| 5× pod at $2.20/h would need 611 tok/s; it delivers 14.5. That is a 42x gap, | |
| not a tuning problem. | |
| What does reach it, from HF configs (throughput figures are bandwidth-derived | |
| estimates, not measured): | |
| ``` | |
| Qwen3-Coder-30B-A3B 30B/3B active ~30 GB FP8 KV 12 GiB@256k 1x A40 ~$0.10-0.25/1M | |
| Qwen3-Coder-Next 80B/3B active ~74 GB FP8 KV 12 GiB@256k 2x A40 ~$0.20-0.80/1M | |
| Qwen3.8-27B dense ~26 GB FP8 KV 32 GiB@256k 256k does not fit on 1 card | |
| ``` | |
| Served by **vLLM, not llama.cpp** — tensor parallelism and continuous batching | |
| are exactly what the concurrency measurement above shows llama.cpp lacking. | |
| ## Operational trap: bare `wait` kills the hot-reload watcher | |
| `llama-server` is started with `&`, so a bare `wait` anywhere later in the | |
| script waits for *it* too, forever. The server keeps serving, the selftest | |
| never finishes, the watcher never starts, and the pod silently ignores every | |
| published update. Symptom: `/props` reports the old config long after a push. | |
| Always `wait "$pid"` with an explicit pid. | |
| --- | |
| # Session 3 — 256k validated, and a measured alternative that dominates it | |
| ## Kimi-K3 at 262144 context: it works, and it costs nothing in decode | |
| `--ctx-size 262144`, all-VRAM on 5x A40, 39.5-42.5 GB used per card: | |
| ``` | |
| n_ctx 262144 (was 16384 -- never a hardware limit, just a default) | |
| decode 7.25 tok/s (unchanged from 16k: 7.2-7.3) | |
| ``` | |
| 16x the context for no throughput cost, exactly as the KDA/MLA geometry | |
| predicted. The DSpark rungs OOM'd at this context and the ladder correctly fell | |
| through to all-VRAM without a draft — no loss, DSpark was measured at zero gain. | |
| ## Prefill degrades gracefully, and the proxy times out before it finishes | |
| A cold 38k prefill takes ~230 s. Runpod's HTTP proxy sits behind Cloudflare, | |
| which cuts the connection at ~125 s (`error code: 524`). The workaround is to | |
| build the prefix in chunks — each request extends the cached prefix — which | |
| also measures the degradation curve a single call would have hidden: | |
| ``` | |
| chunk 1 +8400 tok @ 8k ctx 228.6 tok/s | |
| chunk 2 +8404 tok @ 17k ctx 203.6 tok/s | |
| chunk 3 +8404 tok @ 25k ctx 182.4 tok/s | |
| chunk 4 +8605 tok @ 34k ctx 164.1 tok/s | |
| chunk 5 +8704 tok @ 43k ctx 149.8 tok/s | |
| ``` | |
| -34% across 5x the context, far better than quadratic — 69 of 93 layers are | |
| KDA (linear), only 24 are MLA. 42,517 tokens prefilled in 3 min 55 total. | |
| Re-submitting a cached 32k prompt returns `prompt_n=4` in 0.6 s. | |
| ## vLLM's Kimi-K3 support is real and unreachable | |
| vLLM and SGLang both ship day-0 K3 support, including the two things that | |
| failed here: a hybrid prefix cache over recurrent KDA state, and working DSpark | |
| block-diffusion speculative decoding. None of it is usable on A40s, because | |
| every vLLM-loadable K3 checkpoint is the full unpruned model: | |
| ``` | |
| nvidia/Kimi-K3-NVFP4 1499 GiB + requires Blackwell (sm_100) | |
| RedHatAI/Kimi-K3-NVFP4 1533 GiB + requires Blackwell | |
| RedHatAI/Kimi-K3-FP8-BLOCK 2626 GiB + requires sm_89+ | |
| ``` | |
| No AWQ, no GPTQ, no W4A16. The 181 GiB IQ1_S we run is llama.cpp-only, and | |
| llama.cpp is exactly what caps us. **The unlock would be quantising a REAP'd K3 | |
| to W4A16 for vLLM** — Code145 at 4-bit lands near 260 GB (6x A40 / 4x A100-80). | |
| That build, not any config change, is the only path to the good engine. | |
| ## The measured alternative | |
| Qwen3-Coder-30B-A3B-Instruct, AWQ 4-bit (16.9 GiB), vLLM 0.27.1, **one** A40 at | |
| $0.44/h, 262144 context with fp8 KV. Everything below is measured, not derived: | |
| ``` | |
| concurrency tokens duration aggregate per stream $/1M output | |
| 1 128 1.6 s 77.6 tok/s 77.6 $1.575 | |
| 4 512 1.9 s 265.3 tok/s 66.3 $0.461 | |
| 8 1024 2.1 s 496.8 tok/s 62.1 $0.246 | |
| 16 2048 2.6 s 782.4 tok/s 48.9 $0.156 | |
| 32 4096 3.5 s 1165.6 tok/s 36.4 $0.105 | |
| prefill 38,494 tok cold in 10.42 s = 3693 tok/s; warm 1.07 s | |
| ``` | |
| Head to head: | |
| ``` | |
| Kimi-K3 5x A40 Qwen3-Coder 1x A40 K3 official API | |
| hourly $2.20 $0.44 per-token | |
| decode, 1 stream 7.2 tok/s 77.6 tok/s 29.2 tok/s | |
| prefill @38k ~165 tok/s 3693 tok/s -- | |
| 32k prompt, cold 3 min 55 10 s -- | |
| context 262144 5 cards 1 card -- | |
| $/1M output ~85 $1.58 .. $0.105 ~16 | |
| French broken correct correct | |
| ``` | |
| Qwen is 10.8x on decode, 22x on prefill, on one fifth of the hardware, and it | |
| answers in French. The under-$1/1M target is met from two concurrent streams | |
| and reaches $0.105 at 32. Kimi-K3 self-hosted on A40s is not competitive on any | |
| axis measured here; it remains interesting only where the specific K3 model is | |
| the requirement. | |
| ## Two mistakes worth not repeating | |
| - Creating a pod from `create-pod` without a start command: the schema has no | |
| `args` field, so the container starts with the image default. Use | |
| `create-template` with `dockerStartCmd`, then deploy with `templateId`. | |
| - `--disable-log-requests` was removed in vLLM 0.27; it crash-loops the server. | |
| A pod's `args` cannot be edited after creation — only the template can — so a | |
| bad flag costs a full recreate. | |
| ## The context ceiling on 5x A40, and why q8 KV is not the way out | |
| ``` | |
| ctx 16384 7.2-7.3 tok/s | |
| ctx 262144 7.25 tok/s | |
| ctx 524288 7.29 tok/s <- ceiling, all-VRAM, no draft | |
| ctx 1048576 OOM: 29 GB of f16 MLA cache does not fit alongside the weights | |
| KV q8_0 rejected outright, at every context | |
| ``` | |
| Decode is flat across a 32x range of context. That is the KDA/MLA hybrid doing | |
| exactly what it is for: only 24 of 93 layers cache at all, and the other 69 hold | |
| a recurrent state whose size is independent of context. | |
| **q8_0 KV cannot be used here**, and the failure signature says why. Every | |
| placement rung failed in 1.6-2.5 s with `failed to create llama_context from | |
| model` and *no* `cudaMalloc` line — a parameter rejection. Contrast the genuine | |
| OOM at 512k on the draft rung: 29.9 s, with `cudaMalloc failed: out of memory` | |
| printed. Fast failure with no allocator message means llama.cpp refused the | |
| configuration; slow failure with one means it tried and ran out. The 69 KDA | |
| layers carry a recurrent state llama.cpp will not quantise, so f16 is forced and | |
| the ceiling is set by memory: 512k fits, 1M does not. | |
| TurboQuant (Google, ICLR 2026) would take the cache to ~3 bits and put 1M | |
| within reach, but as of this run it is a paper with an open SGLang feature | |
| request (#21618) and no merged implementation in any engine. | |
| ## Qwen decode against context length | |
| ``` | |
| short context 77.6 tok/s | |
| 119k context 34.3 tok/s (prompt served warm from prefix cache) | |
| ``` | |
| Halves across 119k, and still 4.7x Kimi's short-context rate. Tool calling was | |
| verified on the OpenAI endpoint: two parallel calls, correct arguments, | |
| `finish_reason: tool_calls`. | |
| ## Runpod's HTTP proxy caps a single request at ~125 s | |
| Cloudflare sits in front of the pod proxy and returns `error code: 524`. A cold | |
| prefill longer than that cannot complete in one call — measured at ~40k on | |
| Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached | |
| prefix) or expose a TCP port instead. | |
| --- | |
| # Session 4 — speculative decoding, and why the regime decides everything | |
| n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40, | |
| `--speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'`. | |
| Same model, same hardware, same quant, two pods running side by side so the | |
| comparison is simultaneous rather than sequential: | |
| ``` | |
| A. pure generation (no prompt/output overlap) | |
| baseline 102.21 tok/s | |
| ngram 58.49 tok/s -43% | |
| B. code editing (output re-emits the input) | |
| baseline 130.16 tok/s | |
| ngram 242.39 tok/s +86% | |
| ``` | |
| **Speculation is not a free win — it is a bet on repetition.** With nothing to | |
| guess, every draft is rejected and the verification cost is pure loss; that is | |
| the -43%. Claude Code lives almost entirely in regime B (read a file, emit a | |
| modified version, repeat identifiers), so it is the right default *there* and | |
| the wrong default for prose. | |
| vLLM's telemetry, aggregated over both regimes: | |
| ``` | |
| drafts 207 draft tokens 1035 accepted 706 = 68.2% | |
| accepted per position: 179 / 156 / 131 / 120 / 120 | |
| mean accepted length 3.41 -> ~4.4 tokens emitted per forward pass | |
| ``` | |
| This is the measurement llama.cpp never produced: DSpark there exposed no | |
| `draft_n` at all and moved throughput by 0.0%. The mechanism was never broken — | |
| the engine was. | |
| ## The number that matters for the original goal | |
| Single stream, code editing, one A40 at $0.44/h: | |
| ``` | |
| 242.39 tok/s -> $0.50 per 1M output tokens | |
| ``` | |
| The under-$1/1M target is met **on a single stream**, without needing | |
| concurrency to amortise anything. For reference the same target on Kimi-K3 | |
| required 611 tok/s aggregate against 14.5 measured. | |
| Note also that the baseline itself reads higher here (102-130 tok/s) than the | |
| 77.6 measured earlier: that earlier figure was a 128-token request whose wall | |
| clock was dominated by per-request overhead. Longer outputs amortise it. Quote | |
| 77.6 for short replies and ~130 for sustained generation. | |
| --- | |
| # Session 5 — Kimi-Linear-48B-A3B: genuine Kimi, 1M context, one A40 | |
| `cyankiwi/Kimi-Linear-48B-A3B-Instruct-AWQ-4bit`, 28.4 GiB, vLLM 0.27.1, one | |
| A40 at $0.44/h. `KimiLinearForCausalLM`, `model_type: kimi_linear` — the same | |
| architecture family as K3, with the same MLA geometry (`kv_lora_rank 512`, | |
| `qk_rope_head_dim 64`). | |
| ## Why it fits where K3 could not | |
| ``` | |
| 27 layers = 7 MLA (full_attn_layers [4,8,12,16,20,24,27]) + 20 KDA | |
| KV/token = 7 x (512 + 64) x 2 = 7.88 KB (K3: 24 layers -> 27.6 KB) | |
| weights AWQ4 28.4 GiB | |
| 1,048,576 ctx 7.88 GiB KV => 36.3 GB on a 46 GB card | |
| ``` | |
| vLLM confirmed it at boot: **`GPU KV cache size: 1,416,566 tokens`** on a single | |
| card. K3 needed five A40s at $2.20/h to reach 524,288. | |
| ## Measured | |
| ``` | |
| flux aggregate per stream $/1M | |
| 1 67.4 67.4 1.814 | |
| 4 206.6 51.6 0.592 | |
| 8 352.4 44.0 0.347 | |
| 16 556.1 34.8 0.220 | |
| 32 779.9 24.4 0.157 | |
| prefill cold 6162 tok/s (better than Qwen3-Coder-30B's 3957) | |
| prefill warm 6430 tok/s ratio 1.04 -> NO prefix caching | |
| code editing 89.4 tok/s | |
| tool calling 2 parallel calls, correct arguments | |
| code correct; French fluent | |
| ``` | |
| ## n-gram speculation is BROKEN on this backend | |
| Rejection sampling is supposed to be lossless, so any output difference between | |
| speculative and non-speculative decoding is a bug. Here it corrupts code, | |
| reproducibly and identically across runs: | |
| ``` | |
| with speculation without | |
| add() return a -b):\n return a - return a - b\n\ndef multiply(a, | |
| fibonacci fibonacci(n-1(n-1) + fibonacci(n-2) fibonacci(n-1) + fibonacci(n-2) | |
| ``` | |
| The signature is a just-emitted fragment re-spliced at the wrong place — a bad | |
| draft accept. Plain prose and verbatim copying stayed perfect, which is why it | |
| took targeted probes to find: code is dense in the short repeats (parentheses, | |
| operators) that n-gram matching latches onto. | |
| It is also 3-4x *slower*, because vLLM disables CUDA graphs under speculation | |
| on `TritonMLABackend`: | |
| ``` | |
| flux with spec without ratio | |
| 1 23.1 71.2 3.1x | |
| 4 58.7 249.6 4.3x | |
| 32 203.9 781.0 3.8x | |
| $/1M @32 0.600 0.157 | |
| ``` | |
| **Turn speculation off for Kimi-Linear.** Worth reporting upstream. | |
| ## Two fixes that made it work at all | |
| - **Tokenizer.** `tokenization_kimi.py` imports `bytes_to_unicode` from | |
| `transformers.convert_slow_tokenizer`, removed in transformers >= 5.5.3 which | |
| every vLLM >= 0.24 requires. No Kimi-Linear repo ships a fast `tokenizer.json`, | |
| so the Python tokenizer is mandatory. Fixed by publishing | |
| `patdev/kimi-linear-tokenizer-fix`: same files with a local fallback | |
| definition, verified byte-identical to the transformers reference before | |
| deployment. Point vLLM at it with `--tokenizer`. | |
| - **Tool parser.** `hermes` yields zero calls. The chat template uses | |
| `<|tool_calls_section_begin|>` / `<|tool_call_begin|>`, the K2 token scheme — | |
| so the parser is `kimi_k2`. (`kimi_k3` is for `<|open|>`/`<|close|>`/`<|sep|>` | |
| and does not apply.) It also sets `skip_special_tokens=False`, which those | |
| markers need. | |
| - **Headroom.** `--gpu-memory-utilization 0.93` OOMs at the very last step, in | |
| the speculative rejection sampler asking for 160 MiB with 143 MiB left. 0.90 | |
| leaves room and still yields >1M tokens of KV. | |
| ## Head to head | |
| ``` | |
| Kimi-Linear 48B Qwen3-Coder-30B | |
| hourly $0.44 $0.44 | |
| decode 1 stream 67-71 tok/s 77.6-79.2 tok/s | |
| aggregate @32 780 1166 | |
| prefill cold 6162 3957 | |
| prefill warm 6430 63315 | |
| context 1,048,576 262,144 | |
| $/1M @32 0.157 0.105 | |
| tool calling yes yes | |
| French fluent correct | |
| brand GENUINE KIMI Qwen | |
| ``` | |
| Kimi-Linear wins on context (4x) and cold prefill (1.6x). It loses on decode | |
| and, decisively for an agent client, on prefix caching: **6430 vs 63315, ten | |
| times worse**, because vLLM does not yet support prefix caching over recurrent | |
| hybrid state. A 32k system prompt costs ~5 s every turn instead of ~0.5 s. | |
| ## Qwen3-Coder-Next 80B: no data | |
| Two attempts, both infrastructure failures, zero requests served. | |
| ``` | |
| attempt 1 hung 22 min in NCCL peer-to-peer init on 2x A40 | |
| attempt 2 NCCL_P2P_DISABLE=1 cleared that, then looped ~10 min on | |
| "No available shared memory broadcast block found in 60 seconds" | |
| ``` | |
| `NCCL_P2P_DISABLE=1` is required for tensor parallelism on Runpod A40 pairs. | |
| The second hang is unproven — the hypothesis is torch.compile on a 512-expert | |
| MoE, testable with `--enforce-eager`. Its boot log did confirm one thing: | |
| `enable_prefix_caching=False`, the same hybrid limitation as Kimi-Linear. | |