patdev commited on
Commit
c57032f
Β·
verified Β·
1 Parent(s): 4614aa1

Upload FINDINGS.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. FINDINGS.md +127 -0
FINDINGS.md CHANGED
@@ -261,3 +261,130 @@ renamed_expert_keys 80 040 = 92 layers x 145 experts x 6 tensors
261
  Full width makes it the serious candidate for working French. It does not fit
262
  2Γ— A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4Γ— A40 at
263
  $1.80/h, or an imatrix-guided quantisation down to ~1.7 bit.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
261
  Full width makes it the serious candidate for working French. It does not fit
262
  2Γ— A40 at MXFP4 (347.5 GB against 185 GB of fast memory); either 4Γ— A40 at
263
  $1.80/h, or an imatrix-guided quantisation down to ~1.7 bit.
264
+
265
+ ---
266
+
267
+ # Session 2 β€” what actually bounds decode, and what long context costs
268
+
269
+ ## Decode is serialisation-bound, not bandwidth-bound
270
+
271
+ Six placements had moved 60–88 GB of weights between CPU and GPU without
272
+ changing throughput, and no mechanism had survived measurement. The question
273
+ none of them asked was whether the hardware is *idle* between tokens. It is.
274
+
275
+ Aggregate throughput against concurrent streams, 5Γ— A40, 64 tokens each,
276
+ distinct prefixes so no stream rides another's prompt cache:
277
+
278
+ ```
279
+ 1 stream 5.56 tok/s aggregate 5.56 per stream
280
+ 2 streams 9.85 tok/s aggregate 4.93 per stream 1.77x
281
+ 4 streams 14.53 tok/s aggregate 3.63 per stream 2.61x
282
+ ```
283
+
284
+ Concurrency scales, so single-stream decode leaves silicon idle. The
285
+ bandwidth arithmetic agrees independently: ~7 GB of weights are touched per
286
+ token, and at 138 ms/token that is **~51 GB/s against 696 GB/s available** on
287
+ the active card β€” 7% of the roof. llama.cpp splits layers across GPUs, so at
288
+ batch 1 exactly one card computes while the other four wait.
289
+
290
+ This retires the last open question from session 1 and explains why every
291
+ kernel-level optimisation attempt returned nothing: the expert GEMM was never
292
+ the constraint.
293
+
294
+ ## Sustained decode on real output: 7.22 tok/s
295
+
296
+ The 7.3 tok/s headline was measured on selftests, three of whose five prompts
297
+ produce degenerate repetition (`# [CS_ah]`, `et d'et d'`) β€” and a repetition
298
+ loop decodes cheap. Re-measured on 600 tokens of coherent prose: **7.22 tok/s**.
299
+ The number holds; it is now measured on output worth having.
300
+
301
+ ## Prefill: 167 tok/s cold, ~free warm
302
+
303
+ Every selftest prompt is 6–22 tokens, so its "prefill tok/s" is per-request
304
+ overhead and says nothing. Measured properly, 3408-token prompt:
305
+
306
+ ```
307
+ cold 3408 tok in 20.4 s = 167 tok/s (23x the decode rate β€” prefill batches)
308
+ warm 4 tok in 0.53 s (slot prompt cache reprocessed only the delta)
309
+ ```
310
+
311
+ For a Claude Code client sending a 32k system prompt: **~3 min once**, then
312
+ only the changed tokens on every later turn.
313
+
314
+ ## Long context is nearly free on this architecture
315
+
316
+ From the config, not from assumption: `kv_lora_rank 512`, `qk_rope_head_dim
317
+ 64`, `full_attn_layers` every 4th layer, `max_position_embeddings 1048576`.
318
+
319
+ ```
320
+ 24 of 93 layers cache at all (MLA); the other 69 are KDA β€” recurrent state,
321
+ constant regardless of context length.
322
+
323
+ KV per token = 24 x (512 + 64) x 2 = 27.6 KB
324
+
325
+ 65536 ctx -> 1.8 GB 262144 -> 7.2 GB 1048576 -> 29 GB
326
+ ```
327
+
328
+ Against ~38 GB left free by the weights on a 5-GPU host. **16384 was never a
329
+ hardware limit, just an untested default** β€” and it could not accept a single
330
+ Claude Code request. What constrains context here is the weights sharing the
331
+ cards, never the cache.
332
+
333
+ ## Cost reality check
334
+
335
+ Runpod hosts an official Kimi-K3 endpoint (`moonshot-kimi`). Measured:
336
+ 300 tokens in 10.281 s, billed $0.004797.
337
+
338
+ ```
339
+ speed $/1M output quality
340
+ official 29.2 tok/s ~16 full precision, French works
341
+ ours (1 stream) 7.2 tok/s ~85 1.14 bpw, French broken
342
+ ours (4 streams) 14.5 tok/s ~42
343
+ ```
344
+
345
+ The official endpoint is 4x faster and 5.7x cheaper. Self-hosting buys
346
+ privacy, no per-token billing, and context control β€” not price or quality.
347
+ Worth restating whenever the effort seems to justify itself on cost.
348
+
349
+ ## No smaller GGUF exists
350
+
351
+ Ours is the most compact Kimi-K3 published, because it is the Width50 variant
352
+ (`moe_intermediate_size` 1536 rather than 3072) β€” which is also why French is
353
+ broken.
354
+
355
+ ```
356
+ ours REAP448 Width50 IQ1_S 181 GiB
357
+ 0xTank REAP568 UD-IQ1_S 246 GiB
358
+ mmnga-o REAP50 UD-IQ1_S 305 GiB
359
+ prometheusAIR REAP55 IQ1_M 319 GiB
360
+ hellohazime REAP640ja 411 GiB
361
+ ```
362
+
363
+ A "light" build for a single A40 (46 GB, ~37 GB of weights after a 256k KV)
364
+ does not exist and would have to be produced β€” roughly a 64-expert REAP.
365
+
366
+ ## Under $1 per million tokens is not reachable with K3
367
+
368
+ At $0.44/h per A40, $1/1M output requires **122 tok/s aggregate per card**. The
369
+ 5Γ— pod at $2.20/h would need 611 tok/s; it delivers 14.5. That is a 42x gap,
370
+ not a tuning problem.
371
+
372
+ What does reach it, from HF configs (throughput figures are bandwidth-derived
373
+ estimates, not measured):
374
+
375
+ ```
376
+ Qwen3-Coder-30B-A3B 30B/3B active ~30 GB FP8 KV 12 GiB@256k 1x A40 ~$0.10-0.25/1M
377
+ Qwen3-Coder-Next 80B/3B active ~74 GB FP8 KV 12 GiB@256k 2x A40 ~$0.20-0.80/1M
378
+ Qwen3.8-27B dense ~26 GB FP8 KV 32 GiB@256k 256k does not fit on 1 card
379
+ ```
380
+
381
+ Served by **vLLM, not llama.cpp** β€” tensor parallelism and continuous batching
382
+ are exactly what the concurrency measurement above shows llama.cpp lacking.
383
+
384
+ ## Operational trap: bare `wait` kills the hot-reload watcher
385
+
386
+ `llama-server` is started with `&`, so a bare `wait` anywhere later in the
387
+ script waits for *it* too, forever. The server keeps serving, the selftest
388
+ never finishes, the watcher never starts, and the pod silently ignores every
389
+ published update. Symptom: `/props` reports the old config long after a push.
390
+ Always `wait "$pid"` with an explicit pid.