patdev commited on
Commit
78b5ac3
·
verified ·
1 Parent(s): 51cc4b7

Upload FINDINGS.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. FINDINGS.md +54 -0
FINDINGS.md CHANGED
@@ -531,3 +531,57 @@ Cloudflare sits in front of the pod proxy and returns `error code: 524`. A cold
531
  prefill longer than that cannot complete in one call — measured at ~40k on
532
  Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached
533
  prefix) or expose a TCP port instead.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
531
  prefill longer than that cannot complete in one call — measured at ~40k on
532
  Kimi and ~232k on Qwen. Chunk the prompt (each request extends the cached
533
  prefix) or expose a TCP port instead.
534
+
535
+ ---
536
+
537
+ # Session 4 — speculative decoding, and why the regime decides everything
538
+
539
+ n-gram speculative decoding on Qwen3-Coder-30B-A3B-AWQ, vLLM 0.27.1, 1x A40,
540
+ `--speculative-config '{"method":"ngram","num_speculative_tokens":5,...}'`.
541
+ Same model, same hardware, same quant, two pods running side by side so the
542
+ comparison is simultaneous rather than sequential:
543
+
544
+ ```
545
+ A. pure generation (no prompt/output overlap)
546
+ baseline 102.21 tok/s
547
+ ngram 58.49 tok/s -43%
548
+
549
+ B. code editing (output re-emits the input)
550
+ baseline 130.16 tok/s
551
+ ngram 242.39 tok/s +86%
552
+ ```
553
+
554
+ **Speculation is not a free win — it is a bet on repetition.** With nothing to
555
+ guess, every draft is rejected and the verification cost is pure loss; that is
556
+ the -43%. Claude Code lives almost entirely in regime B (read a file, emit a
557
+ modified version, repeat identifiers), so it is the right default *there* and
558
+ the wrong default for prose.
559
+
560
+ vLLM's telemetry, aggregated over both regimes:
561
+
562
+ ```
563
+ drafts 207 draft tokens 1035 accepted 706 = 68.2%
564
+ accepted per position: 179 / 156 / 131 / 120 / 120
565
+ mean accepted length 3.41 -> ~4.4 tokens emitted per forward pass
566
+ ```
567
+
568
+ This is the measurement llama.cpp never produced: DSpark there exposed no
569
+ `draft_n` at all and moved throughput by 0.0%. The mechanism was never broken —
570
+ the engine was.
571
+
572
+ ## The number that matters for the original goal
573
+
574
+ Single stream, code editing, one A40 at $0.44/h:
575
+
576
+ ```
577
+ 242.39 tok/s -> $0.50 per 1M output tokens
578
+ ```
579
+
580
+ The under-$1/1M target is met **on a single stream**, without needing
581
+ concurrency to amortise anything. For reference the same target on Kimi-K3
582
+ required 611 tok/s aggregate against 14.5 measured.
583
+
584
+ Note also that the baseline itself reads higher here (102-130 tok/s) than the
585
+ 77.6 measured earlier: that earlier figure was a 128-token request whose wall
586
+ clock was dominated by per-request overhead. Longer outputs amortise it. Quote
587
+ 77.6 for short replies and ~130 for sustained generation.