patdev commited on
Commit
28eb400
Β·
verified Β·
1 Parent(s): ce3e5b7

Upload FINDINGS.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. FINDINGS.md +13 -3
FINDINGS.md CHANGED
@@ -48,9 +48,19 @@ by measurement:
48
 
49
  What stands is empirical, not theoretical: **no rebalancing of CPU against GPU
50
  within 92 GB of VRAM improves anything**, across placements spanning 60–88 GB of
51
- VRAM and 4–33 GB of host traffic per token. The practical conclusion holds β€”
52
- 6–8 tok/s needs more VRAM (4Γ— A40, ~$1.80/h), not more tuning β€” but the reason
53
- it holds is not established, and should not be presented as if it were.
 
 
 
 
 
 
 
 
 
 
54
 
55
  ## What did NOT work, and why
56
 
 
48
 
49
  What stands is empirical, not theoretical: **no rebalancing of CPU against GPU
50
  within 92 GB of VRAM improves anything**, across placements spanning 60–88 GB of
51
+ VRAM and 4–33 GB of host traffic per token.
52
+
53
+ **The GPUs are not decorative, though β€” verified.** Running with `-ngl 0`, so
54
+ that essentially nothing sits in VRAM (1265 MiB and 269 MiB), gives **0.65
55
+ tok/s** against 2.51 for the same prompt with autofit. The two A40s are worth
56
+ **3.9Γ—**. That check was run specifically because, if the GPUs had contributed
57
+ nothing, "buy more VRAM" would have been the wrong advice.
58
+
59
+ So the direction β€” fit more of the model in VRAM β€” is sound. The magnitude is
60
+ not predictable from here: a naive linear fit through (0 GB β†’ 0.65) and
61
+ (92 GB β†’ 2.94) lands near 5.4 tok/s at full residency, but removing the host
62
+ path entirely should be super-linear, since it swaps a mixed regime for pure
63
+ VRAM at 696 GB/s. The honest answer is "substantially better, number unknown".
64
 
65
  ## What did NOT work, and why
66