patdev commited on
Commit
a560fd0
·
verified ·
1 Parent(s): 5388b01

Upload FINDINGS.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. FINDINGS.md +14 -6
FINDINGS.md CHANGED
@@ -25,12 +25,20 @@ that seems obvious.
25
 
26
  Throughput is **invariant** across configurations that differ enormously:
27
 
28
- | Configuration | VRAM used | CPU-side bytes/token | decode |
29
- |---|---|---|---|
30
- | autofit, 16 experts | 87.7 GB | ~33 GB | 2.46 / 3.00 |
31
- | `--cpu-moe` + moe-cache | 60.2 GB | ~4 GB | 2.21 / 2.68 |
32
- | `--n-cpu-moe 86` | 69.9 GB | ~4 GB | 2.51 / 3.06 |
33
- | autofit, 8 experts | 87.4 GB | ~17 GB | 3.15 |
 
 
 
 
 
 
 
 
34
 
35
  Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%.
36
 
 
25
 
26
  Throughput is **invariant** across configurations that differ enormously:
27
 
28
+ | Configuration | VRAM used | attention | CPU-side bytes/token | decode |
29
+ |---|---|---|---|---|
30
+ | autofit, 16 experts | 87.7 GB | 48 layers on host | ~33 GB | 2.46 / 3.00 |
31
+ | `--cpu-moe` + moe-cache | 60.2 GB | all in VRAM | ~4 GB | 2.21 / 2.68 |
32
+ | `--n-cpu-moe 86` | 69.9 GB | all in VRAM | ~4 GB | 2.51 / 3.06 |
33
+ | manual balanced `-ot` | 87.8 GB | all in VRAM, 44.3/43.5 split | ~4 GB | 2.31 / 2.98 |
34
+ | autofit, 8 experts | 87.4 GB | 48 layers on host | ~17 GB | 3.15 |
35
+ | `-ngl 0` (no GPU) | 1.5 GB | all on host | all | **0.65** |
36
+
37
+ The manual placement is the cleanest of these: explicit `-ot ...=CUDA0/CUDA1`
38
+ layer ranges fill both cards to 96%/94% with every layer's attention resident
39
+ and only 72 layers' routed experts on the host. It took three calibration
40
+ rounds, each guided by the shortfall rather than a guess — 45256 MiB over, then
41
+ 1038 MiB over, then fitting. It performs exactly like everything else.
42
 
43
  Cutting host traffic by 8× moved nothing. Halving expert compute moved +28%.
44