Winery-Strata / docs /SECOND_GPU.md
Zandy-Wandy's picture
🍷 Winery Strata: fork of Niko1221/Strata (Winery Cuvée family, wine-red UI)
bbb6388 verified
|
Raw History Blame Contribute Delete
7.15 kB
# CUDA1–3 expert experiment
CUDA0 keeps the dense layers, KV/state, MTP and its existing expert cache. Up to
three independent expert caches fill CUDA1, CUDA2 and CUDA3 in that order. Each
expert has one owner: CUDA0, one secondary GPU, or the CPU pool. Each secondary
GPU computes its rows while the CPU pool works; the results return through
private pinned host buffers before CUDA0 continues the layer. The GPU launches
can overlap, but each layer waits for all three.
The shipped profile ranks 8,000 `(layer, expert)` pairs. CUDA0 takes its usual
`--expert-cache auto` allocation. With only CUDA1, it takes the first remaining
pairs. When CUDA2 or CUDA3 is enabled, the remaining profiled pairs are assigned
round robin to the secondary GPUs. The ranked list is then extended with all
other pairs in expert-then-layer order and distributed in the same way. This
fills the additional VRAM and lets all secondary GPUs see frequent experts, but
unranked experts may hardly ever be routed. The startup
log reports slots and GiB per card; each request reports the number of expert
entries actually computed on each secondary GPU. Remote outputs are packed
before returning over PCIe/USB4: the request log compares MiB transferred
with the previous full-row transfer. The input and routing metadata use pinned
host staging. This reduces transfer traffic and launch overhead, but does not
release the expert arena in system RAM or significantly change VRAM use.
The optional `--gpu-placement layer` gives each layer to exactly one secondary
GPU (layer number modulo the number of enabled cards). It fills that card with
the most frequently routed experts of its layers first, then the unranked ones.
This can reduce USB4 device switches and synchronization when several secondary
cards serve the same layer in the default `stripe` mode. All three cards still
keep their requested expert slots; the primary CUDA0 cache and CPU fallback
continue to work. The log counts active layer launches on each card, so the
tradeoff can be measured alongside tokens per second. This placement is
experimental and may be slower if a given layer needs more than one card's
compute capacity.
The experiment requires the corresponding visible CUDA devices and a CUDA
enabled build. Each card must keep at least 512 MiB free. Tiers must be enabled
in order. Under WDDM (Windows, WSL2), CUDA registration of the host expert arena is capped at
8 GiB to leave room for the contexts and MTP on CUDA0 (`STRATA_ARENA_PIN_GIB` overrides it; on Linux the whole
arena is registered). The rest remains
available to the CPU pool; the PCIe expert path is available only for the
registered layers. Without secondary GPUs, the original uncapped registration
behavior applies.
On an existing Linux installation, apply the patch to its source tree, then run:
```sh
./setup.sh --setup --gpu1-experts 5000 --gpu2-experts 5000 --gpu3-experts 5000
```
On Windows, from PowerShell in the installation folder:
```powershell
.\START-HERE.bat --setup --gpu1-experts 5000 --gpu2-experts 5000 --gpu3-experts 5000
```
To test the layer placement instead, add `--gpu-placement layer` to the setup
command. To go back, run the same command with `--gpu-placement stripe`. After
the new engine is built, the mode can also be changed without a rebuild: edit
`--expert-cache-remote-placement` in `strata-iq3_xxs.json` from `layer` to
`stripe` or vice versa and restart the server. If the argument is absent, the
default is `stripe`.
Choose the same model, size, context and image settings as your current
installation when prompted. Setup recompiles the engine, updates its server
configuration, and starts it. Later starts keep those settings. To tune the
slots without recompiling, edit the values after `--expert-cache-device1`,
`--expert-cache-device2` and `--expert-cache-device3` in the stored
`strata-iq3_xxs.json`, then restart with `.\run-iq3_xxs.bat`.
Compare identical requests at one, two and four GPUs, preferably with several
repeats. VRAM use alone does not show useful offload: compare the per-request
CUDA1–3 counts and tokens per second. More GPU contexts and synchronization
may lower the speed. Prompt prefill still uses CUDA0; secondary GPUs serve
decode, including MTP verification. No peer-to-peer access is required.
Longer contexts reserve more KV/state memory on CUDA0, which reduces its
automatic expert cache. On a 64 GB PC `setup.py` limits IQ3_XXS to 128K even
if 262K was selected; the secondary caches still require the host expert arena.
To try 262K on 64 GB, choose Q2_0 or IQ2_XS instead and check that CUDA0 still
has enough free VRAM after loading. Benchmark long prompts separately from
short decode requests.
The existing warning about the CUDA0 expert-cache GPU hit path still applies:
its outputs diverge from cache-off runs. Treat performance as experimental
until the generated tokens have been validated.
## Peer tier (`--peer-device`)
`--peer-device N` puts a second adaptive expert cache on CUDA device N. It
takes the ranked pairs CUDA0's cache does not hold, as many as fit. The peer
computes the rows of its own experts, for decode windows and for prompt
chunks; the activations and the results cross NVLink or another P2P path. The
tier adapts while the server runs, like the primary cache. It is an
alternative to the CUDA1-3 caches above, not a third tier beside them.
- `--peer-device N` (default off): enable the tier on CUDA device N (N >= 1).
- `--peer-reserve-mib M` (default 600): leave M MiB free on the peer card; the
cache takes what remains. The prompt-path buffers need this headroom.
- `--peer-slots N` (default 0): cap the tier at N experts; 0 = as many as fit.
- `--peer-adapt-swaps N` (default -1): swaps per adaptive round on the peer;
-1 uses the primary's `--adapt-swaps`.
- `--peer-prefill-rows N` (default -1): the share of each prompt chunk's rows
the peer computes; -1 is half of chunk x top-k, 0 keeps prompt rows on the
primary.
`--peer-device` requires `--expert-profile` and an enabled expert cache, and
the device must be visible; it refuses otherwise. It also refuses
`--layer-split` (a different second-GPU mode: use one or the other) and
`--expert-cache-device1..3` (the peer tier already caches experts on that card):
strata generate: --peer-device cannot be combined with --layer-split (use one or the other)
strata generate: --peer-device cannot be combined with --expert-cache-device1..3 (the peer tier already caches experts there)
Without `--peer-device` the binary is unchanged; its output is byte-identical
to the release. With `--peer-device` and the same expert set split across the
two cards, the generated tokens are byte-identical to the single-GPU run under
the exactness gate.
With a peer the prompt path keeps the MMQ path; the fused int8 prompt path is
not yet combined with the peer's rows. Mapped host buffers gain
`cudaHostAllocPortable` only with a peer, since only then does a second
context write them. The tier size is manual for now (`--peer-reserve-mib`,
`--peer-slots`); automatic sizing on small cards wants the buffer lending of
#216 and is a follow-up.