Winery-Strata / docs /SECOND_GPU.md
Zandy-Wandy's picture
🍷 Winery Strata: fork of Niko1221/Strata (Winery Cuvée family, wine-red UI)
bbb6388 verified
|
Raw History Blame Contribute Delete
7.15 kB

CUDA1–3 expert experiment

CUDA0 keeps the dense layers, KV/state, MTP and its existing expert cache. Up to three independent expert caches fill CUDA1, CUDA2 and CUDA3 in that order. Each expert has one owner: CUDA0, one secondary GPU, or the CPU pool. Each secondary GPU computes its rows while the CPU pool works; the results return through private pinned host buffers before CUDA0 continues the layer. The GPU launches can overlap, but each layer waits for all three.

The shipped profile ranks 8,000 (layer, expert) pairs. CUDA0 takes its usual --expert-cache auto allocation. With only CUDA1, it takes the first remaining pairs. When CUDA2 or CUDA3 is enabled, the remaining profiled pairs are assigned round robin to the secondary GPUs. The ranked list is then extended with all other pairs in expert-then-layer order and distributed in the same way. This fills the additional VRAM and lets all secondary GPUs see frequent experts, but unranked experts may hardly ever be routed. The startup log reports slots and GiB per card; each request reports the number of expert entries actually computed on each secondary GPU. Remote outputs are packed before returning over PCIe/USB4: the request log compares MiB transferred with the previous full-row transfer. The input and routing metadata use pinned host staging. This reduces transfer traffic and launch overhead, but does not release the expert arena in system RAM or significantly change VRAM use.

The optional --gpu-placement layer gives each layer to exactly one secondary GPU (layer number modulo the number of enabled cards). It fills that card with the most frequently routed experts of its layers first, then the unranked ones. This can reduce USB4 device switches and synchronization when several secondary cards serve the same layer in the default stripe mode. All three cards still keep their requested expert slots; the primary CUDA0 cache and CPU fallback continue to work. The log counts active layer launches on each card, so the tradeoff can be measured alongside tokens per second. This placement is experimental and may be slower if a given layer needs more than one card's compute capacity.

The experiment requires the corresponding visible CUDA devices and a CUDA enabled build. Each card must keep at least 512 MiB free. Tiers must be enabled in order. Under WDDM (Windows, WSL2), CUDA registration of the host expert arena is capped at 8 GiB to leave room for the contexts and MTP on CUDA0 (STRATA_ARENA_PIN_GIB overrides it; on Linux the whole arena is registered). The rest remains available to the CPU pool; the PCIe expert path is available only for the registered layers. Without secondary GPUs, the original uncapped registration behavior applies.

On an existing Linux installation, apply the patch to its source tree, then run:

./setup.sh --setup --gpu1-experts 5000 --gpu2-experts 5000 --gpu3-experts 5000

On Windows, from PowerShell in the installation folder:

.\START-HERE.bat --setup --gpu1-experts 5000 --gpu2-experts 5000 --gpu3-experts 5000

To test the layer placement instead, add --gpu-placement layer to the setup command. To go back, run the same command with --gpu-placement stripe. After the new engine is built, the mode can also be changed without a rebuild: edit --expert-cache-remote-placement in strata-iq3_xxs.json from layer to stripe or vice versa and restart the server. If the argument is absent, the default is stripe.

Choose the same model, size, context and image settings as your current installation when prompted. Setup recompiles the engine, updates its server configuration, and starts it. Later starts keep those settings. To tune the slots without recompiling, edit the values after --expert-cache-device1, --expert-cache-device2 and --expert-cache-device3 in the stored strata-iq3_xxs.json, then restart with .\run-iq3_xxs.bat.

Compare identical requests at one, two and four GPUs, preferably with several repeats. VRAM use alone does not show useful offload: compare the per-request CUDA1–3 counts and tokens per second. More GPU contexts and synchronization may lower the speed. Prompt prefill still uses CUDA0; secondary GPUs serve decode, including MTP verification. No peer-to-peer access is required.

Longer contexts reserve more KV/state memory on CUDA0, which reduces its automatic expert cache. On a 64 GB PC setup.py limits IQ3_XXS to 128K even if 262K was selected; the secondary caches still require the host expert arena. To try 262K on 64 GB, choose Q2_0 or IQ2_XS instead and check that CUDA0 still has enough free VRAM after loading. Benchmark long prompts separately from short decode requests.

The existing warning about the CUDA0 expert-cache GPU hit path still applies: its outputs diverge from cache-off runs. Treat performance as experimental until the generated tokens have been validated.

Peer tier (--peer-device)

--peer-device N puts a second adaptive expert cache on CUDA device N. It takes the ranked pairs CUDA0's cache does not hold, as many as fit. The peer computes the rows of its own experts, for decode windows and for prompt chunks; the activations and the results cross NVLink or another P2P path. The tier adapts while the server runs, like the primary cache. It is an alternative to the CUDA1-3 caches above, not a third tier beside them.

  • --peer-device N (default off): enable the tier on CUDA device N (N >= 1).
  • --peer-reserve-mib M (default 600): leave M MiB free on the peer card; the cache takes what remains. The prompt-path buffers need this headroom.
  • --peer-slots N (default 0): cap the tier at N experts; 0 = as many as fit.
  • --peer-adapt-swaps N (default -1): swaps per adaptive round on the peer; -1 uses the primary's --adapt-swaps.
  • --peer-prefill-rows N (default -1): the share of each prompt chunk's rows the peer computes; -1 is half of chunk x top-k, 0 keeps prompt rows on the primary.

--peer-device requires --expert-profile and an enabled expert cache, and the device must be visible; it refuses otherwise. It also refuses --layer-split (a different second-GPU mode: use one or the other) and --expert-cache-device1..3 (the peer tier already caches experts on that card):

strata generate: --peer-device cannot be combined with --layer-split (use one or the other)
strata generate: --peer-device cannot be combined with --expert-cache-device1..3 (the peer tier already caches experts there)

Without --peer-device the binary is unchanged; its output is byte-identical to the release. With --peer-device and the same expert set split across the two cards, the generated tokens are byte-identical to the single-GPU run under the exactness gate.

With a peer the prompt path keeps the MMQ path; the fused int8 prompt path is not yet combined with the peer's rows. Mapped host buffers gain cudaHostAllocPortable only with a peer, since only then does a second context write them. The tier size is manual for now (--peer-reserve-mib, --peer-slots); automatic sizing on small cards wants the buffer lending of #216 and is a follow-up.