|
Download docs/MULTI_GPU.md from WineryLabs/Winery-Strata: direct link, hf CLI and curl.
- Browser
- Download file 8.98 kB
-
https://huggingface.co/WineryLabs/Winery-Strata/resolve/main/docs/MULTI_GPU.md
- Command line
-
hf download hf://WineryLabs/Winery-Strata/docs/MULTI_GPU.md
-
curl -L -o MULTI_GPU.md https://huggingface.co/WineryLabs/Winery-Strata/resolve/main/docs/MULTI_GPU.md
8.98 kB
| # Strata on two or three GPUs (layer split) | |
| One model can run across several NVIDIA cards in one PC. The layers are split into contiguous ranges, one per GPU: | |
| the first card runs layers 0 to K-1, the next card runs K onward, and so on; the last card also runs the output head | |
| and the draft (MTP) layer. Each card keeps an expert cache for **its own layers only**, so two cards hold about twice | |
| the experts one card holds - for the Coder model on a 16 GB + 24 GB pair, nearly all of them, which is where the | |
| speed comes from (decode then barely touches the CPU pool). | |
| This is pipeline (layer) parallelism, not tensor parallelism: a token crosses from one card to the next once per | |
| verify window (a few hundred KB through pinned RAM), not twice per layer. No NVLink or peer-to-peer access is | |
| needed; cards on x4 or x1 slots work, and the PCIe share of each card is probed on its own link. A `--pcie-frac` you give | |
| is every card's share and skips those probes; there is no per-card setting yet. | |
| ## Using it | |
| **Nothing to type.** `START-HERE.bat` (Linux: `./setup.sh`) lists your NVIDIA cards and says for each one whether | |
| Strata can use it: | |
| ``` | |
| Your NVIDIA GPUs: | |
| GPU 0: NVIDIA GeForce RTX 5080, 16 GB VRAM - can be used | |
| GPU 1: NVIDIA GeForce GTX 1080 Ti, 11 GB VRAM - not supported - older than the RTX 20 series (compute capability 6.1; Strata needs 7.5 or newer) | |
| GPU 2: NVIDIA GeForce RTX 3090, 24 GB VRAM - can be used | |
| ... | |
| 1) GPU 0 (NVIDIA GeForce RTX 5080, 16 GB) + GPU 2 (NVIDIA GeForce RTX 3090, 24 GB) together (recommended) | |
| 2) GPU 2 (NVIDIA GeForce RTX 3090, 24 GB) only | |
| 3) GPU 0 (NVIDIA GeForce RTX 5080, 16 GB) only | |
| Which GPUs? [1]: | |
| ``` | |
| When two or more cards can share the model, the two best together are recommended (the newest generation first: | |
| it becomes the main card). A model installed on one card asks once, at its next start, whether to use both from | |
| now on; the answer is kept. | |
| **Choosing yourself** (at setup or at any start): | |
| ``` | |
| --gpus 0,2 these cards together, as nvidia-smi numbers them; the first is the main one. Remembered. | |
| --gpus all every card that can share the model | |
| --gpu 0 one card (at a start: for that start only) | |
| --layer-split auto (default) or the first layer of each later card, e.g. 18 or 16,32 | |
| ``` | |
| **Not supported** (setup says so and names the cards that can be used instead): | |
| - a card older than the RTX 20 series (compute capability below 7.5: GTX 10 and older); | |
| - a card with less than 8 GB of VRAM, together with others (each card holds a copy of the dense weights and its | |
| own prompt buffers) - unless you name it with `--gpus`: then setup says the risk and asks (`--yes` with the named | |
| cards goes ahead); | |
| - Intel GPUs, and a mix of NVIDIA and AMD cards. (AMD cards share a model among themselves: `./setup.sh --backend | |
| hip --gpus 1,0`, see [AMD_HIP.md](AMD_HIP.md).) | |
| Or edit an existing config (`strata-*.json`), then restart: | |
| ```json | |
| "gpu": [0, 2], | |
| "layer_split": "auto" | |
| ``` | |
| **Skip the split when the first card holds everything** (opt-in, 0.1.31): `"split_skip_if_fits": true` in the config | |
| (engine flag `--split-skip-if-fits`, with `--layer-split auto`) runs on the first card alone when it holds every | |
| profiled expert plus the context's KV, the draft layer and the reserve, and says so in the log; otherwise the split | |
| stays. On an R9700 32 GB + RX 9070 XT the R9700 holds all of the Coder's experts: with the flag 4K prompts read at | |
| 1,776 tok/s instead of 1,244 (split) and decode runs at ~60 tok/s instead of ~51 (16K prompts ~5% slower than split). | |
| **Short prompts on a split (0.1.32, #340).** 0.1.30 gave each card's prompt path a loan from its own expert cache, | |
| refilled after every request; on cards that hold nearly all their experts that cost short prompts up to a third of | |
| their speed. 0.1.32 refills all cards at once, uses a smaller streaming ring on a split, and lets a card with free | |
| VRAM keep its own prompt buffers - the same output as 0.1.31, measured on an R9700 + RX 9070 XT: 2K prompts 993 -> | |
| 1,265 tok/s, 16K 1,852 -> 1,950, decode unchanged. `STRATA_SPLIT_OWN=1` (opt-in) gives every card its own buffers: | |
| 2K 1,450 and 16K 2,227 tok/s there, but a full card then keeps a different set of experts resident, so the output | |
| differs from the default's (stable and coherent); `STRATA_SPLIT_OWN=auto` does that only where the buffers are at | |
| most 12% of each card's VRAM. | |
| The engine flags behind it: `--layer-split K1[,K2..]|auto` and `--split-device D1[,D2..]` (the later stages' | |
| devices; default the next visible ones). `--layer-split K --split-device 0` runs both stages on one card sharing | |
| everything - the bit-exact check of the hand-off, not a speed mode. | |
| **auto** tries every placement (all of them for two or three cards; proportional to the free VRAM beyond that) and | |
| keeps the one whose caches would hold the most of the expert profile, hottest pairs weighted most; ties go to the | |
| placement that leaves the fullest card the most room. The startup log prints the choice: | |
| ``` | |
| strata generate: layer split auto: K=19 - the caches hold 11767 of 12288 profiled pairs (fullest device 100%) | |
| strata serve: layer split: layers 0-18 (CUDA0), 19-47 (CUDA1), one hand-off per window | |
| ``` | |
| ## What each card holds | |
| - **every card**: a copy of the dense weights (~3.4 GB for the Coder), its own session state (the KV cache of the full | |
| context), its verify window and its prompt-path buffers, and an expert cache for its layers filled from the profile; | |
| - **the last card**: also the output head and the draft layer (~0.8 GB); | |
| - **host RAM**: the expert arena once, shared by all cards (the CPU pool computes whatever no card holds). | |
| Prompts are read in chunks that flow through the cards in turn; while a later card reads chunk c, the first card | |
| already reads chunk c+1. Conversation checkpoints save and restore every card's state; the adaptive expert swaps copy | |
| into the card that owns the layer. | |
| ## Limits (for now) | |
| - **Works across cards** (bench/results/2026-09-29-layer-split-limits): | |
| - images (`--vision`): each card keeps its own image-position table; | |
| - control vectors and the experimental speed projection: each card holds the vector's tables, switched on and | |
| off per request on all of them; | |
| - KV streaming (`--kv-resident`): each card streams the KV of its own session; | |
| - mid-prompt checkpoints (`--prompt-cache-every`): each card saves its part of a checkpoint when it has read that | |
| chunk; | |
| - the older helper-GPU caches (`--expert-cache-remote`, docs/SECOND_GPU.md): they take the visible GPUs no stage | |
| runs on, and hold only experts no stage's cache holds. On the test rig, a 2080 Ti helper made decoding slower, | |
| as it did without a split: its per-layer round trip costs more than the CPU pool needs for those experts. | |
| - `--mmap-experts` needs a canonical pack (`experts.bin`), with or without a split; a native (IQ) pack says so at | |
| start. | |
| - The prompt path has its own buffers on every card (1.5 GB each at the default 2048-token chunk; `--prefill 1024` | |
| halves that) instead of borrowing cache slots as one card does. An explicit `--expert-cache` on the first card is | |
| capped to leave room for them. | |
| - Under WDDM (Windows, and WSL2) only 8 GiB of the expert arena is pinned (more, mapped into two GPU contexts, | |
| leaves WDDM refusing allocations); the rest streams through the pinned staging ring. A Linux driver has no such | |
| limit, so there the whole arena is pinned (since 0.1.31; the cap cost a 4090 + 3060 split two thirds of its | |
| prompt speed, #253). `STRATA_ARENA_PIN_GIB=N` pins at most N GiB, `0` the whole arena, on any OS. | |
| - Every card needs compute capability 7.5 (RTX 20 or newer). The pre-sm_80 QSA scorer path is fp32 FMAs, so a | |
| Turing card runs the same kernels instead of the tensor-core prompt attention. | |
| ## Measured | |
| The Coder on an RTX 5080 + RTX 3090 (Ryzen 9 9950X3D), 32K context; details in | |
| `bench/results/2026-09-29-layer-split/`: | |
| | | Prompt 16K / 28K tok/s | Decode story / code tok/s | | |
| |---|---|---| | |
| | 5080 alone | 1,726-2,017 / 1,970 | 83-87 / 88-105 | | |
| | 5080 + 3090, best split (K=26) | 2,039 / 2,357 | 84 / 110 | | |
| | 5080 + 3090, auto (K=22) | 2,037 / 2,073 | 80 / 109 | | |
| - **Prompts gain the most** (+18-20%): each card reads its own layers of the chunk while the other reads the next. | |
| - **Decode is on par with the faster card alone**, and ahead on code. Once both caches hold nearly every routed | |
| expert, the per-layer GPU time decides. | |
| - **Correctness:** one GPU is byte-identical to 0.1.20, and the hand-off itself is bit-exact. | |
| **Which cards and in what order:** | |
| - Put the fastest card first; auto gives it as many layers as its cache allows. | |
| - Leave out a much slower card when two already hold the model. An RTX 2080 Ti as a third card made the 5080 + | |
| 3090 pair slower (68 / 90 tok/s decode): every extra card costs its own round per window. | |
| - More cards pay off when the model's routed experts do not fit the faster ones. | |