File size: 76,355 Bytes
bbb6388 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 | # Strata - the details
The technical side of Strata: every measured number, the API, images, all settings and how the engine works.
New here? Start with the [README](../README.md); installing step by step is in [INSTALL.md](INSTALL.md), the models in
[MODELS.md](MODELS.md), common problems in [TROUBLESHOOTING.md](TROUBLESHOOTING.md).
> **On this page:** [Speed](#speed-measured) · [Other GPUs](#other-gpus-estimated) · [Which model?](#which-model) ·
> [Requirements](#before-you-start) · [Windows](#windows) · [Linux](#linux) · [API](#using-it) ·
> [MCP tools](#tools-from-mcp-servers) · [MCP server](#manage-strata-from-your-ai-assistant-mcp-server) ·
> [Images](#images-vision) ·
> [Troubleshooting](#troubleshooting) · [How it works](#how-it-works)
---
## Speed (measured)
RTX 5070 **12 GB**, Ryzen 5 7600 (6 cores), 64 GB DDR5-5200, Windows, engine 0.1.26 with the settings setup writes
(`--prefill auto`, 8-bit KV above 4K, KV streaming from 64K). One code-agent prompt per length, 256 generated tokens,
MTP speculative decoding on. "262K" is the model's full context window (a 259,943-token prompt). The IQ2_XS row was
measured with Swift 1.5's IQ2_XS, which runs at the original's speed.
**Engine 0.1.36 (#136), the same PC:** Q2_0's prompt experts run on fused int8 tensor-core kernels (RTX 30 and newer):
4K 1,294 -> 1,570, 32K 2,170 -> 2,653, 128K 2,123 -> 2,468 tokens/s (+16-22%), as close to an FP16 reference as the
previous kernels (closer at 32K: teacher-forced KL 0.009 vs 0.012). The decode path's block selection and greedy
argmax run on thread-block clusters (RTX 50, sm_90+; other cards keep the previous kernels; the same tokens): Q2_0 output at 4K 89 -> 93.5, at 128K
64.5 -> 76.4 tokens/s. `STRATA_PF_FUSED=0` keeps the previous prompt kernels (byte-identical answers to 0.1.35);
`STRATA_PF_FUSED=1` also runs the native IQ packs' fused kernels (opt-in: IQ2_XS prompts +12% at 4K, +3% at 32K, the
IQ3 packs about even); `STRATA_QSA_CLUSTER=0` / `STRATA_ARGMAX_MULTI=0` turn the decode kernels off. The tables
below are 0.1.26's.
### Prompt processing (tokens/s)
| Model | 1K | 4K | 32K | 64K | 128K | 262K |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| **Q2_0** | 536 | 1,299 | 2,171 | 2,126 | 2,107 | 1,304† |
| **IQ2_XS** | 534 | 1,256 | 2,092 | 1,754 | 1,752 | 1,181*† |
| **IQ3_XXS** | 482 | 1,007 | 1,745 | 1,609 | 1,602 | - |
| **IQ3_S** | 427 | 913 | 1,624 | 1,640 | 1,443 | - |
| **Coder** | 656 | 1,583 | 2,177 | 2,236 | 2,208 | 1,034** |
Engine 0.1.26; `bench/results/2026-09-29-speed-0126`. At 32K-128K that is 8-28% faster than 0.1.22. † not measured
again: 0.1.22. \* measured with images on (the image encoder's VRAM reserve leaves fewer experts cached). \*\* not
measured again: 0.1.14.
### Output (tokens/s)
| Model | 1K | 4K | 32K | 64K | 128K | 262K |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| **Q2_0** | 87.3 | 93.0 | 81.8 | 76.2 | 73.7 | 60.3† |
| **IQ2_XS** | 79.6 | 78.6 | 76.3 | 63.7 | 62.7 | 52.8† |
| **IQ3_XXS** | 61.9 | 61.6 | 58.5 | 57.2 | 49.0 | - |
| **IQ3_S** | 52.4 | 53.3 | 48.3 | 46.3 | 45.5 | - |
| **Coder** | 58.9 | 55.1 | 54.9 | 53.2 | 43.0 | 42.8† |
Engine 0.1.26, the same runs. † not measured again: 0.1.14.
Output speed depends on the text as well: speculative decoding runs faster when more of the drafted tokens are
accepted, so a different answer to the same prompt moves it by several percent. Run back to back on the 4K prompt,
0.1.14 writes 88.5 tokens/s and 0.1.12 85.7. The numbers before 0.1.13 (prompts about half as fast):
[`bench/results/2026-09-24-final`](../bench/results/2026-09-24-final/matrix.md); these:
[`bench/results/2026-09-28-speed-0114`](../bench/results/2026-09-28-speed-0114/README.md).
IQ3_XXS and IQ3_S at 262K are not measured: with their 43 / 50 GB of experts, a 260K-token context brings a 64 GB PC
to its memory limit by setup's estimate (the experts + the context's KV cache + 24 GB), so setup recommends up to 128K
with them on 64 GB. A longer context you choose (`--context 262144`, or a pick in its list) is kept, with a note: users
ran IQ3_S at 256K on 64 GB with RAM to spare (#406). In the low-RAM mode the KV cache stays in VRAM and the context
does not count against RAM. IQ3_S (engine 0.1.4 or newer) is only published for the original model, not for Swift 1.5.
**KV streaming (engine 0.1.5):** at 64K and more, setup keeps the context's KV cache in RAM and only the part the
attention reads in VRAM (`--kv-resident 32768`), so more experts fit on the GPU. Q2_0 at 262K: 50.9 -> 62.6 tokens/s
(1,589 -> 3,872 experts in VRAM); at 128K about +6%. The attention reads exactly the same values (only where the KV lives changes); it
costs ~13.7 KB of RAM per context token (1.7 GB at 128K). Existing installs: run `START-HERE.bat --setup` once to turn
it on. Setup turns it on when the RAM has room for it; `--kv-streaming on|off` overrides that (on past the RAM test with a
note; never with `--kv k8v4` or under WSL, which cannot stream).
**4-bit KV cache (engine 0.1.8, optional):** `START-HERE.bat --setup` asks above 8K context (or pass `--kv q4_0`). It
halves the KV cache's memory with a Hadamard rotation before 4-bit rounding (PR #21), about 4% faster at 128K, but it
is measurably less precise on long documents (perplexity +8-12%; needle tests still pass). 8-bit stays the default.
Details: [`bench/results/2026-09-27-kv-q4`](../bench/results/2026-09-27-kv-q4/README.md).
**Hybrid K8V4 KV cache (engine 0.1.25, optional, PR #120):** `--kv k8v4` (`START-HERE.bat --setup --kv k8v4`) keeps
the keys at 8 bits and stores the values as rotated 4-bit: 23% less KV memory than 8-bit, so more experts fit in
VRAM. RTX 3090, the Coder at 198K context: 99 instead of 85 tokens/s output, the same needle results, prompts 2-5%
slower. It does not stream its KV cache (KV streaming is on by default from 64K), so it pays off mostly on large
cards at long contexts.
**Reproducible greedy output (0.1.30, opt-in, `STRATA_IQ_MT_MIN=1`):** with the IQ models, the CPU computes an
expert for one token with ggml's dot product and for several tokens with Strata's multi-token kernels, which round
slightly differently. How many tokens share an expert depends on the drafts in a verify window, so the same prompt
at temperature 0 can end in a different (equally good) answer when the drafting, the cache state or a resumed
conversation differ (issue #152). `STRATA_IQ_MT_MIN=1` (in the config's `env`) uses the multi-token kernels for
every group: the answer then no longer depends on the drafting. Measured on a Ryzen 7600 (AVX-512): IQ3_S decode
-1..-3%, the other models the same; the default stays the fastest rule. Through the server, two more things carry
over from one request to the next (#410): the adaptive tier moves experts between RAM and VRAM (the GPU and the CPU
round an expert differently), and the prompt cache resumes a repeated prompt and reads only its tail through the
decode path. For byte-identical repeats add `--prompt-cache 0 --adapt-swaps 0 --pcie-frac 0` to the engine's args
as well (#410): the PCIe share of the missed experts (computed on the GPU instead of the CPU) still made the first
answer after a start differ from the next ones. Measured here (IQ3_XXS, a 3.6K-token prompt, 4 repeats): with all
three switches 1 answer of 4, without `--pcie-frac 0` 2 of 4 (the first one differs), with the defaults 2 of 4.
`--pcie-frac 0` costs decode speed (the missed experts all run on the CPU), so keep it for A/B runs.
**The draft layer's tokens (0.1.27, `--draft-vocab`):** the MTP draft layer can only propose tokens from a subset
of the vocabulary (`mtp/rt/draft_vocab.bin`). Since 0.1.27 the subset includes every Chinese, Japanese and Korean
token (106,299 ids), so answers in those languages are 15-38% faster (Q2_0, RTX 5070). Its head takes ~180 MiB of
VRAM, which the expert cache leaves free for it (0.1.28). `START-HERE.bat --setup --draft-vocab en` keeps the
English/code subset from before (40,525 ids, ~110 MiB less VRAM, English answers 1-2% faster; CJK answers get
almost no drafts). `--draft-vocab cyrillic` takes the English/code subset plus the whole Cyrillic script (58,963
ids): the shipped subsets hold 142 of the vocabulary's 18,580 Cyrillic tokens, so Ukrainian or Russian answers got
1.4 tokens a round; with it 2.1, and 83 -> 109 tokens/s (RTX 5090, the NVFP4 fork), English unchanged.
`tools/draft_vocab.py` builds and inspects subsets. When the start stops with "the draft head does not fit" (a
12 GB card with a long context, #474), the engine says how much the head needs, how much VRAM is free and which
smaller subset fits, and the server's start error repeats it; setup suggests `--draft-vocab en` on cards under
14 GB (only a suggestion: nothing changes unless you pass it).
**Low-RAM mode (engine 0.1.26, chosen by setup):** normally all of a model's experts are copied into RAM (23-50 GB,
pinned) and the GPU holds a copy of the most-used ones. On a PC whose RAM cannot hold them beside the system (the
experts plus ~10 GB), setup instead maps them from one file in the model's folder (`--mmap-experts`, the pack's
`experts.bin`, +23-50 GB of disk). The OS file cache holds what the GPU does not, and it can give that memory back.
On the Coder the engine's committed memory drops from 36 to ~13 GB, with the same answers. With a big GPU (an RTX
5090 holds all of the Coder's experts, most of Q2_0's) it runs at nearly the usual speed. With a small one, most
experts come from the SSD and it is much slower (setup says so). `START-HERE.bat --setup --low-ram on|off` overrides
the choice.
**Low-RAM mode, resident (engine 0.1.30):** when the experts the GPU does not hold fit the RAM (with the same ~10 GB
beside them), setup picks the resident variant instead (`--resident-experts`): at start the engine copies exactly those
experts from `experts.bin` into RAM (page-locked when the driver allows, else locked in RAM), so while it answers
nothing is read from the SSD, however little RAM the OS leaves for its file cache. Examples with setup's context: a
32 GB PC with a 24 GB GPU runs Q2_0, IQ2_XS and the Coder this way (~16-18 GB of experts in RAM, the GPU holds the
other ~18 GB), a 32 GB PC with a 12-16 GB GPU the Coder; IQ3_XXS on a 32 GB PC stays mapped. The details:
- The prompt path borrows room in the GPU's expert cache for its buffers and puts those experts back after the prompt;
as far as the RAM allows, their experts are kept in RAM too (so a prompt reads nothing from the SSD either).
- The cache still follows the conversation (`--adapt-every`): a swap copies the evicted expert back from VRAM into the
RAM place of the one that replaces it, so the RAM copy keeps holding exactly what the GPU does not.
- The answers are the plain mapped mode's for the same expert placement: the bytes are the file's. With a page-locked
copy the GPU also takes its usual share of the misses over PCIe (`--pcie-frac`), as with enough RAM; `--pcie-frac 0`
(or `STRATA_RESIDENT_PIN=0`) gives the mapped mode's exact tokens.
- The engine leaves 4 GB of the RAM it finds free (`STRATA_RESIDENT_HEADROOM_GIB`); when even the experts the GPU does
not hold do not fit, it says so and runs the plain mapped mode. The server log shows, per request, how many expert
reads went to the file (`resident RAM: ... blob reads from the file`: 0 in steady use).
- `--low-ram resident|mmap` forces one variant (also on a PC with enough RAM, e.g. to try it).
- Several GPUs (#364, #384): setup recommends one GPU in the low-RAM mode (the resident variant has no layer split
yet), and asks; `--gpus 0,1` (or answering 2) shares the model across them with the mapped variant
(`--mmap-experts`): the cards together hold more of the experts, and two users measured it 1.3-1.6x faster than
one card, but the OS file cache can fill the RAM to 0 free during long prompts. `--yes` keeps one GPU. A config
with `--resident-experts` started with `--gpus` switches to `--mmap-experts` with a note, and the engine runs that
pair as `--mmap-experts` with a warning instead of refusing it.
**Low-RAM mode without `experts.bin` (engine 0.1.31):** for the native packs (IQ2_XS, IQ3_XXS, IQ3_S, the Coder, Swift,
Q2_0 packed by `tools/iq_pack.py`; not the canonical Q2_0 pack setup makes for AVX-512 CPUs) the mapped mode no longer
needs the pack's `experts.bin`: when the pack has none, the engine
maps the model's GGUF files themselves and reads each expert's gate, up and down rows from where `native_experts.txt`
says they are (the files are checked against it first: every tensor's name, type, shape, offset and bounds). That
saves the 23-50 GB copy on the disk. The answers are the same: on the Coder, 64 greedy tokens from `experts.bin` and
from the GGUF gave identical tokens and logits. An expert read from the GGUF is three reads instead of one, so the
engine fetches a layer's missing experts on 8 threads (`STRATA_FETCH_THREADS`) with one batched page request
(Windows `PrefetchVirtualMemory`). With an `experts.bin` in the pack, nothing changes. Setup does not use this yet.
**A RAM budget (engine 0.1.31, `--resident-budget-gib N`):** the resident variant for a model whose experts do not all
fit: the N GiB of experts the GPU cache does not hold that the expert profile ranks hottest are copied into RAM at
start (locked; page-locked when the driver allows the whole budget), and the rest are read from the files through the
OS file cache. It implies `--mmap-experts` and leaves 4 GB of free RAM (a larger N is clamped to that less 256 MiB,
with a message; #403: a clamped budget no longer fails the safety check that follows, and a budget that cannot be
kept at all is a warning, with every expert read from the files). Setup sets N with `--resident-budget-gib N`. With
the GGUF read in place it also warms the next layer's likely experts: while the CPU works on a layer, a thread applies
the next layer's router to this layer's input and asks the OS for the pages of the predicted experts that neither the
GPU nor the RAM budget holds (only pages - the experts computed are the same; `STRATA_LOOKAHEAD=0` turns it off). This
is what runs [Unsloth's UD-Q4_K_XL](UNSLOTH_Q4.md) (72 GiB of experts) on a 64 GB PC: 7-8.5 tokens/s at N = 40 on an
RTX 5070, against ~3 tokens/s before these changes.
**How much came from where:** with `--stats` the engine prints the tiers of the decode (`expert tiers`: blobs from the
RAM copy, blobs and MB from the files, the time spent reading them; `routing prefetch`: how many of the file reads had
been warmed). The server log has the same per request (`expert tiers: GPU ... hits ...; RAM ... blobs, files ...
blobs ... MB read`), and `GET /metrics` lists `ram_blobs`, `file_blobs` and `file_mb` for each recent request (with
engine 0.1.31 or newer). It also lists each request's speculative drafts, `drafts_offered` and `drafts_accepted`
(`null` when the engine did not report them), and their sums since the server started in `totals` (#457).
Time to first token is prompt length / prompt speed: with Q2_0 about 4 s at 4K, 25 s at 32K, under 2 minutes at 128K
and 4.5 minutes at 262K (engine 0.1.13 made long prompts about twice as fast, below).
**Faster prompts (engine 0.1.13):** the prompt is read in chunks of up to 8,192 tokens instead of 2,048 (`--prefill
auto`: the largest chunk whose buffers fit in the expert-cache slots it borrows, and a request borrows only what its
prompt needs); the experts are multiplied by llama.cpp's quantized MMQ kernels instead of being expanded to FP16
first; the next layer's experts stream over PCIe while the current layer's attention runs; the PLE block runs for the
whole chunk at once; unpinned experts are copied by helper threads. Measured on the RTX 5070 12 GB, 64 GB RAM,
32K-token prompt: Q2_0 572 -> 1,290 tokens/s, IQ3_S 383 -> 1,208. Through the server (Q2_0, 128K context): 999 tokens
353 -> 438 tokens/s, 6,927 tokens 529 -> 1,077, 28,584 tokens 584 -> 1,249. Output speed is unchanged. Needles 5/5
(1K-262K). Details and the quality check:
[`bench/results/2026-09-28-prefill-speed`](../bench/results/2026-09-28-prefill-speed/README.md). Existing installs
switch to `--prefill auto` the next time START-HERE / setup.sh starts them. The raw numbers:
[`bench/results/`](../bench/results/). The [paper](paper/Strata-Paper.pdf) explains every number.
## Other GPUs (estimated)
Not measured - estimated from the runs above (same CPU and 64 GB RAM): the GPU part scaled by memory bandwidth, the CPU
part by how many more experts the card's VRAM holds. Treat as **±20%**. Numbers are *prompt / output* tokens/s.
The prompt figures predate engine 0.1.13, which about doubled prompt speed on the measured card; how much of that a
card gains depends on its PCIe link (the experts stream over it), so they are still the older estimates.
| GPU | Model | 1K | 4K | 32K | 64K | 128K | 262K |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| RTX 5060 Ti 16GB | Q2_0 | ~341 / ~80 | ~472 / ~87 | ~501 / ~81 | ~492 / ~72 | ~476 / ~62 | ~435 / ~53 |
| | IQ2_XS | ~291 / ~80 | ~406 / ~77 | ~434 / ~63 | ~426 / ~62 | ~413 / ~51 | ~383 / ~47 |
| | IQ3_XXS | ~249 / ~66 | ~359 / ~65 | ~381 / ~56 | ~374 / ~54 | ~363 / ~45 | - |
| RTX 3090 24GB | Q2_0 | ~355 / ~128 | ~491 / ~140 | ~521 / ~130 | ~512 / ~115 | ~495 / ~100 | ~453 / ~85 |
| | IQ2_XS | ~303 / ~131 | ~422 / ~128 | ~451 / ~103 | ~444 / ~102 | ~430 / ~85 | ~398 / ~78 |
| | IQ3_XXS | ~260 / ~106 | ~374 / ~103 | ~396 / ~89 | ~390 / ~85 | ~378 / ~71 | - |
More VRAM matters more than a faster GPU: every extra GB holds ~700 more experts, and every expert on the GPU is one the
CPU does not have to compute. A 3090's 24 GB takes most of the CPU work away. (Since 0.1.14 the expert profile ranks
all 24,576 experts; before, the cache stopped at 8,000, about 10-14 GB. `tools/make_profile.py` builds a profile from
your own prompts: run the engine once with `--dump-routing trace.bin`, see the tool's help.)
## Which model?
All three are [ISTA-DASLab's GSQ-RCO quantizations](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF)
of [Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
| Model | Download | RAM it uses | Speed | Quality |
| --- | ---: | ---: | --- | --- |
| **Q2_0** | 66 GB | ~34 GB experts + ~6 GB | fastest | good |
| **IQ2_XS** | 68 GB | ~36 GB experts + ~6 GB | close to Q2_0 | a bit better |
| **IQ3_XXS** | 76 GB | ~43 GB experts + ~6 GB | slower (more CPU work) | best |
With 64 GB of RAM all three fit (close the browser for IQ3_XXS, and keep its context at 128K or less). With 48 GB only Q2_0 / IQ2_XS may fit. With 32 GB: the Coder (below).
### Or: the Coder (half the experts, for code)
**[Qwen3.8-Flash-Next GSQ-RCO Coder](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF)** is
ISTA-DASLab's expert-pruned release: 256 of each layer's 512 experts are kept (still 10 active per token), chosen with
RCO on code, agentic and vision calibration data; its authors report 91.3% of the full model's SWE-bench Verified and
98.7% of LiveCodeBench v6. One size, named IQ1_M for its 1.89 bits per *original* parameter; the kept experts are
stored like IQ3_S (IQ2_S-IQ4_XS gate/up, IQ4_NL/Q2_0 down). Shard 1 is 29.6 GB (experts: 23 GB of RAM), so it runs
on **32 GB of RAM**, and at 262K on 64 GB. Its shard 2 and its vision encoder are the original's files: with the
original installed, setup downloads only shard 1. Strata ships its expert profile (`data/expert-profile-coder.bin`,
the shipped ranking mapped onto the kept experts through the release's `rco-allocation.txt`: 72% of the expert
reads hit the GPU on a 12 GB card). Images work; the experimental speed projection loads and runs on it (it was made
for the full model).
```
START-HERE.bat --setup --family coder
```
### Or: Swift 1.5 (a fine-tune that thinks shorter)
The setup's first question also offers **[Swift 1.5](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF)**,
UkisAI's fine-tune of Qwen3.8-Flash-Next, trained to reach the answer with much less thinking (its authors: 63% fewer
thinking tokens, 1.8x sooner answers, under 1% accuracy loss). Same architecture, the same three sizes, its own
vision encoder; Strata runs it at the same speed (4K, IQ2_XS: 465 prompt / 78.7 output tokens/s, vs 467 / 78.3 for
the original). Its authors recommend **IQ2_XS** (their Q2_0 is marked experimental). Its license is the Swift Open
License 1.0 - read it on the model page.
Our small check (8 reasoning questions, default thinking, IQ2_XS): both models got **8/8**; Swift used **1,234**
output tokens in 28 s, the original **2,682** in 46 s - most of the difference from one question the original
thought about for 1,524 tokens. Not a benchmark, but consistent with the claim.
```
START-HERE.bat --setup --family swift --model IQ2_XS
```
### Experimental: Unsloth's UD-Q4_K_XL
A 4-bit quantization of the same model (111 GB, 72 GiB of experts). Setup offers it from engine 0.1.32
(`--family unsloth --model UD-Q4_K_XL`: a RAM budget of your RAM less 24 GB, the rest read from the SSD); what it
does and the manual workflow are in **[docs/UNSLOTH_Q4.md](UNSLOTH_Q4.md)**. On a 64 GB PC with a 12 GB RTX 5070 it writes 7-8.5
tokens/s, most experts read from the SSD; it picks the same tokens as llama.cpp on the same file at 97.5-99% of
the positions of short greedy answers, 90-91% after a 16K prompt, differing mostly at near-ties
([measured](UNSLOTH_Q4.md#quality-against-llamacpp-on-the-same-file)).
## Before you start
You need **only a graphics driver**: NVIDIA 580 or newer (update it with the NVIDIA App or from
[nvidia.com/drivers](https://www.nvidia.com/drivers)), or for AMD the one in [INSTALL.md](INSTALL.md#what-you-need).
Everything else is installed for you the first time.
| | |
| --- | --- |
| GPU | NVIDIA **RTX 20, 30, 40 or 50 series**, **12 GB VRAM or more** (8 GB runs, slowly). Measured on an RTX 5070 and an RTX 3090; RTX 20 (Turing, since 0.1.27) was tested by a contributor on an RTX 2070. Or AMD **Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series**: [AMD_HIP.md](AMD_HIP.md). |
| RAM | **64 GB** recommended (see the table above). |
| CPU | x86-64 with AVX2 (any Intel/AMD desktop CPU from the last ~8 years). AVX-512 (Ryzen 7000/9000) is a bit faster. |
| Disk | ~70-80 GB free for the model, ~6 GB for the MTP layer (+1 GB with images). **Q2_0 on an AVX-512 CPU** also writes a one-time ~40 GB copy of its experts for the fast CPU kernel. An NVMe SSD is strongly recommended. |
| OS | Windows 10/11, or Linux (Ubuntu 22.04/24.04 get everything installed automatically). |
What the first start installs: in this folder `.venv/`, `engine/` and `third_party/`; the model files (`models/`,
`packs/`, `mtp/`, 70-120 GB) in **`Strata-data` next to this folder**, so a new copy of Strata (an update unzipped
elsewhere) finds them and sets itself up the same way. The place is remembered per user (`%APPDATA%\Strata\settings.json`,
`~/.config/strata/settings.json`); `--data-dir` chooses another. Installs from before 0.1.16 are moved there by the next
start (a rename on the same drive; files on another drive are used where they are).
Python 3.12 if you have none (for your user account, no admin), a private Python environment, NVIDIA's CUDA libraries
(from pip, ~0.4 GB), the ready-made Strata engine for RTX 20/30/40/50, the model and the MTP draft layer. If no
ready-made engine fits your PC, it offers to install the build tools (Visual Studio Build Tools + CUDA Toolkit on
Windows, `build-essential` + CUDA on Ubuntu) and compiles the engine for your GPU (asks first; 20-40 minutes once).
---
## Windows
### Double-click `START-HERE.bat`
**The first time** it asks four questions and does the rest:
1. **Which model?** Qwen3.8-Flash-Next (the original) or Swift 1.5 (the fine-tune that thinks shorter).
2. **Which size?** Q2_0, IQ2_XS or IQ3_XXS (it recommends one for your RAM).
3. **How much context?** 8K to 256K tokens (it recommends one for your VRAM).
4. **Images?** yes / no (see [Images](#images-vision)).
Then it downloads and prepares everything (the model is 66-76 GB, so the first start takes a while; an interrupted
download continues where it stopped) and **starts the model**: your browser opens `http://127.0.0.1:8080`, the Strata
app. It has three tabs:
- **Chat:** streaming answers, the model's thinking (folded away once it answers), code with a copy button, pictures when
images are on, and sampling and thinking-level settings. Chats stay in your browser.
- **Monitor:** what the model is doing (reading the prompt, with progress, or writing, at how many tokens/s); GPU load,
VRAM, temperature, power and PCIe traffic; CPU, RAM and disk; the context in use; the last requests.
- **About:** the model and engine settings, and the addresses to connect other apps.
`http://127.0.0.1:8080/?q=your question` opens it with a new chat already asking. The API is at
`http://127.0.0.1:8080/v1` for your apps.
**Every time after that**, `START-HERE.bat` just starts the model (30-90 s to load 34-43 GB into RAM). Nothing is
downloaded again. Closing the window stops the model.
```
START-HERE.bat --setup install another model, or change context / images
SETUP.bat the same (double-click it)
START-HERE.bat --model IQ2_XS --context 32768 --vision yes --yes no questions
START-HERE.bat --gguf-dir D:\models\IQ2_XS use GGUF files you already have
START-HERE.bat --data-dir E:\Strata-data keep the model files somewhere else
START-HERE.bat --port 8081 another port
START-HERE.bat --gpu 1 another GPU (numbered as nvidia-smi; setup picks the one with the most VRAM)
START-HERE.bat --calibrate tune the engine for this PC (about 5-10 minutes), then start
```
With more than one model installed, it asks which one to start. `run-<model>.bat` starts a model directly.
**Tuning for your PC (`--calibrate`, engine 0.1.19).** Three engine settings depend on the PC more than on the model:
- the share of the experts missing from VRAM that are copied to the GPU instead of computed by the CPU
(`--pcie-frac`: a fast PCIe link and a slower CPU want more, a laptop's narrower link less);
- how sure the draft layer must be to add another guess to a check (`--spec-min-p`);
- how many CPU threads compute experts (`--pool-workers`: on CPUs with efficiency cores, fewer can be faster).
The defaults were measured on a Ryzen 5 7600 with an RTX 5070. Setup offers to measure them on your PC after an
install; `START-HERE.bat --calibrate` (Linux: `./setup.sh --calibrate`) does it any time. It measures the output
speed with each setting and keeps one only when it is more than 3% faster. The result is remembered per PC and model
(in the settings file next to the data folder's record), so updates keep it.
### Running it at startup (Task Scheduler)
To have the model up at logon, people start the serve from **Task Scheduler** (or a service). Beware: Windows
throttles such contexts, and the model's ~40 GB expert load then crawls at **~0.05 GiB/s (13-14 minutes)**
instead of **~1.4-1.5 GiB/s (~35 seconds)** - a 24x slower start. Measured on an RTX 5070 Ti + Ryzen 7 9800X3D
+ NVMe, same binary, same args, same cache state:
| How the serve starts | Expert load |
| --- | ---: |
| Double-click / terminal / SSH | 1.42-1.52 GiB/s (~35 s) |
| Task Scheduler with its defaults | 0.05 GiB/s (821-841 s) |
| Task Scheduler with the two settings below | 1.42 GiB/s (35 s) |
In the task's properties set both of these (the defaults are the opposite):
- **Priority level: Normal** (Options tab; the default is Below normal), and
- **Run with highest privileges** (General tab; without it the task runs with a limited user token - which
also strips `SeLockMemoryPrivilege`, the privilege Windows large pages need).
(Both were changed at once, so the isolated effect of each is not measured.) If the model still starts
slowly, the engine prints a hint under its `loaded ... GiB at ...` line naming this cause.
### Chat in the terminal (optional)
```
.venv\Scripts\python chat.py
```
---
## Linux
```bash
./setup.sh
```
The same questions, the same automatic install (it uses `sudo apt` for Python and, only if it has to compile,
for the build tools), and the same start: `http://127.0.0.1:8080`. Later runs of `./setup.sh` (or `./run-<model>.sh`)
start the model directly. Options as on Windows (`./setup.sh --setup`, `--model Q2_0 --yes`, `--gguf-dir /data/Q2_0`).
Terminal chat: `.venv/bin/python chat.py`.
- **Updating:** `git pull`, then `./setup.sh`: it compiles the engine again when its source changed (a minute or
two for the changed files). If that compile fails, it says so and starts the engine you had.
- **Other distributions** (Arch, Fedora, ...): install the C++ compiler and the CUDA Toolkit 13 with your package
manager first (Arch: `sudo pacman -S base-devel cuda`); setup finds `nvcc` on PATH, in `/usr/local/cuda*` and in
`/opt/cuda*`, and does the rest.
- **WSL** works (Ubuntu 24.04 tested), with one limit: the NVIDIA driver pins only about 1 GB of RAM there, so KV
streaming (`--kv-resident`) is off and the KV cache stays in VRAM, and the experts are copied to the GPU from
unpinned RAM (slower prompts than native Linux).
---
## Sharing the GPU with other programs (optional)
By default the model stays loaded until you close Strata. On a PC that also games, renders or runs another model
server, three server options (all off by default; also as keys in `strata-<model>.json`) give the VRAM back:
| Option | Config key | What it does |
| --- | --- | --- |
| `--idle-unload 600` | `"idle_unload_s": 600` | unload the model after 600 s without requests; the next request loads it again |
| `--min-free-vram-mib 11000` | `"min_free_vram_mib": 11000` | load an unloaded model only when that much VRAM is free (it waits up to 15 s for memory being given back), else answer **503** "the GPU is in use by another program" instead of starting into what a game left (with several GPUs it checks the first one) |
| `--before-load "cmd"` | `"before_load": "cmd"` or `["cmd", "arg"]` | a command run before the model is loaded again, e.g. one that unloads another server's model |
`POST /unload` unloads it now (`409` while a request is running) and `POST /load` loads it ahead of a request (both
with `Content-Type: application/json`, e.g. `curl -X POST -H "Content-Type: application/json" localhost:8080/unload`);
`/health` says `"loaded"`, `/v1/models` lists it as `unloaded` (like llama.cpp's router), `/props` sets
`is_sleeping` and the Monitor shows the state. Unloading ends the engine process - and the image encoder, when images
are on; it is started again first, as at a start - so their VRAM and RAM go straight back. The model files stay in
the OS file cache, so loading again takes seconds while that RAM is not needed elsewhere. Measured on an RTX 5060 Ti
16 GB with Q2_0 in the low-RAM mode: unloading takes ~0.3 s, and a request to an unloaded model answered after
4.6 s (text) or 14.7 s (a picture, image encoder on the CPU).
**Keep what the expert cache learned across restarts (opt-in, engine 0.1.36, #477):** a start fills the GPU's expert
cache from the shipped profile, and the adaptive tier (`--adapt-every`) then moves in the experts your requests use.
With `"expert_profile_save": "expert-profile-learned.bin"` in `strata-<model>.json` the engine saves that as a
profile - the experts in VRAM first, then the routing it counted since the start, then the shipped order - on a
clean exit and every 10 minutes between requests (`"expert_profile_save_every": 5` for another interval, `0` for
exit only), written to a temporary file and renamed, so a crash never leaves half a file. The next start begins from
it instead of the config's `--expert-profile` when it is a profile of the same model (else from the config's, as
before). A relative path is in the Strata folder; one file per model, and a profile per project works the same way
(point the key at another file). The file is a fingerprint of what you used the model for: it stays on your PC.
Without the key nothing is counted or written. Setup rewrites the config when run again: add the key again then.
---
## Using it
The server listens on `http://127.0.0.1:8080` (change with `--port` in setup, or edit the run script).
| API | Endpoint |
| --- | --- |
| OpenAI Chat Completions (stream and non-stream, tools) | `POST /v1/chat/completions` |
| Anthropic Messages (stream and non-stream, tools) | `POST /v1/messages` |
| Model list / health | `GET /v1/models`, `GET /models`, `GET /health` |
| Model properties | `GET /props` (also accepts `?model=<loaded-model-id>`) |
| What the model is doing right now | `GET /status`, `GET /slots` (single slot, busy or idle) |
| Everything the Monitor tab shows (engine, live state, last requests, hardware) | `GET /metrics` |
| The MCP servers, their state and tools ([below](#tools-from-mcp-servers)) | `GET /mcp` |
`/models` and `/v1/models` list only the loaded model, with its context limit and input modalities. `/props` exposes the original chat template, context limit, configured generation defaults (shared settings take precedence), model path and engine version when available. Context means the full engine context, not the resident KV window. `n_predict: -1` means no fixed output cap. Unconfigured sampling fields are omitted. `autoload` has no effect; an unknown `model` returns 404. These metadata endpoints and `/slots` require the API key when one is configured. They do not load, unload or restart models.
```bash
curl http://127.0.0.1:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "strata", "messages": [{"role": "user", "content": "Write a haiku about GPUs."}], "max_tokens": 512 }'
```
```python
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="none")
r = client.chat.completions.create(model="strata", messages=[{"role": "user", "content": "Hello!"}])
print(r.choices[0].message.content)
```
- **Thinking levels: none, low, medium, high.** The model thinks before it answers (streamed as
`reasoning_content`, Anthropic: `thinking` blocks). Choose how much per request - in the chat page (the "Thinking"
menu), in `chat.py` (`/think low`), or over the API:
| API | how |
| --- | --- |
| OpenAI | `"reasoning_effort": "none" \| "low" \| "medium" \| "high"` (also `"reasoning": {"effort": ...}`, or `"chat_template_kwargs": {"enable_thinking": false}`) |
| Anthropic | `"output_config": {"effort": "low" \| "medium" \| "high"}`, `"thinking": {"type": "disabled"}`, or `"thinking": {"type": "enabled", "budget_tokens": N}` (under 2K = low, under 8K = medium, more = high) |
Without a setting the model uses its own default, **high**. `none` answers at once (fastest); `low` keeps the thinking
short. The levels are instructions the model was trained with, not a hard token limit: on easy questions all three
think briefly, on hard ones `high` thinks longest and is most accurate.
- **A hard thinking budget (opt-in).** `"reasoning_budget_tokens": N` in a request (OpenAI or Anthropic) caps the
thinking at N tokens: when it gets there the server ends it with a short wrap-up line and `</think>`, and the model
answers from there (the engine continues from what it already holds, so nothing is read again). The wrap-up is
part of the thinking the client sees and counts as output tokens. `"reasoning_budget_tokens": N` in
`strata-<model>.json` sets it for every request; a request's own value wins, and `0` means no budget. Off by default;
Anthropic's `"thinking": {"budget_tokens": N}` still only chooses the level, as above.
- **Anthropic requests that don't ask for thinking (opt-in, 0.1.32, #278).** By default a `/v1/messages` request
with no `"thinking"`, effort or budget thinks as the model's template does. `"anthropic_thinking": "on_request"` in
`strata-<model>.json` renders such a request without thinking - Anthropic's own rule, and what Claude Code's short
helper calls (a session title in a few dozen tokens) need; its real turns ask for thinking when it is on there.
- **Streaming.** With `"stream": true` everything arrives as it is made: the thinking, the answer, and tool calls
(the tool's name first, then its arguments piece by piece, like OpenAI and Anthropic do). While the model reads a
long prompt the stream sends keep-alives, so agents do not time out; the server window prints progress every
15 s, and `GET /status` says what it is doing (`reading the prompt`, `answering`, tokens so far). Closing the
connection or pressing stop in your app really stops the model, so the next request starts at once.
- **Chat apps.** Any app with an "OpenAI-compatible" provider works: base URL `http://127.0.0.1:8080/v1`, any API key.
- **OpenCode** (#543). A starting point for `opencode.jsonc` (in your project, or `~/.config/opencode/`); the field
names are OpenCode's, so check its config docs if your version differs:
```jsonc
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"strata": {
"npm": "@ai-sdk/openai-compatible",
"name": "Strata (local)",
"options": { "baseURL": "http://127.0.0.1:8080/v1", "apiKey": "none" }, // or your api_key
"models": {
"strata": {
"name": "Qwen3.8-Flash-Next (Strata)",
// context: what you chose in setup; output: what one reply may use (prompt + output must fit)
"limit": { "context": 262144, "output": 32768 },
"options": { "reasoningEffort": "high" }, // sent as reasoning_effort
"variants": { // switch between them in OpenCode
"low": { "reasoningEffort": "low" },
"medium": { "reasoningEffort": "medium" },
"none": { "reasoningEffort": "none" }
}
}
}
}
},
"model": "strata/strata"
}
```
Set `limit.context` to the context you chose in setup: OpenCode compacts the conversation before it gets there.
Keep `limit.output` well under it: a request whose prompt plus `max_tokens` runs past the context is refused (see
**Context** below), or add `"fit_max_tokens": true` to `strata-<model>.json`. For a hard cap on the thinking, add
`"reasoning_budget_tokens": N` to `strata-<model>.json` (see above).
- **Claude Code** (Strata 0.1.17 or newer): set `ANTHROPIC_BASE_URL=http://127.0.0.1:8080` and
`ANTHROPIC_MODEL` to a Claude model name it knows (it refuses names it doesn't; Strata ignores the name), plus any
`ANTHROPIC_AUTH_TOKEN` (or your `api_key`, if you set one).
- **Context.** Chosen in setup (8K-262K). Requests longer than that are refused, never silently cut. A request whose
`max_tokens` would run past the context is refused too (400); agents that always ask for their full output cap
can instead get it shortened to the room left: add `"fit_max_tokens": true` to `strata-<model>.json` (or pass
`--fit-max-tokens` to `serve/server.py`). A prompt that leaves no room at all is still refused.
- **Model aliases** (0.1.32). `"aliases": ["qwen", "local-model"]` in `strata-<model>.json` lists the model under
those names too in `/v1/models` (each with its own `id`, and in the model's `aliases`), like llama-server's
`--alias`; a request naming one is answered under that name. Any other name is still served, as before.
- **From other devices on your network.** The server listens on your PC only (`127.0.0.1`) unless you say otherwise:
run setup with `START-HERE.bat --setup --host 0.0.0.0 --api-key some-long-secret` (or add `"host": "0.0.0.0"` and
`"api_key": "..."` to `strata-<model>.json`). The server window then prints this PC's addresses
(`from other devices: http://192.168.x.x:8080/`); open that on the other device, or use `.../v1` as an API base URL.
On Windows the firewall blocks it until you allow it: accept its prompt for Python (private networks), or run
`New-NetFirewallRule -DisplayName "Strata 8080" -Direction Inbound -Protocol TCP -LocalPort 8080 -Action Allow -Profile Private`
in an admin PowerShell, and make sure the network is set to Private.
- **From the internet.** Put a tunnel in front of it, for example [cloudflared](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/):
`cloudflared tunnel --url http://127.0.0.1:8080`. **Set a key first**, or anyone with the link can use your PC:
add `"api_key": "some-long-secret"` to `strata-<model>.json` (or set the `STRATA_API_KEY` environment variable);
clients then send it as their API key. Streamed answers carry `X-Accel-Buffering: no`, so nginx-style proxies pass
each token on at once. The web app's settings and MCP tools only answer Strata's own page: when you open it through
a proxy or tunnel whose address differs, add that address, e.g. `"trusted_origins": ["https://strata.example.com"]`.
With the key set, any `Host` name reaches the server (see Host names below).
- **From web apps in a browser (CORS).** Off by default. `"cors_origins": ["https://chat.example.com"]` lets pages of
those origins call `/v1/*` from the browser (Open WebUI's direct connections, browser extensions); `["*"]` lets any
page do it - only sensible with an API key. It never opens `/settings`, `/unload` or the MCP tools.
- **Host names (DNS rebinding).** A web page of another site can point its own name at `127.0.0.1` and then reach
this server as if it were its own, so without an API key the server answers only requests whose `Host` is a name
it knows (with a key the check is off: such a page cannot send the key, and tunnels and proxies that pass their
own name on keep working):
`localhost` (and `*.localhost`), any IP address (`127.0.0.1`, `[::1]`, `192.168.x.x`, ...), the address it
listens on and, when it listens beyond this PC (`0.0.0.0` or a LAN address), this PC's name (`mypc`, `mypc.local`)
and `host.docker.internal`; any port. Others get **403** naming the setting, and the server window prints one line
for each. Reach it under another name (a reverse proxy that keeps the name, a tunnel, a DNS name on your network,
another container's name for it)? Add the name: `"allowed_hosts": ["strata.example.com"]` in
`strata-<model>.json` or `STRATA_ALLOWED_HOSTS=strata.example.com` (comma-separated); `".example.com"` allows that
name and every name below it, and `["*"]` turns the check off (so does setting `api_key`). The hosts of
`trusted_origins` count as allowed. Requests without a `Host` header (HTTP/1.0 clients) pass.
- **Web pages without an API key.** Without `api_key`, a `POST` to `/v1/*` that carries an `Origin` header (a
browser page sent it) is answered only for Strata's own page, pages on `localhost` or an allowed host name (any
port), the origins in `trusted_origins` or `cors_origins`, and browser extensions and desktop apps
(`chrome-extension://`, `moz-extension://`, `app://`: no web site can send those), and only with a JSON body; any
other page, and `Origin: null`, gets **403**. Clients that send no `Origin` (curl, the OpenAI and Anthropic SDKs,
other servers) are not affected. With
an API key, the key decides. `POST /unload` and `POST /load` take `Content-Type: application/json` from Strata's
own page (or no `Origin`), like `/settings`.
**Conversation cache.** A request that continues a chat reads only the part after what the engine already holds: the
live session, or one of the checkpoints it keeps in RAM (up to 6, ~118 MB each, taken at the start of each new
assistant turn and every 16K prompt tokens). A checkpoint is used only when the prompt starts with exactly its tokens
and pictures. The oldest checkpoint - in practice the end of the system prompt, which every chat of the same client
shares - is kept for good while the rest rotates by least recent use, so a NEW chat that shares that prefix starts
reading after it instead of from token 0. A prompt read from the start is also checkpointed at the end of its system
prompt when that is 2,048 tokens or more (engine 0.1.20; PR #62 + #65), so that root exists for agent clients with long
system prompts and tool lists. Engine options: `--prompt-cache N` (0 = off), `--prompt-cache-every N`,
`--prompt-cache-root N` (0 = no system-prompt checkpoint), `--turn-token ID`.
**Multiple conversations (opt-in).** Add `--conversation-cache-mib 8192
--conversation-cache-slots 4` to the engine arguments to park up to four conversations
in a bounded 8 GiB host-RAM cache. This preserves controller/worker histories when
their requests alternate; it does not execute requests concurrently. No client session
ID is required: only exact token/image prefixes with matching steering mode are reused.
The default budget is 0 (disabled); `--prompt-cache 0` also disables parking.
The initial shared-core integration supports a single session GPU: combining
enabled parking with `--layer-split` is rejected before model loading. Ordinary
upstream layer-split checkpoints remain available with parking disabled. FP16,
INT8, Q4_0 and identity-layout K8V4 snapshots are supported; the K8V4 draft ring
remains INT8, as in upstream. Windows/HIP and multi-GPU runtime coverage must be
reported separately from Linux/CUDA evidence.
Snapshots contain running state, checkpoints, used K/V pages, and draft-layer K/V.
They add host RAM, not another model or VRAM allocation. The byte budget also counts
an incoming snapshot during a switch. After a restore, unchanged K/V pages can be
retained for the next parking operation; growth appends storage without copying
the existing pages. Rewinds refresh the affected pages, and running state and
checkpoints are captured again. Retained active K/V counts against the same byte
budget and is discarded before evicting parked entries under memory pressure.
If reserving space for growth would evict another conversation, parking uses a
full capture instead.
Oldest parked entries are evicted first.
Oversized snapshots or host allocation failures fall back to ordinary prompt processing.
`--conversation-cache-min-free-mib N` (default 2560) additionally requires that
physical-RAM headroom remain available: the engine checks before allocation and
again after capture. Unknown telemetry or insufficient RAM skips parking. Windows
uses `GlobalMemoryStatusEx`, Linux uses `MemAvailable`; these are host-level samples,
not a reservation or enforcement of container/job memory limits. An 8 GiB budget
is a cap, not a recommendation for every machine.
The shared snapshot core validates all layers and checkpoints before applying any
state. Invalid entries are discarded; transfer/synchronization failure is fatal
rather than permission to continue with partial state. Indexer spare keys and the
moving spare row are preserved, including checkpoint rewinds.
The engine log reports parking, restoration, bytes, evictions, individual snapshot
sizes and K/V bytes reused during capture. `STRATA_SNAPSHOT_FULL_CAPTURE=1` disables
retention for diagnostic comparisons. Snapshots are not
persisted across restarts.
**Current limits (v1):** one request at a time, and one conversation cached at a time (switching between two chats
re-reads the other one unless the opt-in cache above is enabled); images only when set up with them (below); no video. **Temperature / top_p / top_k / min_p /
seed** are honored per request (OpenAI and Anthropic fields); with the default adaptive expert tier a sampled result
is not reproducible run to run - for seed-reproducible output add `--adapt-every 100000` (static residency) to the
engine arguments. The run config's optional `sampling` block sets the defaults for requests that leave the fields out
(`"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20}`); a request's own fields always win, and with no
block at all a request without sampling keys decodes greedy. The penalties (`presence_penalty`, `frequency_penalty`,
`repetition_penalty`, with `penalty_last_n` capping how many recent tokens they count over, default 64 when any
penalty is set) ride the same path; they count the tokens the request has consumed, so a repetition penalty
suppresses what the model itself just said, not the prompt alone. Since engine 0.1.19 they apply to every token
the speculative decoding checks at once, exactly as if it decoded one token at a time (before, only the first of
each batch got them). That makes requests with penalties 1-11% slower than in 0.1.18: the draft layer guesses
without penalties, so more of its guesses are now rejected. Requests without penalties are unchanged. `top_k` keeps at most 64 candidates: `0` ("off") or anything above 64 uses all 64.
---
## Tools from MCP servers
The chat page can give the model tools from [MCP](https://modelcontextprotocol.io) servers, as LM Studio and Claude
Desktop do: reading your files, fetching web pages, searching, anything an MCP server offers. List the servers in
`strata-<model>.json` under `"mcp_servers"` - the same shape as Claude Desktop's `mcpServers` block, which you can
also paste as it is (key `"mcpServers"`):
```json
"mcp_servers": {
"files": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "C:\\Users\\me\\Documents\\notes"]},
"search": {"url": "http://127.0.0.1:3000/mcp", "headers": {"Authorization": "Bearer ..."}}
},
"mcp": {"timeout_s": 60, "max_result_chars": 20000, "max_rounds": 8}
```
Or keep them in their own file and start the server with `--mcp-config path\to\claude_desktop_config.json` (a file
with an `mcpServers` block; add it to the `serve/server.py` line of your run script). Restart Strata after a change.
- **A program** (`command`, `args`, optional `env` and `cwd`) is started by Strata and spoken to over its
stdin/stdout; `npx`, `uvx`, `python` and friends are found on `PATH` as usual (Node.js is needed for `npx`
servers). **An address** (`url`, optional `headers`) uses MCP's Streamable HTTP transport (the older SSE-only
transport is not supported). `"disabled": true` leaves an entry out.
- The servers start with Strata, in the background; the server window says what each one offers
(`MCP server 'files': 14 tools (...)`), or why it did not start - its tools are then left out and the chat works
without them. The Monitor tab lists them, and the Sampling drawer has **Use tools from MCP servers** (on by
default). A server that stops later is started again at its next call.
- In the chat each call shows as a small block (tool, arguments, result); the model reads the result and goes on,
up to `max_rounds` calls in a row per answer. A tool that fails or takes longer than `timeout_s` (default 60 s)
gives the model an `error: ...` result instead of ending the chat. Results longer than `max_result_chars`
(default 20,000 characters) are cut, with a note, before the model reads them. Stop stops a running tool too.
- Only the chat page uses them. API clients (omp, Claude Code, OpenAI and Anthropic SDKs) see the API exactly as
before and keep their own tools; a request to `/v1/chat/completions` opts in with `"strata_mcp": true` (it then
gets `strata_mcp` tool events in the stream).
**Security.** MCP tools run on your PC with your user's rights, and **the model decides when to call them** - also
because of what it reads (a web page or a file can contain instructions). Give a filesystem server only the folders
it needs, prefer read-only tools, and don't add servers you don't trust. The tools can only be used from the chat
page itself (a request with another site's Origin or without a JSON content type is refused); if Strata is reachable
from other devices, set an API key.
**Context extension past 262K (rope scaling, EXPERIMENTAL, off unless you pick it).** The model was trained on
262,144 positions (rotary base 1e7). Rope scaling rescales the rotation angles so that longer contexts stay usable,
with llama.cpp's types and flag names. `linear` is Position Interpolation: every angle is shrunk by the factor.
`yarn` keeps the high-frequency angles, interpolates the low-frequency ones, and adds the magnitude correction
that keeps the attention temperature where training put it. **Without the flags nothing changes:** an unscaled
run computes exactly what it did before the feature existed, bit for bit. Scaled contexts need proportionally
more VRAM/RAM for the KV cache and the rope tables (~13 KB and ~0.26 KB per token).
What was measured (contributors' runs, RTX 5080 + IQ3_S, native path, in PR #84): per-position perplexity on the
same tokens, with only the scaling flag changed. At 293K tokens (1.12x the trained length), `yarn` with factor 2
lowered the NLL by 0.18 nats against both `none` and `linear` 2 (2.04 vs 2.22 / 2.22). That is 3-4x the path noise
measured at the same length. `linear` 2 was indistinguishable from `none`. At 2.7K and 32K no arm separated from the
noise. Needle tests do not tell the arms apart: the unscaled model also finds a needle at 413K. Long real-document
Q&A worked with `yarn` 2 at 421K and `yarn` 4 at 714K (8/8 each), and a 1M-token `yarn` 4 run read end to end.
Taken together, use **yarn**. It is still experimental: the numbers come from one machine and one quant.
- setup: `START-HERE.bat --setup --context 393216` asks nothing extra - it picks the method (yarn; one
question when run interactively) and derives the factor from the final context for you (final context /
262,144, at least 1: 1.5 at 393K, 2 at 512K, 1 inside the trained range; `--rope-scaling`/`--rope-scale`
override; an explicit `--rope-scale` is kept as given even when it is too small for the context actually
served, so check it if you set one). An explicit `--rope-scaling none` for a context past
262,144 is refused: the setup will not configure a run with the stock angles past the trained range. If
the RAM check reduces a chosen 384K/512K back inside the trained range, an omitted method adds no
scaling, and an explicitly chosen one stays at factor 1 - the trained angles, no expansion (not a
switch for rope as a whole: explicitly supplied rope settings keep their behavior).
- engine: `--rope-scaling none|linear|yarn`, `--rope-scale F`, and the raw ggml knobs `--rope-freq-base`,
`--rope-freq-scale`, `--yarn-orig-ctx` (default 262,144), `--yarn-ext-factor`, `--yarn-attn-factor`,
`--yarn-beta-fast` (32), `--yarn-beta-slow` (1). The model file's `rope.scaling.*` keys, when a
fine-tune ships them, are the defaults the flags override.
The scaling is fixed for the whole run - the engine stores keys in its cache after rotating them, so one
cache must never mix two scalings, and there is no per-request form. Within the trained 262,144 a scaled
run is a slightly different model: the rescaled angles, and `yarn`'s magnitude correction, apply at every
position, not only past the trained end. That is why the setup turns scaling on only for a context past
262,144. Pictures read the same scaled table (their (t, h, w) positions feed it). That should work, but it is
unmeasured: all the runs above are text.
---
## Manage Strata from your AI assistant (MCP server)
`tools/strata_mcp.py` is an MCP server for Claude Code, Claude Desktop, Cursor, VS Code, Codex and other assistants.
Once it is added, you can ask your assistant "install Strata for this PC", "start Strata" or "is Strata running?".
In Claude Code, add it with:
```bash
claude mcp add strata -- python C:\Users\you\Strata\tools\strata_mcp.py
```
It has eight tools: status (the running model, what is installed, the hardware, a recommended size), the model
list, install, start, stop, logs, a speed test, and connection settings for other apps.
Install runs `setup.py` with `--yes` in the background. Before it downloads anything, it shows the plan and waits
for your OK. Start and stop work like the run scripts and the server's own unload. The MCP server only ends
processes it started itself. It uses only Python's standard library, so it works before `.venv` exists.
The config snippets for every client, the tool arguments and the safety rules are in
[docs/MCP_SERVER.md](MCP_SERVER.md). This is the opposite direction from
[Tools from MCP servers](#tools-from-mcp-servers) above, where the Strata model calls *your* MCP tools.
---
## Images (vision)
The model has a vision encoder: [`mmproj-Qwen3.8-Flash-Next-BF16.gguf`](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF)
(0.9 GB, a 27-layer ViT plus the projector into the language model). It is **optional**: say yes when the setup asks
"Images?", or run it again with `--vision gpu` (or `--vision cpu`). The setup downloads the encoder, builds a small
helper (`strata-vision`, from llama.cpp's `mtmd` library) and adds it to your start script. Nothing else changes.
| Encoder on | Time per picture | Cost |
| --- | --- | --- |
| **GPU** (recommended) | **0.1-0.5 s** (up to 1,024 image tokens) | ~1.4 GB of VRAM is kept free for it, so the expert cache is smaller: text output is a few % slower (table below) |
| CPU | 10-30 s (pictures are scaled down to ~300 image tokens) | nothing on the GPU |
A picture becomes up to 1,024 tokens of the context (a 640x480 photo: 300). The same picture sent again, as chat apps
do on every turn, is encoded only once.
**A spare GPU for the encoder (0.1.33, #408):** with a card the engine doesn't use, add `"cuda_device": 2` (numbered
like `nvidia-smi`) to the `"vision"` section of `strata-<model>.json`: the encoder then runs on that card alone. Lower
`--vram-reserve-mib` in `"args"` to 700 as well, so the engine's cards keep that VRAM for the expert cache. The
encoder's card needs code in the ready-made encoder (RTX 20/30/40/50).
### Sending a picture
**Terminal chat:** type `/image <path to a picture>`, press Enter, then type your question.
```
you> /image C:\Users\me\Pictures\receipt.jpg
(picture attached: receipt.jpg - now type your question)
you> What is the total on this receipt?
```
**OpenAI API** (an `image_url` part: a `data:` URL, an `http(s)://` URL or a local file path):
```python
import base64
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="none")
img = base64.b64encode(open("photo.jpg", "rb").read()).decode()
r = client.chat.completions.create(model="strata", messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img}"}},
{"type": "text", "text": "What is in this picture?"}]}])
print(r.choices[0].message.content)
```
**Anthropic API:** an `image` block with a `base64` (or `url`) source, as usual.
JPEG, PNG, BMP, GIF, WebP, TIFF and AVIF work (the last ones are converted to PNG first; agents such as omp send
WebP). Chat apps with image upload work the same way.
### Speed with images on (4K context, measured)
| Model | Prompt tok/s, images off | Prompt tok/s, images on | Output tok/s, images off | Output tok/s, images on |
| --- | ---: | ---: | ---: | ---: |
| **Q2_0** | 541 | 531 | 90.7 | 86.9 |
| **IQ2_XS** | 468 | 458 | 77.1 | 74.0 |
| **IQ3_XXS** | 411 | 401 | 64.7 | 59.5 |
"Off" is the published setup; "on" runs with the encoder loaded on the GPU and its VRAM kept free, so ~1,000 fewer
experts fit in VRAM and the CPU computes a few more per token: 2-8% slower. For text, turning images on changes
nothing else (the output is bit-identical when the VRAM is the same).
A question about a picture (a 640x480 newspaper page = 300 image tokens, a 328-token prompt; the whole request, encoder
on the GPU, measured through the API):
| Model | Answer ("MEN WALK ON MOON") | Picture encoded | Prompt | Output |
| --- | ---: | ---: | ---: | ---: |
| **Q2_0** | 1.9 s | ~0.1 s | 207 tok/s | 65-73 tok/s |
| **IQ2_XS** | 2.3 s | ~0.1 s | 173 tok/s | 50-60 tok/s |
| **IQ3_XXS** | 2.7 s | ~0.1 s | 144 tok/s | 41-50 tok/s |
(Short prompts run below the 4K prompt speed: a 2,048-token chunk is where the prompt path is efficient. Short answers
run below the long-output speed: the first rounds have no draft yet.)
**How it works inside:** the encoder turns the picture into rows of the same width as the model's word embeddings;
Strata puts them where the prompt has `<|image_pad|>` tokens and gives each one its 2-D position (row and column in the
picture; the model uses interleaved M-RoPE). Answers match llama.cpp's multimodal implementation token for token on our
test images.
---
## Experimental speed projection (EXPERIMENTAL, off by default)
**This is an experiment, not a finished feature.** It ships with Strata but stays off unless you turn it on.
A 480 KB control vector for Qwen3.8-Flash-Next (`data/experimental-speed-projection/`, see its README). After each
of layers 4-44 the engine removes one direction from every hyper-connection stream of the residual: `h -= (h . v) v`,
one unit vector `v` per layer, exactly as llama.cpp does with the package's `--cvec-mode project` patches.
**What it changes.** The vector's own package describes it as a **refusal-direction projection**: with it the model
declines far fewer requests (it reports 1 of 50 vs 50 of 50 on its test set), and removing refusals removes a safety
behaviour - you are responsible for what the model writes with it on. It also shifts ordinary answers a little
(measured below). It is not an optimization in the engine: on the same text it costs 0.2-0.4% per token. What a
chat's tokens/s does with it on depends on the text the model writes (length, repetition, how well the drafts land),
so measure it on your own prompts; the Monitor marks every request ESP or stock.
**Turning it on (at setup).** `START-HERE.bat --setup` asks "Turn on the experimental speed projection?" (default:
no), or pass `--experimental-speed-projection on` (`off`, or a path to another vector GGUF). Only for the original
Qwen3.8-Flash-Next, not Swift 1.5. It writes these engine flags (llama.cpp's) into `strata-<model>.json`:
```
--control-vector-scaled <Strata>\data\experimental-speed-projection\Qwen3.8-Flash-Next-experimental-speed-projection.gguf:1.0
--control-vector-layer-range 4 44 --cvec-mode project --cvec-dir per-layer
```
The engine log then says `control vector mode = project, dir = per-layer, layers 4..44 (41 steered)`, and the web
app's About tab lists it. (`--cvec-mode add` is llama.cpp's stock additive mode, for additive vectors.)
**Per request.** A loaded vector is on for every request unless it says otherwise: the web app's Sampling drawer has
a switch, and the API takes `"experimental_speed_projection": false` in the request body (OpenAI and Anthropic; a
config default goes in `"sampling": {"experimental_speed_projection": false}`). Switching drops the conversation
cache once, since the model state was computed the other way. Switched off, the output is token-for-token the stock
model's.
**Measured here** (Q2_0, fixed experts, 2,557 teacher-forced tokens of code, a document and a chat): the top-1 token
changes at 10% of positions, mean KL from the stock model 0.063 nats (max 4.1), perplexity +15% on code, +2.3% on
the document, +0.4% on the chat. Details: `bench/results/2026-09-27-esp/`.
---
## Troubleshooting
| Symptom | What to do |
| --- | --- |
| `the NVIDIA driver is too old` | Update the driver (NVIDIA App or nvidia.com/drivers), restart, run `START-HERE.bat` again. |
| Python or the build tools could not be installed | Install what it names (links are printed), then run it again. Everything already done is kept. |
| `port 8080 is already in use` | Strata is already running (look for its window), or another program uses the port: `START-HERE.bat --port 8081`. |
| `cudaHostRegister ... out of memory` in the log | Normal on Windows: the engine pins the experts in per-layer slices instead. Only a problem if the load then fails. |
| `ExpertCache: cudaMalloc(...) failed: out of memory` although VRAM is free | Windows' page file is off or tiny: every allocation on the graphics card is also charged to Windows' commit (RAM + page file). Set the page file to "System managed" (System > About > Advanced system settings > Performance > Advanced > Virtual memory) and restart. Since 0.1.19 the engine retries with a smaller cache instead of stopping, and setup warns about a page file under 4 GB (issue #60). |
| The first start takes minutes | It is reading 34-55 GB into RAM; the second start is faster while the files are in the OS cache. |
| The PC freezes for a few minutes at the start | Normal, most of all the first time (the server window says when it happens): the engine loads the experts into RAM, pins part of it for the GPU and sizes the expert cache. Wait; don't close the window. Still frozen after 10 minutes: restart the PC, close other programs, try again, or pick a smaller size. |
| `the engine stopped unexpectedly (exit code ...)` | The engine process ended mid-answer - usually out of RAM (Linux ends the biggest program: `sudo dmesg \| grep -i -E 'killed process\|out of memory'`). The next request starts it again by itself. If it repeats: close other programs or pick a smaller size. The server also warns at start when the model's experts leave less than ~6 GB of RAM for everything else. |
| Slow output, disk light busy | Not enough free RAM: close other programs, or choose Q2_0 / IQ2_XS. |
| `prompt ... exceeds the context` | The request is longer than the context you chose: run setup again with a bigger `--context`. |
| `the setup refuses --rope-scaling none for a past-trained context` | A context past the trained 262,144 needs the rotary angles rescaled (experimental rope scaling), and the setup will not configure one with the stock angles there. Let it pick (`START-HERE.bat --setup --context 393216` adds yarn and a covering factor), or pass `--rope-scaling linear` or `yarn` yourself. |
| Slower than the tables | The monitor plugged into the GPU and other GPU programs take VRAM from the expert cache; RAM running below its rated speed (enable EXPO/XMP in the BIOS) slows the CPU half. |
| `this server was started without the vision encoder` | The model was set up for text only: run setup again with `--vision gpu`. |
| A picture is refused or `cannot read the image` | The file is not a picture Pillow can open (JPEG, PNG, WebP, GIF, BMP, TIFF, AVIF work). |
| Pictures are slow (10-30 s) | The encoder runs on the CPU: run setup again with `--vision gpu` (needs ~1.4 GB of VRAM). |
| A request never finishes: "reading the prompt", GPU "100%" at low power | The GPU ran out of VRAM (engines before 0.1.9 could end with ~30 MiB free at large contexts). Run `START-HERE.bat` once to get engine 0.1.9 or newer; the log then says `... MiB of VRAM free with everything loaded` (a few hundred) and names the `--vram-reserve-mib` to add if it is low. |
| Generation stops mid-answer, GPU "100%", one CPU core busy | Fixed in engine 0.1.12 (issue #29, a race in the CPU expert pool on big-VRAM cards). Since then a request that stops moving ends with an error instead of hanging (after 2 minutes; 1 minute from 0.1.13): the log says `no progress for ... s ... (issue #29)` with where it stopped, and the next request starts the engine again. If you see that line, please open an issue with it. Engine 0.1.13 adds a stall report under it (what every expert-pool thread and the GPU handshake were doing, memory and page faults) and, on Windows, a `strata-stall-<pid>.dmp` file with every thread's stack: attach both. (`STRATA_WATCHDOG_S` sets the time in seconds; 0 turns it off.) Engine 0.1.14 fixes the stall those reports found (issue #31: with the IQ packs the host could wait forever inside the NVIDIA driver while copying experts in a verify window; the experts are now copied by a GPU kernel, `--pcie-mode dma` restores the old way). |
| `the engine said nothing for ... s during the request` or `... did not finish the request after it was stopped (STOP)` | Issue #481: the engine and the server lost step (the engine waits for its next command, the server for the request's end; GPU at 0 %, nothing in the log). The server ends the engine after 300 s without a line from it during a request (while a prompt is read: each chunk may take three times the previous one's time, the first one up to its tokens at 50 tok/s more), the request ends with an error and the next request starts the engine again. `"engine_silence_s": 600` in `strata-<model>.json` sets the time (0 = wait forever, as before). If you see it, please add the end of the engine log to #481. |
| `out of memory: cudaFuncSetAttribute` in the log (IQ3_XXS, long prompt) | Fixed in engine 0.1.15: CUDA loaded a kernel's code when it was first needed, and mid-prompt there was no VRAM left for it. Run `START-HERE.bat` (Windows) or `./setup.sh` (Linux) once to update. |
| Anything else | The engine log is `strata-<model>.log` in this folder. |
---
## How it works
<p align="center"><img src="paper/tiers.svg" width="760" alt="memory tiers"></p>
- **GPU (VRAM):** attention and DeltaNet mixers, the gated-residual weights, routers, shared experts, output head, the MTP
draft layer, the KV cache (from 64K: only its most-read part, the rest streams from RAM), and an **expert cache** that fills the rest of VRAM with the most-used experts (it adapts to
the conversation while you chat).
- **RAM:** all 24,576 experts, pinned. The CPU computes the experts that are not on the GPU **in place**, at the same time
as the GPU works on the cached ones (AVX-512 / AVX2 kernels, ggml's for the i-quants).
- **SSD:** the 28.8 GB n-gram table, read a few rows per token through the OS cache.
- **Speculation:** the model's own MTP layer drafts up to 3 tokens; one pass over all 48 layers checks them. 2.4-3.2
tokens per pass on average. When the reply repeats the context (code edits, quoted text), **prompt lookup** (engine
0.1.7) drafts up to 5 tokens from the earlier copy, but only where its measured acceptance and cost say it pays:
code edits 6-11% faster, other text unchanged. The drafts are checked like the MTP's, so the output is the same.
- **Prompts** are processed in 2,048-token chunks with the experts streamed to the GPU over PCIe.
The full story, with measurements, bottlenecks and what comes next: **[docs/paper/Strata-Paper.pdf](paper/Strata-Paper.pdf)**.
---
## Credits and licenses
Strata itself: [MIT](../LICENSE). The model files are not part of it; their licenses apply to them (below).
- Model: [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) by the Qwen team; quantizations:
[ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF).
The Coder: [ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF)
(Apache-2.0 per its card); its support in Strata came from @pjgmobile's PR #54.
Swift 1.5: [ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF)
by UkisAI. Their licenses apply to the weights.
- [llama.cpp / ggml](https://github.com/ggml-org/llama.cpp) (MIT): the i-quant formats, the GPU dot products and
dequantizers transcribed in `src/kernels/cuda/iq_kernels.cu`, the CPU backend linked for the i-quant experts, the
`mtmd` library behind the image encoder (`tools/vision/`), and `gguf-py` used by the tools. See
`third_party/ggml/LICENSE`.
- Ideas from [Splash](https://github.com/incoai/splash), [ninfer](https://github.com/Neroued/ninfer) and
[HyperQwen](https://github.com/syv-ai/HyperQwen); references in the paper.
- The web app's font: [Outfit](https://github.com/Outfitio/Outfit-Fonts) (SIL Open Font License 1.1, see
`serve/web/fonts/OFL.txt`). Its Monitor tab started from @code-martin's dashboard idea (PR #22).
- The experimental speed projection's vector (`data/experimental-speed-projection/`): Qwen Community License 1.0,
made from the model's activations (see its README).
### Start the text API without occupying the GPU
`serve/server.py --engine strata --config strata-<model>.json --lazy` (or `"lazy_load": true` in that
config) starts the lightweight HTTP API without spawning the native engine. The first generation request
loads it through the existing reload path, including `before_load` and `min_free_vram_mib`. Eager startup
remains the default. This option is text-only: a vision configuration with lazy startup is rejected explicitly.
`POST /v1/load` and `/v1/unload` are JSON control aliases for integrations, accepting `{}` or
`{"model":"<configured model>"}` and returning model status. They require the configured API key,
`application/json`, and no foreign browser Origin. They return **409** while a request is active or queued,
and **404** for an unknown model. Existing `/load` and `/unload` behavior is preserved. `/api/health` aliases
`/health`; `/v1/status` exposes `loaded` and `auto_load`. The unloaded model remains discoverable.
Unloading and shutdown close the native engine's stdin after sending `QUIT`, allowing Windows' detached
stdin reader to see EOF. Cleanup waits for process exit before releasing handles; if forced shutdown still
times out, the server keeps ownership and reports an error rather than claiming the model was unloaded.
### JSON response formats
`POST /v1/chat/completions` accepts `response_format: {"type":"json_object"}` or
`{"type":"json_schema","json_schema":{"name":"answer","strict":true,"schema":{"type":"object","properties":{"answer":{"type":"integer"}},"required":["answer"],"additionalProperties":false}}}`.
The schema must describe an object at its root. Local `#` references work; remote references are refused.
`json_schema` is checked with the Python package `jsonschema` when it is installed (`python -m pip install
"jsonschema>=4.23,<5"`; setup does not add it); without it the answer is only checked to be one JSON object, and the
server says so once.
This is **schema prompting followed by server validation**, not grammar-constrained decoding. One generation
is made per request, with no hidden retry. Successful responses contain a validated JSON object. Malformed JSON,
duplicate keys, non-finite numbers, schema violations and incomplete generations return **502** with
`error.code: structured_output_failed`; invalid request schemas return **400**. JSON formats combined with
tools/MCP are refused explicitly. Without `response_format`, ordinary text and tool behavior stays the same.
Structured SSE buffers the answer while sending keep-alive comments. It emits content only after validation,
then usage/timings and `[DONE]`; failures emit an SSE error and `[DONE]` without invalid content deltas.
`/v1/status.structured_output` advertises the formats, validation method and buffered streaming behavior.
### API request monitor
Off by default, since it keeps prompts and answers in memory: turn it on with `"api_monitor": true` in
`strata-<model>.json` (or `serve/server.py --api-monitor`); otherwise nothing is recorded and the two endpoints below
answer 404.
Open `/api-monitor` to inspect API traffic without opening a chat. It shows the model state, safe
load/unload controls, active/queued requests, original request bodies, output, separate reasoning and
non-stream response bodies. Total wall-clock includes FIFO waits and automatic loading; load, queue,
first-token, prompt/output tokens and engine decode timing are shown separately.
`GET /api/requests` returns compact summaries; `GET /api/requests?id=<id>` returns one retained request.
Both use the existing API-key check. The monitor retains the newest **100 requests in memory** until restart,
with **262,144 characters per input/output/reasoning/response field** and visible truncation flags. The actual API
responses are unaffected. Headers are not recorded, and the monitor key is kept in this tab's session storage.
Treat request history as sensitive input/output when exposing Strata on a network: set an API key as above.
The page uses relative URLs and works through the existing host binding or a reverse proxy.
|