Strands Decider v21, compressed for WebGPU

An unofficial, compressed build of strands-decider-2B-hobson-v21 by Alex Nahas, for the hand-written WebGPU engine in alxnahas/strands-decider-web. It is not a Strands release. Live demo: https://alxnahas.github.io/strands-decider-web/

The browser downloads about 472 MB (437 MB of weights, a 27 MB embedding bundle, 8 MB of tokenizer and head), against 1,085 MB for the earlier int4 build of v19 that the demo used.

What changed from v21

  • v21's LoRA adapter is merged into Qwen/Qwen3.5-2B-Base (revision b1485b2fa6dfa1287294f269f5fb618e03d52d7c).
  • Decoder layers 14 to 22 (9 of 24) are removed. One linear block, h + W rms_norm(h) + b, takes their place after layer 13. It was trained to match v21 on 1,516 rows built from v21's public training recipes, none of them JevBench items. This is the same block as in alxnahas/strands-decider-1B-v21-pruned.
  • The residual stream is rotated by a fixed Hadamard matrix, then every linear layer is quantized with GPTQ to int4 (group 32, symmetric), calibrated on 256 rows from the same recipes.
  • The embedding keeps the full 248,320-token vocabulary at int3 (group 32). The 32,126 most common rows (by token counts over v21's training prompts, plus the special and byte tokens) ship up front. The engine fetches any other row with an HTTP range request the first time a prompt uses it.
  • Every group scale is stored in 8 bits on a per-row log scale instead of fp16 (at most 0.9% error per scale), and the weights and tokenizer are gzipped. The engine expands them to the fp16 layout its kernels read when it loads.
  • The readout head and its temperatures are v21's, unchanged.

Results

Accuracy was measured through the WebGPU engine itself: the JevBench harness called a local shim that runs each task in headless Chrome. Stock Chrome (portable kernels) and Chrome with WebGPU flags (subgroup-matrix kernels) gave the same answers on all 231 tasks. An MLX simulation of the same quantized weights agrees on 230 of 231, and on-demand embedding rows give bit-identical outputs to the full table (40 prompts in 9 languages).

JevBench public set (231 tasks):

correct answers changed vs v21 v1.6.1 rescore: Capability (Intelligence, Calibration)
v21, bfloat16 (MLX) 176 53.2 (32.3, 74.2)
this build, WebGPU engine 176 11 50.8 (30.4, 71.1)
this build, MLX simulation 175 12 52.2 (30.4, 74.0)
v21 in the earlier demo's format (int4, all 24 layers, MLX) 163 31 48.8 (25.4, 72.3)

The rescore applies the JevBench v1.6.1 rules to these 231 items only (no sealed half), so it is not a board number. It moves by a point or two with very small probability changes (the engine and the simulation differ by a mean KL of 0.0002), mostly through the noul cutoffs at 0.2 and 0.8 and calibration bins: the same 176 answers from stock Chrome rescore to 51.6 (30.6, 72.5). The same items were used to compare builds during the work, so the scores may read a little high.

Other languages, 48 public tasks translated into each language, correct of 48 (v21 on MLX):

en zh ja ko ar hi uk de pl
v21 45 43 43 44 40 31 42 43 39
this build, WebGPU engine 44 45 42 41 39 30 42 42 39

Held-out tests, never used to choose a build (MLX simulation; accuracy in percent, and the gap to v21 in points with a 95% paired-bootstrap interval):

tasks v21 this build gap
OpenBookQA (choice: each question bare, with its supporting fact, with an unrelated fact) 3,000 64.1 63.7 โˆ’0.3 [โˆ’1.3, +0.5]
BoolQ validation (noul) 3,270 86.8 85.9 โˆ’0.9 [โˆ’1.5, โˆ’0.3]

The BoolQ loss is small but its interval excludes zero, so the 176 above probably reads a little high.

GPU time per forward pass in headless Chrome on an Apple M4 Pro: 29 ms at 68 tokens, 161 ms at 512, 646 ms at 2,048 (the earlier v19 format: 46, 259 and 1,041 ms). Stock Chrome, without subgroup-matrix: 37, 201 and 800 ms. GPUs without 32-wide subgroups (most Intel, AMD and mobile GPUs) get subgroup-free kernels; forced on the M4 Pro, they give the same answers on all 231 tasks at 39, 212 and 844 ms.

Files

  • engine-weights/manifest.json: model config and every tensor's shard, offsets, shape and format.
  • engine-weights/weights.0.wire.gz: the int4 layers, norms and the linear block, gzipped. Tensors are stored back to back; each scale tensor is N rows of (max |scale| as fp16, step as uint16), then one byte per group, sign << 7 | k, for |scale| = max 2^(-k step / 4096).
  • engine-weights/embed_bundle.bin: the bundled embedding rows (token ids as uint32, then one row each).
  • engine-weights/embed_rows.bin: every embedding row (836 bytes each: 768 bytes of 3-bit codes, the 4-byte scale header, then 64 scale bytes), read only in ranges.
  • model/: v21's tokenizer (gzipped), pointer head and decider config.

License

Apache-2.0, like v21 and Qwen/Qwen3.5-2B-Base; see LICENSE.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for alxnahas/strands-decider-v21-webgpu

Finetuned
(105)
this model