Strands Decider 2B: int4 weights for WebGPU

StrandsAgents/strands-decider-2B-hobson-v19 with its LoRA merged into Qwen/Qwen3.5-2B-Base (revision b1485b2), then quantized for a hand-written WebGPU engine that runs the model entirely in the browser.

  • engine-weights/: int4 symmetric round-to-nearest, block 32, MatMulNBits layout, fused projections (1.06 GB, 2 shards). manifest.json lists every tensor's offset and shape.
  • model/: tokenizer, pointer head and decider config, unchanged from the original model.

Code, demo and conversion scripts: https://github.com/alxnahas/strands-decider-web. Live demo: https://alxnahas.github.io/strands-decider-web/

On a 27-item benchmark the browser engine picks the same top answer as the official PyTorch implementation on 27/27. Mean absolute probability error is 0.019, all of it from int4 quantization.

License: Apache-2.0, the same as the source models.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alxnahas/strands-decider-2B-webgpu

Finetuned
(98)
this model