Where Can Qwen3.5 Decision Models Afford Low Precision? A Sub-3-Bit Sensitivity Map, Runtime-Priced Byte Accounting, and an Equal-Size Check of Hand Recipes (working draft)

Working draft v9.1 (2026-10-07), working paper, not peer reviewed. Numbers may still change as remaining runs finish.

๐Ÿ“„ kev-d3-paper-draft-v9.1.pdf (current)

Per-module sub-3-bit sensitivity of two Qwen3.5 hybrid decision models (Gated DeltaNet linear attention interleaved 3:1 with full attention), measured as option-distribution KL under binary / ternary / 3-bit / 4-bit GPTQ, with a measured-KL allocation (a multiple-choice knapsack priced in the target runtime's storage format; called HyBitQ in earlier drafts) that turns the measurements into bit allocations.

Related release: Jakevin/kev-4b-ternary-mlx (revision v2.0 uses this allocation; v1.0 is the earlier hand-designed recipe).

Kev-4B decision-v7 retention held-out retention size
v1.0 (hand recipe + rank-16 adapter) 96.8% 86.5% 1.679 GB
v2.0 (measured allocation, no adapter) 99.2% 94.5% 1.678 GB

Earlier versions

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support