Instructions to use Jakevin/kev-d3-paper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jakevin/kev-d3-paper with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Jakevin/kev-d3-paper --local-dir kev-d3-paper
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Where Can Qwen3.5 Decision Models Afford Low Precision? A Sub-3-Bit Sensitivity Map, Runtime-Priced Byte Accounting, and an Equal-Size Check of Hand Recipes (working draft)
Working draft v9.1 (2026-10-07), working paper, not peer reviewed. Numbers may still change as remaining runs finish.
๐ kev-d3-paper-draft-v9.1.pdf (current)
Per-module sub-3-bit sensitivity of two Qwen3.5 hybrid decision models (Gated DeltaNet linear attention interleaved 3:1 with full attention), measured as option-distribution KL under binary / ternary / 3-bit / 4-bit GPTQ, with a measured-KL allocation (a multiple-choice knapsack priced in the target runtime's storage format; called HyBitQ in earlier drafts) that turns the measurements into bit allocations.
Related release: Jakevin/kev-4b-ternary-mlx (revision v2.0 uses this allocation; v1.0 is the earlier hand-designed recipe).
| Kev-4B | decision-v7 retention | held-out retention | size |
|---|---|---|---|
| v1.0 (hand recipe + rank-16 adapter) | 96.8% | 86.5% | 1.679 GB |
| v2.0 (measured allocation, no adapter) | 99.2% | 94.5% | 1.678 GB |
Earlier versions
- kev-d3-paper-draft-v8.pdf: superseded draft v8 (2026-10-07).
- kev-d3-paper-draft-v6.pdf: superseded draft v6 (2026-10-07).
- kev-d3-paper-draft-v2.pdf: superseded draft v2 (2026-10-07, earlier title "Where Can a Hybrid Linear-Attention Model Afford One Bit?", still contained [TBD] sections). Kept so existing links keep working.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support