Instructions to use Jakevin/kev-4b-ternary-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jakevin/kev-4b-ternary-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Jakevin/kev-4b-ternary-mlx --local-dir kev-4b-ternary-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Kev-4B Mixed Ternary/3-bit for MLX (v2.0)
📄 Paper (working draft): the v2.0 bit allocation is described in Where Can Qwen3.5 Decision Models Afford Low Precision? A Sub-3-Bit Sensitivity Map, Runtime-Priced Byte Accounting, and an Equal-Size Check of Hand Recipes, not peer reviewed. The paper page always links the latest draft PDF.
An experimental, standalone compression of Jared Palmer's Kev-4B decision model, published by Jakevin. v2.0 replaces the v1.0 recipe with a sensitivity-guided mixed-precision layout. It has the same file size as v1.0 or slightly smaller, and scores higher than v1.0 on every suite we measured. The full distribution is about 1.71 GB. It includes the compressed backbone, pointer head, configuration, tokenizer and the required Kev runtime source.
| Suite (clean scored questions) | bf16-MLX | v1.0 | v2.0 |
|---|---|---|---|
| decision-v7 development (1,264) | 0.8718 | 0.8441 (96.8%) | 0.8647 (99.2%) |
| transfer-v9, non-MMLU part (766) | 0.8316 | 0.7193 (86.5%) | 0.7859 (94.5%) |
| transfer-v9, MMLU part (280) | 0.6071 | 0.2250 (37.1%) | 0.3964 (65.3%) |
Backbone file (model.safetensors) |
8.412 GB | 1.679 GB | 1.678 GB |
Percentages are accuracy relative to bf16. Read the table with these limits in mind. Decision-v7 is closely related to Kev's training recipe. The MMLU part is small, and the compressed models lose a large share of their general-knowledge accuracy on it. See "Evaluation" below.
This is a decision scorer, not a text-generating chat model. It implements
Kev's existing state / questions -> answers / usage System One API.
You need Apple Silicon and the included custom MLX runtime. Stock
Transformers, stock mlx-lm text generation and llama.cpp cannot load this
artifact.
Earlier version: v1.0 (ternary + 4-bit + CLoQ rank-16 compensation) is
still available with revision='v1.0'. It is identical to the earlier
v0.1.0-experimental tag.
What Changed in v2.0, in Plain Language
v1.0 used one rule for the whole network. Most weight matrices were stored as ternary (each weight is -1, 0 or +1 times a scale). A few attention-related matrices were stored as 4-bit. A small correction adapter (CLoQ) recovered some of the lost accuracy.
v2.0 measures where precision matters before deciding how to store anything:
- Measure sensitivity. We compressed each of the 200 backbone weight matrices one at a time, leaving the others untouched. Each one was tried as ternary, 3-bit and 4-bit. For every attempt we measured how far the model's answer probabilities moved, using KL divergence on 32 calibration questions. That gives a cost table: how much damage each matrix does at each precision.
- Spend the bytes where they help most. Next we picked one precision per matrix to minimise the total predicted damage. The constraint was that the whole layout costs no more bytes than v1.0. This is a multiple-choice knapsack, solved with a Lagrangian search plus a greedy fill.
- Quantize for real. The chosen layout was quantized with GPTQ-style error compensation and exported to MLX's native quantized format. There is no correction adapter.
The resulting layout has 108 ternary matrices, 91 3-bit matrices and one 4-bit matrix. The early layers are much more sensitive than the late ones:
| Layers | ternary | 3-bit | 4-bit |
|---|---|---|---|
| 0-7 | 5 | 45 | 0 |
| 8-15 | 9 | 40 | 1 |
| 16-23 | 44 | 6 | 0 |
| 24-31 | 50 | 0 | 0 |
recipe/sensitivity.json holds the per-matrix cost table.
recipe/allocation.json holds the chosen precision for each matrix.
The method combines established ideas: GPTQ, and sensitivity-guided mixed-precision bit allocation in the spirit of HAWQ-family methods. It does not claim a new quantizer. Sensitivity is measured directly as the KL of Kev's decision probabilities rather than through a Hessian proxy. Relevant work: GPTQ, HAWQ-V3.
Download and Run
Download this repository, including runtime/. It contains the packed-aware
Kev loader and serving source. The v2.0 runtime adds 3-bit support to
kev/packed_model.py and to the exporter; the upstream Kev release alone does
not include this loader. Python 3.12 or 3.13 is required.
python3 -m pip install huggingface_hub
python3 -c "from huggingface_hub import snapshot_download; snapshot_download('Jakevin/kev-4b-ternary-mlx', revision='v2.0', local_dir='kev-4b-ternary-mlx')"
cd kev-4b-ternary-mlx
python3.13 -m venv .venv
.venv/bin/python -m pip install './runtime[serve]'
HF_HUB_OFFLINE=1 HF_HOME=/tmp/kev-empty-hf-cache KEV_BACKEND=mlx \
KEV_PREFIX_CACHE=1 KEV_PREFIX_MAX_TOKENS=4096 \
.venv/bin/python -m kev.serve --run . --host 127.0.0.1 --port 8010
curl http://127.0.0.1:8010/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"state":"I was charged twice for order 8812.","questions":{"team":{"type":"choice","instructions":"Which team should handle this ticket?","criteria":{"billing":"Charges and payments","shipping":"Delivery and delays"}}}}'
The library interface has not changed:
from kev.checkpoint import Checkpoint
tokenizer, model = Checkpoint("path/to/downloaded-repository").load("mps")
To get v1.0, use revision='v1.0' in snapshot_download.
Once dependencies are installed, loading and inference work offline with an empty Hugging Face cache. They never load the original backbone or the original PEFT adapter, and they never rebuild a dense bf16 backbone.
Compression Recipe
Source checkpoint: jaredpalmer/kev-4b,
revision 6cfce5c2fa4b4bd64026336ab649c5ca78857d52, based on
Qwen/Qwen3.5-4B-Base, revision
1001bb4d826a52d1f399e183466143f4da7b741b.
- Source and calibration. The source Kev LoRA was merged before compression. All calibration data comes from 128 records of the decision-v7 calibration partition, RNG 1234. GPTQ-style error compensation uses all 128 records with damping 0.01. Sensitivity probing uses the first 32 records.
- Sensitivity probing details. If a matrix's ternary KL was below 3e-5, 3- and 4-bit were not probed. If its 3-bit KL was below 1e-4, 4-bit was not probed.
- Byte budget. The budget is v1.0's MLX storage cost over the same 200 matrices: 2.943 bits per weight. Ternary and 3-bit codes are counted at 2.5 and 3.5 bits per weight, 4-bit at 4.5, including group scales. The chosen layout uses 2.940 bits per weight.
- Storage formats. Ternary codes use native 2-bit MLX affine storage with group size 64. 3-bit and 4-bit matrices use MLX affine quantization with group size 64. The embedding is 4-bit.
- Unchanged parts. Small
in_proj_a/b, norms, conv kernels and DeltaNet parameters keep their saved source precision. The pointer head and source temperature are unchanged, and no new temperature was fitted. - What is not used. No CLoQ or other compensation adapter, no gradient training and no online rotation.
Activations are bf16. Only requested embedding rows are dequantized. This is a
weight compression method. Activations and the attention/prefix cache are not
claimed to be low-bit. (packed.json still contains the v1 runtime's generic
"arithmetic" string. In v2.0, source_cloq_sha256 is null and no
compensation is applied.)
Evaluation
All suites were evaluated through the native packed MLX path from this exact
model.safetensors, with zero rejected records and zero truncations.
bf16 and v1.0 (downloaded from this repository) were scored with the same harness.
| Suite | Metric | bf16-MLX | v1.0 | v2.0 |
|---|---|---|---|---|
| decision-v7 dev (1,468 rows, 1,264 clean) | accuracy | 0.87184 | 0.84415 | 0.86472 |
| Brier | 0.18248 | 0.23251 | 0.19754 | |
| ECE | 0.0117 | 0.0432 | 0.0147 | |
| transfer-v9 non-MMLU (948 rows, 766 clean) | accuracy | 0.83159 | 0.71932 | 0.78590 |
| Brier | 0.21711 | 0.38167 | 0.29290 | |
| transfer-v9 MMLU (316 rows, 280 clean) | accuracy | 0.60714 | 0.22500 | 0.39643 |
| Brier | 0.53717 | 0.81927 | 0.71195 |
How the suites were used:
- Decision-v7 development was used to select recipes, both in the original v1.0 work and as one of the two publication gates for v2.0. It is a development diagnostic, not a blind test.
- Transfer-v9 was not used to build the v2.0 layout. The allocation depends only on decision-v7 calibration records. However, its non-MMLU part was the second publication gate (publish only if v2.0 beats v1.0 at equal or smaller size).
- MMLU part. It is small (280 clean questions), and both compressed versions lose much of bf16's general-knowledge accuracy. Do not rely on this artifact for knowledge-heavy questions without testing your own task.
Native vs simulated arithmetic. Native MLX quantized matmul differs slightly from the merged-weight simulation used during the experiment. On decision-v7 dev, 5 of 1,468 argmaxes changed, with a mean per-row maximum probability difference of 0.0021 (max 0.032). Transfer non-MMLU had 6 flips out of 948. Reloading the folder in a fresh process gave identical outputs on 20 records.
Aggregate records are in evaluation/v2.0/. v1.0's original records were moved
to evaluation/v1.0/. No evaluation sentences, labels or per-example
predictions are uploaded. No NAER / Chinese Linguipedia data, derivatives or
predictions were used in this export or included in this repository.
Size and Memory
| v1.0 | v2.0 | |
|---|---|---|
model.safetensors |
1,678,909,327 B | 1,677,874,359 B |
| MLX active memory after load | 1.679 GB | 1.678 GB |
Sizes are in decimal GB and exclude the tokenizer and pointer head. Measurements were taken on an Apple M4 Pro with 64 GiB RAM, macOS 26.6.2, with MLX 0.32.2, mlx-lm 0.31.3, torch 2.8.0 and transformers 5.17.0. The MLX peak during the batched evaluation run was 3.61 GB.
v1.0's single-request memory probes (evaluation/v1.0/memory-*.json) were not
repeated for v2.0. Weight memory is about the same, but 3-bit kernels may have
different latency. No 8-GB hardware was tested.
Reproduce
The export and evaluation used the Kev experiment harness. That harness is not redistributed here because it needs the decision-v7 suite. The steps, run from a Kev checkout with an MLX environment:
KEV=<path to jaredpalmer/kev-4b @ 6cfce5c2...>
# 1. per-matrix sensitivity table (ternary / 3-bit / 4-bit, 32 probe records, 128 Hessian records)
python sens4b.py --run $KEV --out sens-fwd.json --kinds tern,affine3,affine4 --n 32 --eps 3e-5 --eps4 1e-4
# 2. knapsack allocation at v1.0's MLX byte budget (2.9426 bits/weight)
python allocate.py --sens sens-fwd.json --budget 2.9426470588235296 --costs mlx --kinds tern,affine3,affine4 --out allocation.json
# 3. GPTQ quantization of that layout -> packed container
python d3.py --run $KEV --layout alloc --alloc allocation.json --save_packed d3-v1size.safetensors
# 4. standalone MLX export (included in this repo)
PYTHONPATH=runtime python runtime/scripts/build_packed_model.py \
--packed d3-v1size.safetensors --checkpoint $KEV \
--config <Qwen/Qwen3.5-4B-Base @ 1001bb4d...>/config.json --out kev-4b-ternary-mlx
You can reuse the outputs of steps 1-2 (recipe/sensitivity.json,
recipe/allocation.json) directly. The packed source hash is recorded in
packed.json (source_packed_sha256). All file hashes are in
release-manifest.json. The runtime's included tests pass with the 3-bit
changes.
Limitations
- Decision-v7 development was used for recipe selection. Transfer non-MMLU was used as a publication gate. No new blind benchmark is claimed.
- A simple depth rule (more bits for early layers, fewer for late layers) at the same file size, evaluated the same way, scores 0.8584 on decision-v7, 0.7846 on transfer-v9 non-MMLU and 0.4429 on transfer-v9 MMLU (v2.0: 0.8647 / 0.7859 / 0.3964). That depth-rule layout is not released here, and no paired significance test has been run yet. v2.0's gain over v1.0 comes mainly from moving bits toward the early layers, not from measuring KL per se.
- KL sensitivities are measured one matrix at a time. Measured whole-model damage is larger than the sum of the single-matrix values, but the ranking between layouts held in our Kev-0.8B checks.
- Brier/ECE are worse than bf16, but better than v1.0. Confidence-sensitive users should evaluate their own task.
- Native inference has not been validated on non-Apple devices or llama.cpp.
- This is an experimental derivative of Kev-4B. It is not a new independently pretrained foundation model, and it is not the Taiwan/mainland wording model.
Attribution and License
Apache-2.0. The original Kev decision model and source are by Jared Palmer.
The base backbone is from Qwen. The compression experiments (v1.0 and
v2.0), standalone exports and publication are by Jakevin. The runtime
distribution modifies Kev's checkpoint and MLX loaders and adds the packed
model implementation. v2.0 adds 3-bit support. See LICENSE and NOTICE for
attribution.
Paper
See Jakevin/kev-d3-paper for the working-draft paper behind v2.0.
- Downloads last month
- 114
Quantized