Clef-Flash mixed-bit (measured-KL, text-only)

Version v2.0 (2026-10-08). main is this build. Version v1.0 (T-cloq16, 2026-10-06) stays on the v1.0 tag.

Unofficial post-training quantization of Cloudflare/clef-flash (revision 17f0b0ad64efb65d273590632833508766b2aae6). This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team. Licensed Apache-2.0 like the original (see LICENSE, NOTICE.md for the changes).

Text-only package: the vision tower is removed. The separate vision repo is unchanged.

Versions

version method packed text decision-v7 acc (% of bf16 0.8861) transfer-v9 acc (% of bf16 0.8011) how to load
v2.0 (this, main) measured-KL mixed-bit GPTQ 3.13 GB 0.8766 (98.9%) 0.7361 (91.9%) snapshot_download("Jakevin/clef-flash-ternary-mlx")
v1.0 T-cloq16 (ternary + 4-bit + rank-16 CLoQ) 3.13 GB 0.8497 (95.9%) 0.6788 (84.7%) snapshot_download("Jakevin/clef-flash-ternary-mlx", revision="v1.0")

Paired cluster bootstrap of accuracy, v2.0 − v1.0, 10,000 resamples, seed 1234, 95% percentile CI, clusters = (source, group) within source. Same items, same order, existing per-question rows (no re-run).

suite n (clean) Δ acc 95% CI v2-only correct v1-only correct
decision-v7 1264 +0.0269 [+0.0127, +0.0427] 59 25
transfer-v9 1046 +0.0574 [+0.0344, +0.0803] 105 45

What changed (v2.0)

  • Backbone (Qwen3.5-9B text model, 32 layers: 24 Gated DeltaNet + 8 gated attention): GPTQ post-training quantization, layer-sequential, 128 calibration records (seed 1234). Per-module bit-width is chosen by a knapsack on measured output KL under a packed-file byte budget matching v1.0's 3.13 GB text size. Allocation b31: 81 binary, 34 ternary, 63 3-bit, 22 4-bit modules (200 linears). No CLoQ.
  • Token embedding and lm_head: 4-bit affine (group 64), same as v1.0. The joint-schema head only reads lm_head rows of the option tokens.
  • Unchanged (release bytes): joint-schema head, norms, DeltaNet conv / A_log / dt_bias / in_proj_a / in_proj_b, tokenizer, joint_schema_model.py.
  • Format clef-mkl-v1 (model.safetensors): binary 8 codes/byte + bf16 scale; ternary 5 trits/byte + bf16 scale; 3-bit bitstream + bf16 scale/bias; 4-bit two codes/byte + bf16 scale/bias; group 64. See packed_format.py. Runs on Apple Silicon with MLX (binary and ternary as 2-bit QuantizedLinear, 3-bit and 4-bit native).

Size

bytes
model.safetensors 3.13 GB
whole folder (incl. 0.24 GB head, tokenizer, code) 3.39 GB
original bf16 release 19.06 GB (18.82 GB shards + 0.24 GB head)

Retained capability (text, measured on Apple M4 Pro with this exact runtime; bf16 release scored the same way)

suite (dev) clean questions bf16 release acc this model acc % retained Brier bf16 → this KL(bf16‖this)
decision-v7 1264 0.8861 0.8766 98.9% 0.1694 → 0.1838 0.056
transfer-v9 1046 0.8011 0.7361 91.9% 0.2873 → 0.3633 0.188

Retention depends on the task mix. decision-v7 (whose calibration partition was used for quantization) is the in-distribution number. transfer-v9 is the out-of-distribution check.

Accuracy is on the clean questions of the author's private frozen eval suites (kev; development splits); these are not the card's Decision Index / Typesafe benchmarks.

Use

Requires mlx, mlx-lm (tested 0.31.3 / 0.32.2), torch, transformers, safetensors.

import sys
from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-mlx")  # main = v2.0; or a local folder
sys.path.insert(0, pkg)
import torch
from transformers import AutoTokenizer
from clef_stream import text_args, Head, encode_record, forward_logits
from packed_format import PackedBackbone
bb = PackedBackbone(f"{pkg}/model.safetensors", text_args(pkg))   # ~5.5 GB peak
head = Head(pkg, None, bb.lm_head_q)
tok = AutoTokenizer.from_pretrained(pkg)
record = {"state": "Our checkout started returning errors and orders are blocked.",
          "questions": {"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
                        "outage": {"type": "noul", "instructions": "Is a service down?"}}}
enc = encode_record(tok, record)
zs = forward_logits([enc], bb.embed, bb, bb.norm, head, modes=("q4",))["q4"][0]
print({q.question_id: dict(zip(q.option_ids, torch.softmax(z, -1).tolist())) for q, z in zip(enc.questions, zs)})

To load v1.0 instead:

from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-mlx", revision="v1.0")

Limitations

  • Quantization is post-training only (no recovery training). Calibration used 128 records from kev's decision-v7 calibration partition only (one seed, 1234).
  • The KL reference for the allocation was an 8-bit copy of every decoder linear. The bf16 reference pass exceeded the swap budget on this machine, so that is a deviation from a bf16 KL measurement.
  • The knapsack used v1-style packed-file byte costs (bin 1.25, tern 1.85, 3-bit 3.5, 4-bit 4.5 bits/weight including scales), not in-memory MLX 2-bit storage. Loader still unpacks into MLX QuantizedLinear.
  • No same-size depth-rule control (for example "4-bit the last k layers, ternary the rest") was run at 3.13 GB. The gain versus v1.0 is a real paired difference on these suites; it is not identified as specifically due to measured-KL versus a simpler bit-width rule.
  • Smaller 2.2 GB and 2.5 GB measured-KL builds collapsed (decision-v7 0.36 / 0.64) and are not released.
  • Knowledge-heavy multiple choice is still the weak spot relative to bf16 (transfer-v9 retains 91.9%). Use this model for Clef-style routing / classification decisions close to the calibration mix.
  • The joint-schema head runs in fp32 on CPU here (the release runs it in bf16 on CUDA).
  • Text-only. The vision tower is not in this repo; Jakevin/clef-flash-ternary-vision-mlx is unchanged.

License and attribution

Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself post-trained from Qwen/Qwen3.5-9B (Apache-2.0). Distributed under the Apache License 2.0 (LICENSE); NOTICE.md lists the modifications. Not an official Cloudflare release; "Clef" and "Cloudflare" identify the source model only.

Reproduce

Method and deviations: runs/clef-mkl-20261007/ in the kev repo (not public). Recipe: {"alloc": "b31", "sha256": "df63ead9e19bbefd806534dcd9bde7bdfac106b35b672675d64b261cdd76f344", "calib": 128, "seed": 1234, "group": 64, "damp": 0.01, "emb_bits": 4}.

Downloads last month
62
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
I32
·
U8
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jakevin/clef-flash-ternary-mlx

Finetuned
Qwen/Qwen3.5-9B
Quantized
(42)
this model