Clef-Flash ternary (T-cloq16, with vision)

Version v1 (2026-10-06).

Unofficial ternary post-training quantization of Cloudflare/clef-flash (revision 17f0b0ad64efb65d273590632833508766b2aae6). This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team. Licensed Apache-2.0 like the original (see LICENSE, NOTICE.md for the changes).

Package with vision: the release's vision tower is kept in bf16, unmodified; the text backbone is identical to the -text package.

What changed

  • Backbone (Qwen3.5-9B text model, 32 layers: 24 Gated DeltaNet + 8 gated attention): GPTQ post-training quantization, layer-sequential, 128 calibration records. Ternary (group 64) for DeltaNet in_proj_qkv/in_proj_z and all MLP linears; 4-bit affine for DeltaNet out_proj and attention q/k/v/o; plus a rank-16 closed-form low-rank compensation adapter (CLoQ, bf16) on every quantized linear, kept unfolded (y = W_q x + B(Ax)).
  • Token embedding and lm_head: 4-bit affine (group 64). The joint-schema head only reads lm_head rows of the option tokens.
  • Unchanged (release bytes): joint-schema head, norms, DeltaNet conv / A_log / dt_bias / in_proj_a / in_proj_b, tokenizer, joint_schema_model.py, processor config, vision tower (bf16).
  • Format clef-ternary-v1 (model.safetensors): ternary codes 5 per byte + bf16 scale per 64 weights; 4-bit codes 2 per byte + bf16 scale/bias per 64; see packed_format.py. Runs on Apple Silicon with MLX (ternary as 2-bit QuantizedLinear).

Size

bytes
model.safetensors 4.04 GB
whole folder (incl. 0.24 GB head, tokenizer, code) 4.30 GB
original bf16 release 19.06 GB (18.82 GB shards + 0.24 GB head)

Retained capability (text, measured on Apple M4 Pro with this exact runtime; bf16 release scored the same way)

suite (dev) clean questions bf16 release acc this model acc % retained Brier bf16 → this KL(bf16‖this)
decision-v7 1264 0.8861 0.8497 95.9% 0.1694 → 0.2056 0.090
hard-v1 1083 0.6473 0.4801 74.2% 0.4749 → 0.6284 0.386
transfer-v9 1046 0.8011 0.6788 84.7% 0.2873 → 0.4225 0.292

Retention depends on the task mix. decision-v7 (whose calibration partition was used for quantization) is the in-distribution number. On transfer-v9 the drop is concentrated in knowledge-heavy multiple choice (MMLU / MMLU-Pro lose ~33 points) while classification / NLI / policy-style sources lose 0-5 points; on hard-v1 (hard multi-step decision records) every source drops, most in multi-hop (-27 pt), probability (-20 pt) and judge (-19 pt) questions.

Accuracy is on the clean questions of the author's private frozen eval suites (kev; development splits); these are not the card's Decision Index / Typesafe benchmarks. Per-example flips exist: on a SystemOne-style routing example ("Our checkout started returning errors and orders are blocked.", billing vs technical) the bf16 release says technical (0.961) and this model says billing (0.788). Images: smoke test only (synthetic solid colours / shapes: 6/6 correct); image accuracy was not measured.

Use

Text-only use is the same as the -text package (PackedBackbone ignores the vision tensors). Images additionally need mlx-vlm>=0.7.6, torchvision, pillow:

import sys
from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-vision-mlx")  # from huggingface_hub; or a local folder
sys.path.insert(0, pkg)
from PIL import Image
from clef_vision import ClefVision
m = ClefVision(pkg)   # ~5.8 GB MLX peak
print(m.predict({"state": "Look at the attached image.", "images": [Image.open("photo.jpg")],
                 "questions": {"colour": {"type": "choice", "instructions": "What is the dominant colour of the object or image?",
                                          "criteria": {"red": None, "green": None, "blue": None, "yellow": None}}}}))

Limitations and planned v2

  • Quantization is post-training only (no recovery training). Calibration used 128 records from kev's decision-v7 calibration partition only (sentiment / topic / NLI / intent classification plus synthetic policy & composition records), which is why retention is highest there.
  • Knowledge-heavy multiple choice degrades most (transfer-v9: MMLU 0.784 → 0.466, MMLU-Pro 0.650 → 0.315 for T+CLoQ16). Multi-step reasoning degrades too (hard-v1: 74.2% retained; multi-hop 0.605 → 0.339). Use this model for Clef-style routing / classification decisions close to the calibration mix, not for knowledge or multi-step reasoning decisions.
  • The joint-schema head runs in fp32 on CPU here (the release runs it in bf16 on CUDA); probabilities are a little less confident than the release (mean confidence 0.903 → 0.870, ECE 0.020 → 0.030 on decision-v7).
  • Planned v2 (not done): mixed calibration set (decision-v7 + train/validation splits of MMLU, MMLU-Pro, PAWS, emotion, QNLI + generic text), optionally 4-bit MLP down_proj (+~0.53 GB), re-evaluated on all three suites.

License and attribution

Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself post-trained from Qwen/Qwen3.5-9B (Apache-2.0). Distributed under the Apache License 2.0 (LICENSE); NOTICE.md lists the modifications. Not an official Cloudflare release; "Clef" and "Cloudflare" identify the source model only.

Reproduce

Full method, commands and logs: runs/ternary-clef-flash-20261006/README.md in the kev repo (not public). Recipe: {"scale": "mse", "group": 64, "hi_names": "out_proj,q_proj,k_proj,v_proj,o_proj", "hi_bits": 4, "emb_bits": 4, "cloq_rank": 16, "calib": 128, "damp": 0.01, "suite": "evals/v7/decision-v7"}.

Downloads last month
59
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
I32
·
U8
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jakevin/clef-flash-ternary-vision-mlx

Finetuned
Qwen/Qwen3.5-9B
Quantized
(42)
this model