Instructions to use Jakevin/clef-flash-ternary-vision-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jakevin/clef-flash-ternary-vision-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Jakevin/clef-flash-ternary-vision-mlx --local-dir clef-flash-ternary-vision-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Clef-Flash ternary (T-cloq16, with vision)
Version v1 (2026-10-06).
Unofficial ternary post-training quantization of Cloudflare/clef-flash
(revision 17f0b0ad64efb65d273590632833508766b2aae6). This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team.
Licensed Apache-2.0 like the original (see LICENSE, NOTICE.md for the changes).
Package with vision: the release's vision tower is kept in bf16, unmodified; the text backbone is identical to the -text package.
What changed
- Backbone (Qwen3.5-9B text model, 32 layers: 24 Gated DeltaNet + 8 gated attention): GPTQ post-training quantization, layer-sequential, 128 calibration records. Ternary (group 64) for DeltaNet in_proj_qkv/in_proj_z and all MLP linears; 4-bit affine for DeltaNet out_proj and attention q/k/v/o; plus a rank-16 closed-form low-rank compensation adapter (CLoQ, bf16) on every quantized linear, kept unfolded (y = W_q x + B(Ax)).
- Token embedding and lm_head: 4-bit affine (group 64). The joint-schema head only reads lm_head rows of the option tokens.
- Unchanged (release bytes): joint-schema head, norms, DeltaNet conv / A_log / dt_bias / in_proj_a / in_proj_b, tokenizer,
joint_schema_model.py, processor config, vision tower (bf16). - Format
clef-ternary-v1(model.safetensors): ternary codes 5 per byte + bf16 scale per 64 weights; 4-bit codes 2 per byte + bf16 scale/bias per 64; seepacked_format.py. Runs on Apple Silicon with MLX (ternary as 2-bitQuantizedLinear).
Size
| bytes | |
|---|---|
model.safetensors |
4.04 GB |
| whole folder (incl. 0.24 GB head, tokenizer, code) | 4.30 GB |
| original bf16 release | 19.06 GB (18.82 GB shards + 0.24 GB head) |
Retained capability (text, measured on Apple M4 Pro with this exact runtime; bf16 release scored the same way)
| suite (dev) | clean questions | bf16 release acc | this model acc | % retained | Brier bf16 → this | KL(bf16‖this) |
|---|---|---|---|---|---|---|
| decision-v7 | 1264 | 0.8861 | 0.8497 | 95.9% | 0.1694 → 0.2056 | 0.090 |
| hard-v1 | 1083 | 0.6473 | 0.4801 | 74.2% | 0.4749 → 0.6284 | 0.386 |
| transfer-v9 | 1046 | 0.8011 | 0.6788 | 84.7% | 0.2873 → 0.4225 | 0.292 |
Retention depends on the task mix. decision-v7 (whose calibration partition was used for quantization) is the in-distribution number. On transfer-v9 the drop is concentrated in knowledge-heavy multiple choice (MMLU / MMLU-Pro lose ~33 points) while classification / NLI / policy-style sources lose 0-5 points; on hard-v1 (hard multi-step decision records) every source drops, most in multi-hop (-27 pt), probability (-20 pt) and judge (-19 pt) questions.
Accuracy is on the clean questions of the author's private frozen eval suites (kev; development splits); these are not the card's Decision Index /
Typesafe benchmarks. Per-example flips exist: on a SystemOne-style routing example ("Our checkout started returning errors and orders are blocked.", billing vs technical) the bf16 release says technical (0.961) and this model says billing (0.788). Images: smoke test only (synthetic solid colours / shapes: 6/6 correct); image accuracy was not measured.
Use
Text-only use is the same as the -text package (PackedBackbone ignores the vision tensors). Images
additionally need mlx-vlm>=0.7.6, torchvision, pillow:
import sys
from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-vision-mlx") # from huggingface_hub; or a local folder
sys.path.insert(0, pkg)
from PIL import Image
from clef_vision import ClefVision
m = ClefVision(pkg) # ~5.8 GB MLX peak
print(m.predict({"state": "Look at the attached image.", "images": [Image.open("photo.jpg")],
"questions": {"colour": {"type": "choice", "instructions": "What is the dominant colour of the object or image?",
"criteria": {"red": None, "green": None, "blue": None, "yellow": None}}}}))
Limitations and planned v2
- Quantization is post-training only (no recovery training). Calibration used 128 records from kev's decision-v7 calibration partition only (sentiment / topic / NLI / intent classification plus synthetic policy & composition records), which is why retention is highest there.
- Knowledge-heavy multiple choice degrades most (transfer-v9: MMLU 0.784 → 0.466, MMLU-Pro 0.650 → 0.315 for T+CLoQ16). Multi-step reasoning degrades too (hard-v1: 74.2% retained; multi-hop 0.605 → 0.339). Use this model for Clef-style routing / classification decisions close to the calibration mix, not for knowledge or multi-step reasoning decisions.
- The joint-schema head runs in fp32 on CPU here (the release runs it in bf16 on CUDA); probabilities are a little less confident than the release (mean confidence 0.903 → 0.870, ECE 0.020 → 0.030 on decision-v7).
- Planned v2 (not done): mixed calibration set (decision-v7 + train/validation splits of MMLU, MMLU-Pro, PAWS, emotion, QNLI + generic text), optionally 4-bit MLP down_proj (+~0.53 GB), re-evaluated on all three suites.
License and attribution
Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself
post-trained from Qwen/Qwen3.5-9B (Apache-2.0). Distributed under the Apache License 2.0 (LICENSE); NOTICE.md lists the
modifications. Not an official Cloudflare release; "Clef" and "Cloudflare" identify the source model only.
Reproduce
Full method, commands and logs: runs/ternary-clef-flash-20261006/README.md in the kev repo (not public). Recipe:
{"scale": "mse", "group": 64, "hi_names": "out_proj,q_proj,k_proj,v_proj,o_proj", "hi_bits": 4, "emb_bits": 4, "cloq_rank": 16, "calib": 128, "damp": 0.01, "suite": "evals/v7/decision-v7"}.
- Downloads last month
- 59
8-bit