Instructions to use Jakevin/clef-flash-ternary-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jakevin/clef-flash-ternary-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Jakevin/clef-flash-ternary-mlx --local-dir clef-flash-ternary-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Clef-Flash mixed-bit (measured-KL, text-only)
Version v2.0 (2026-10-08). main is this build. Version v1.0 (T-cloq16, 2026-10-06) stays on the v1.0 tag.
Unofficial post-training quantization of Cloudflare/clef-flash
(revision 17f0b0ad64efb65d273590632833508766b2aae6). This is not an official Cloudflare release and is not endorsed by Cloudflare or the Qwen team.
Licensed Apache-2.0 like the original (see LICENSE, NOTICE.md for the changes).
Text-only package: the vision tower is removed. The separate vision repo is unchanged.
Versions
| version | method | packed text | decision-v7 acc (% of bf16 0.8861) | transfer-v9 acc (% of bf16 0.8011) | how to load |
|---|---|---|---|---|---|
v2.0 (this, main) |
measured-KL mixed-bit GPTQ | 3.13 GB | 0.8766 (98.9%) | 0.7361 (91.9%) | snapshot_download("Jakevin/clef-flash-ternary-mlx") |
| v1.0 | T-cloq16 (ternary + 4-bit + rank-16 CLoQ) | 3.13 GB | 0.8497 (95.9%) | 0.6788 (84.7%) | snapshot_download("Jakevin/clef-flash-ternary-mlx", revision="v1.0") |
Paired cluster bootstrap of accuracy, v2.0 − v1.0, 10,000 resamples, seed 1234, 95% percentile CI, clusters = (source, group) within source. Same items, same order, existing per-question rows (no re-run).
| suite | n (clean) | Δ acc | 95% CI | v2-only correct | v1-only correct |
|---|---|---|---|---|---|
| decision-v7 | 1264 | +0.0269 | [+0.0127, +0.0427] | 59 | 25 |
| transfer-v9 | 1046 | +0.0574 | [+0.0344, +0.0803] | 105 | 45 |
What changed (v2.0)
- Backbone (Qwen3.5-9B text model, 32 layers: 24 Gated DeltaNet + 8 gated attention): GPTQ post-training quantization, layer-sequential, 128 calibration records (seed 1234). Per-module bit-width is chosen by a knapsack on measured output KL under a packed-file byte budget matching v1.0's 3.13 GB text size. Allocation b31: 81 binary, 34 ternary, 63 3-bit, 22 4-bit modules (200 linears). No CLoQ.
- Token embedding and lm_head: 4-bit affine (group 64), same as v1.0. The joint-schema head only reads lm_head rows of the option tokens.
- Unchanged (release bytes): joint-schema head, norms, DeltaNet conv / A_log / dt_bias / in_proj_a / in_proj_b, tokenizer,
joint_schema_model.py. - Format
clef-mkl-v1(model.safetensors): binary 8 codes/byte + bf16 scale; ternary 5 trits/byte + bf16 scale; 3-bit bitstream + bf16 scale/bias; 4-bit two codes/byte + bf16 scale/bias; group 64. Seepacked_format.py. Runs on Apple Silicon with MLX (binary and ternary as 2-bitQuantizedLinear, 3-bit and 4-bit native).
Size
| bytes | |
|---|---|
model.safetensors |
3.13 GB |
| whole folder (incl. 0.24 GB head, tokenizer, code) | 3.39 GB |
| original bf16 release | 19.06 GB (18.82 GB shards + 0.24 GB head) |
Retained capability (text, measured on Apple M4 Pro with this exact runtime; bf16 release scored the same way)
| suite (dev) | clean questions | bf16 release acc | this model acc | % retained | Brier bf16 → this | KL(bf16‖this) |
|---|---|---|---|---|---|---|
| decision-v7 | 1264 | 0.8861 | 0.8766 | 98.9% | 0.1694 → 0.1838 | 0.056 |
| transfer-v9 | 1046 | 0.8011 | 0.7361 | 91.9% | 0.2873 → 0.3633 | 0.188 |
Retention depends on the task mix. decision-v7 (whose calibration partition was used for quantization) is the in-distribution number. transfer-v9 is the out-of-distribution check.
Accuracy is on the clean questions of the author's private frozen eval suites (kev; development splits); these are not the card's Decision Index / Typesafe benchmarks.
Use
Requires mlx, mlx-lm (tested 0.31.3 / 0.32.2), torch, transformers, safetensors.
import sys
from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-mlx") # main = v2.0; or a local folder
sys.path.insert(0, pkg)
import torch
from transformers import AutoTokenizer
from clef_stream import text_args, Head, encode_record, forward_logits
from packed_format import PackedBackbone
bb = PackedBackbone(f"{pkg}/model.safetensors", text_args(pkg)) # ~5.5 GB peak
head = Head(pkg, None, bb.lm_head_q)
tok = AutoTokenizer.from_pretrained(pkg)
record = {"state": "Our checkout started returning errors and orders are blocked.",
"questions": {"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"}}}
enc = encode_record(tok, record)
zs = forward_logits([enc], bb.embed, bb, bb.norm, head, modes=("q4",))["q4"][0]
print({q.question_id: dict(zip(q.option_ids, torch.softmax(z, -1).tolist())) for q, z in zip(enc.questions, zs)})
To load v1.0 instead:
from huggingface_hub import snapshot_download
pkg = snapshot_download("Jakevin/clef-flash-ternary-mlx", revision="v1.0")
Limitations
- Quantization is post-training only (no recovery training). Calibration used 128 records from kev's decision-v7 calibration partition only (one seed, 1234).
- The KL reference for the allocation was an 8-bit copy of every decoder linear. The bf16 reference pass exceeded the swap budget on this machine, so that is a deviation from a bf16 KL measurement.
- The knapsack used v1-style packed-file byte costs (bin 1.25, tern 1.85, 3-bit 3.5, 4-bit 4.5 bits/weight including
scales), not in-memory MLX 2-bit storage. Loader still unpacks into MLX
QuantizedLinear. - No same-size depth-rule control (for example "4-bit the last k layers, ternary the rest") was run at 3.13 GB. The gain versus v1.0 is a real paired difference on these suites; it is not identified as specifically due to measured-KL versus a simpler bit-width rule.
- Smaller 2.2 GB and 2.5 GB measured-KL builds collapsed (decision-v7 0.36 / 0.64) and are not released.
- Knowledge-heavy multiple choice is still the weak spot relative to bf16 (transfer-v9 retains 91.9%). Use this model for Clef-style routing / classification decisions close to the calibration mix.
- The joint-schema head runs in fp32 on CPU here (the release runs it in bf16 on CUDA).
- Text-only. The vision tower is not in this repo;
Jakevin/clef-flash-ternary-vision-mlxis unchanged.
License and attribution
Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself
post-trained from Qwen/Qwen3.5-9B (Apache-2.0). Distributed under the Apache License 2.0 (LICENSE); NOTICE.md lists the
modifications. Not an official Cloudflare release; "Clef" and "Cloudflare" identify the source model only.
Reproduce
Method and deviations: runs/clef-mkl-20261007/ in the kev repo (not public). Recipe:
{"alloc": "b31", "sha256": "df63ead9e19bbefd806534dcd9bde7bdfac106b35b672675d64b261cdd76f344", "calib": 128, "seed": 1234, "group": 64, "damp": 0.01, "emb_bits": 4}.
- Downloads last month
- 62
8-bit