clef-flash-mlx-4bit

Cloudflare/clef-flash converted to MLX for Apple Silicon, with the backbone quantized to 4-bit. Clef-Flash is a decision model: given a state and a schema of typed questions (choice, score, noul), it returns a probability for every allowed option of every question from one prefill pass, with no text generation.

This is an unofficial conversion. It is not made or endorsed by Cloudflare. All credit for the model goes to the Clef authors; see the announcement.

What is in this repo

File Contents
model*.safetensors, config.json Qwen3.5-9B backbone from Clef-Flash, converted with mlx_lm.convert (affine, 4-bit, group size 64). Vision encoder dropped.
joint_head.safetensors, joint_head_config.json Clef's joint schema head, unchanged from the original release (bf16)
clef_mlx.py MLX port of the release's joint_schema_model.py (record encoding and joint schema head); the head runs in float32
tokenizer, chat template, LICENSE From the original release

Text only. The vision encoder is not included, so image and video inputs are not supported.

Usage

pip install mlx-lm huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("TrevorJS/clef-flash-mlx-4bit")
sys.path.insert(0, path)
from clef_mlx import load, decide

clef = load(path)
print(decide(clef, {
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {"type": "choice", "instructions": "Which team should handle the message?",
                       "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
}))
# {'department': {'billing': 0.043, 'technical': 0.957}, 'urgency': {'0': 0.096, '1': 0.072, '2': 0.832},
#  'outage': {'true': 0.818, 'false': 0.182}}   (4-bit output)

decide takes the same record shape as the original encode_record / systemone (state as text or JSON, questions keyed by ID) and returns {question_id: {option_id: probability}}. All questions in a record are scored jointly.

Verification

  • Head: on identical inputs, clef_mlx.JointSchemaHead matches the release's torch JointSchemaHead to within 4e-6 on the logits.
  • Encoding: clef_mlx.encode_record produces the same token IDs, spans and option IDs as the release's encode_record on 11 of 11 sampled records.
  • Backbone: quantization is the only source of drift. No bf16 reference was run, so the end-to-end check is task accuracy and the agreement between the two quantizations (below).

JevBench public items (231 items, pinned commit bb05a335), scored by this repo's code on an Apple M2 (24 GB):

Variant All Easy Standard Hard Peak memory Median latency (M2)
8-bit 188/231 48/48 71/72 69/111 10.1 GB 4.8 s
4-bit 182/231 48/48 68/72 66/111 5.9 GB 2.9 s

The two variants pick the same option on 216 of 231 items (median max per-option probability difference 0.011). Latency is for one M2 and reflects that machine, not the model on a GPU server.

Other variant: TrevorJS/clef-flash-mlx-8bit.

Conversion

mlx-lm 0.31.3, mlx 0.32.3:

python -m mlx_lm convert --hf-path Cloudflare/clef-flash --mlx-path clef-flash-mlx-4bit -q --q-bits 4 --q-group-size 64

then joint_head.safetensors, joint_head_config.json and LICENSE copied from the original repo.

License

Apache-2.0, following Cloudflare/clef-flash and its base model Qwen/Qwen3.5-9B. clef_mlx.py is a port of Cloudflare's Apache-2.0 joint_schema_model.py.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TrevorJS/clef-flash-mlx-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(5)
this model