Instructions to use TrevorJS/clef-flash-mlx-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TrevorJS/clef-flash-mlx-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir clef-flash-mlx-8bit TrevorJS/clef-flash-mlx-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
clef-flash-mlx-8bit
Cloudflare/clef-flash converted to MLX for Apple Silicon,
with the backbone quantized to 8-bit. Clef-Flash is a decision model: given a state and a schema of typed questions
(choice, score, noul), it returns a probability for every allowed option of every question from one prefill pass,
with no text generation.
This is an unofficial conversion. It is not made or endorsed by Cloudflare. All credit for the model goes to the Clef authors; see the announcement.
What is in this repo
| File | Contents |
|---|---|
model*.safetensors, config.json |
Qwen3.5-9B backbone from Clef-Flash, converted with mlx_lm.convert (affine, 8-bit, group size 64). Vision encoder dropped. |
joint_head.safetensors, joint_head_config.json |
Clef's joint schema head, unchanged from the original release (bf16) |
clef_mlx.py |
MLX port of the release's joint_schema_model.py (record encoding and joint schema head); the head runs in float32 |
tokenizer, chat template, LICENSE |
From the original release |
Text only. The vision encoder is not included, so image and video inputs are not supported.
Usage
pip install mlx-lm huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("TrevorJS/clef-flash-mlx-8bit")
sys.path.insert(0, path)
from clef_mlx import load, decide
clef = load(path)
print(decide(clef, {
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
}))
# {'department': {'billing': 0.043, 'technical': 0.957}, 'urgency': {'0': 0.096, '1': 0.072, '2': 0.832},
# 'outage': {'true': 0.818, 'false': 0.182}} (4-bit output)
decide takes the same record shape as the original encode_record / systemone (state as text or JSON, questions keyed
by ID) and returns {question_id: {option_id: probability}}. All questions in a record are scored jointly.
Verification
- Head: on identical inputs,
clef_mlx.JointSchemaHeadmatches the release's torchJointSchemaHeadto within 4e-6 on the logits. - Encoding:
clef_mlx.encode_recordproduces the same token IDs, spans and option IDs as the release'sencode_recordon 11 of 11 sampled records. - Backbone: quantization is the only source of drift. No bf16 reference was run, so the end-to-end check is task accuracy and the agreement between the two quantizations (below).
JevBench public items (231 items, pinned commit bb05a335), scored by this repo's code on an Apple M2 (24 GB):
| Variant | All | Easy | Standard | Hard | Peak memory | Median latency (M2) |
|---|---|---|---|---|---|---|
| 8-bit | 188/231 | 48/48 | 71/72 | 69/111 | 10.1 GB | 4.8 s |
| 4-bit | 182/231 | 48/48 | 68/72 | 66/111 | 5.9 GB | 2.9 s |
The two variants pick the same option on 216 of 231 items (median max per-option probability difference 0.011). Latency is for one M2 and reflects that machine, not the model on a GPU server.
Other variant: TrevorJS/clef-flash-mlx-4bit.
Conversion
mlx-lm 0.31.3, mlx 0.32.3:
python -m mlx_lm convert --hf-path Cloudflare/clef-flash --mlx-path clef-flash-mlx-8bit -q --q-bits 8 --q-group-size 64
then joint_head.safetensors, joint_head_config.json and LICENSE copied from the original repo.
License
Apache-2.0, following Cloudflare/clef-flash and its base model
Qwen/Qwen3.5-9B. clef_mlx.py is a port of Cloudflare's Apache-2.0 joint_schema_model.py.
- Downloads last month
- -
8-bit