mlx-community/clef-4bit

Cloudflare/clef converted to MLX (4-bit) for Apple Silicon.

Clef turns a state (text, JSON, images, or video) plus a schema of typed questions into a probability for every allowed option, in a single forward pass. It is not a chat model — mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the joint schema head.

Usage

pip install mlx-vlm huggingface_hub   # no torch needed
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("mlx-community/clef-4bit")
sys.path.insert(0, path)
import clef_mlx

model = clef_mlx.load(path)
response = model.systemone({
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

Images (PIL) and videos (frame arrays) go in images / videos, as in the original:

from PIL import Image
model.predict({
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
})

See the original model card for the input format, question types, and benchmarks.

Conversion

  • Backbone: mlx_vlm.convert -q --q-bits 4 --q-group-size 64 (vision tower kept in bf16).
  • Joint schema head: joint_head.safetensors copied unchanged (bf16) and run by clef_mlx.py.
  • processor_config.json is the original from Cloudflare/clef; prompt/token layout matches the reference joint_schema_model.py exactly (images and video).

Parity vs. official PyTorch implementation (bf16)

Inputs Top answer agrees Max abs Δprob
Text (4 records, 10 questions) 10/10 0.037
Images + video (5 records, 9 questions) 9/9 0.097

Measured on an M5 Max (128 GB). Small spot-check, not a full benchmark run.

License

Apache-2.0, following Cloudflare/clef.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/clef-4bit

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(5)
this model