Instructions to use mlx-community/clef-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/clef-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/clef-4bit") config = load_config("mlx-community/clef-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/clef-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/clef-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use mlx-community/clef-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/clef-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mlx-community/clef-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/clef-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/clef-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
mlx-community/clef-4bit
Cloudflare/clef converted to MLX (4-bit) for Apple Silicon.
Clef turns a state (text, JSON, images, or video) plus a schema of typed questions into a
probability for every allowed option, in a single forward pass. It is not a chat model —
mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce
meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the
joint schema head.
Usage
pip install mlx-vlm huggingface_hub # no torch needed
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("mlx-community/clef-4bit")
sys.path.insert(0, path)
import clef_mlx
model = clef_mlx.load(path)
response = model.systemone({
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
Images (PIL) and videos (frame arrays) go in images / videos, as in the original:
from PIL import Image
model.predict({
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
})
See the original model card for the input format, question types, and benchmarks.
Conversion
- Backbone:
mlx_vlm.convert -q --q-bits 4 --q-group-size 64(vision tower kept in bf16). - Joint schema head:
joint_head.safetensorscopied unchanged (bf16) and run byclef_mlx.py. processor_config.jsonis the original from Cloudflare/clef; prompt/token layout matches the referencejoint_schema_model.pyexactly (images and video).
Parity vs. official PyTorch implementation (bf16)
| Inputs | Top answer agrees | Max abs Δprob |
|---|---|---|
| Text (4 records, 10 questions) | 10/10 | 0.037 |
| Images + video (5 records, 9 questions) | 9/9 | 0.097 |
Measured on an M5 Max (128 GB). Small spot-check, not a full benchmark run.
License
Apache-2.0, following Cloudflare/clef.
- Downloads last month
- -
4-bit