File size: 10,663 Bytes
65781a8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
---
license: apache-2.0
base_model: Cloudflare/clef
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: mlx
tags:
  - mlx
  - mlx-vlm
  - qwen3.5
  - multimodal
  - clef
  - cloudflare
---

# Clef MLX

MLX quantizations of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), a 27B multimodal decision model post-trained from Qwen3.8-27B.

## MLX Files

| Quantization | File | Size |
| --- | --- | ---: |
| 4-bit | [Clef-MLX-4bit](Clef-MLX-4bit) | 15.5 GB |
| 6-bit | [Clef-MLX-6bit](Clef-MLX-6bit) | 21.6 GB |
| 8-bit | [Clef-MLX-8bit](Clef-MLX-8bit) | 27.7 GB |

## Usage with MLX-VLM

### Installation

```bash
pip install -U mlx-vlm huggingface_hub
```

### Python API

Download the target quantization and load directly with MLX-VLM:

```python
from huggingface_hub import snapshot_download
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

# 1. Download target quantization (Clef-MLX-4bit, Clef-MLX-6bit, or Clef-MLX-8bit)
subfolder = "Clef-MLX-4bit"
model_dir = snapshot_download("abenzerps/Clef-MLX", allow_patterns=f"{subfolder}/*")
model_path = f"{model_dir}/{subfolder}"

# 2. Load model and processor
model, processor = load(model_path)
config = model.config

# Text prompt
prompt = "Explain why reproducible builds matter."
formatted_prompt = apply_chat_template(processor, config, prompt)

output = generate(model, processor, formatted_prompt, verbose=True)
print(output)
```

For image input:

```python
prompt = "Describe this image."
image = ["image.jpg"]

formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=len(image)
)

output = generate(model, processor, formatted_prompt, image, verbose=True)
print(output)
```

### Command Line Interface

Download the quantization folder and run inference via CLI:

```bash
# Download 4-bit quantization folder
huggingface-cli download abenzerps/Clef-MLX --include "Clef-MLX-4bit/*" --local-dir ./Clef-MLX

# 4-bit Text
python -m mlx_vlm.generate \
    --model ./Clef-MLX/Clef-MLX-4bit \
    --prompt "Explain why reproducible builds matter."

# 4-bit Image
python -m mlx_vlm.generate \
    --model ./Clef-MLX/Clef-MLX-4bit \
    --image image.jpg \
    --prompt "Describe this image."
```

<details>
<summary>Original Python / Transformers Usage</summary>

Tested with `torch` 2.11 and `transformers` 5.10.2 on a single H200. Image and video inputs also need `pillow`.

```python
import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))
```

### Jev / SystemOne API
`systemone` takes a Jev/SystemOne `POST /v1/systemone` request body and returns the same response body: `model`, `answers` keyed by question ID, and `usage`. A `choice` answer has `choice`, `confidence`, and `probabilities`; a `score` answer has the expected `score`, `confidence`, `legend`, and `probabilities`; a `noul` answer has the probability of true. `instructions` is optional, and `images` and `videos` may be added to the request.

```python
from joint_schema_model import systemone

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])
```

### Images and video
Add `images` (PIL images) or `videos` (frame arrays) to the record and pass the processor to `encode_record`. Optional processor arguments go in `media_kwargs`.

```python
from PIL import Image

record = {
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
```

Text-only and multimodal records can be mixed in the same batch.

| Field | Description |
|---|---|
| `state` | Any string or JSON value describing the situation to decide on |
| `images`, `videos` | Optional lists of images or video frame arrays |
| `media_kwargs` | Optional keyword arguments for the image/video processor |
| `questions` | Mapping of question ID to question |

Each question has:

- `type`: `noul` (true/false), `choice` (named options), or `score` (ordered options)
- `instructions`: what to decide; optional, and the question ID is used when it is omitted
- `criteria`: for `choice`, a mapping of option ID to description; for `score`, a list of option descriptions indexed from 0; for `noul`, optional descriptions for `true` and `false`

`encode_record` accepts `max_length` (default 16,384 tokens) and `max_state_tokens` to bound the input.

</details>

## Benchmarks

<details>
<summary>Decision Index & Workflow Evals</summary>

### Decision Index
Per-benchmark results from our internal run of the [Decision Index](https://clef-evals.workers-ai-mle.workers.dev) 0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.

| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL (case exact accuracy) | 98.5 | **98.8** | 95.8 | 96.5 | 94.5 | 38.1 |
| ToolRet (nDCG@10) | **69.2** | 66.4 | 65.3 | 61.2 | 64.3 | 12.8 |
| API-Bank (accuracy) | 91.9 | **93.1** | 88.2 | 83.7 | 56.3 | 11.5 |
| BANKING77 (macro-F1) | **94.2** | 90.9 | 79.7 | 74.3 | 84.8 | 14.3 |
| CLINC150+OOS (macro-F1) | **97.4** | 66.8 | 89.3 | 83.5 | 79.0 | 3.2 |
| RouterBench (selected quality) | 79.7 | 79.9 | 79.9 | 79.0 | **80.0** | 57.1 |
| Home appliance simulator (case exact accuracy) | 83.0 | **97.7** | 52.3 | 42.0 | 25.0 | 0.0 |
| SGD/SGD-X (macro-F1) | 43.8 | 34.2 | 43.0 | 40.6 | **64.0** | 42.4 |
| ContractNLI (macro-F1) | 81.4 | **84.3** | 71.7 | 76.0 | 57.8 | 29.0 |
| ANLI (macro-F1) | 69.8 | 59.1 | **74.8** | 66.4 | 56.3 | 48.7 |
| BPoMP (accuracy) | **96.9** | 95.4 | 90.6 | 86.9 | 67.0 | 51.6 |
| Humicroedit (accuracy) | 66.7 | **75.1** | 61.9 | 63.0 | 55.8 | 47.2 |
| POP909-CL (accuracy) | 15.8 | 1.6 | **18.1** | 2.5 | 10.8 | 5.1 |
| cfcolor (accuracy) | **66.0** | 65.8 | 64.7 | 58.2 | 56.3 | 52.3 |
| MMLU (accuracy) | 90.3 | **91.8** | 91.7 | 79.3 | 75.3 | 30.7 |
| GPQA Diamond (accuracy) | 48.0 | 51.0 | **78.3** | 44.9 | 38.8 | 27.6 |
| ARC-Easy (accuracy) | 99.0 | **99.5** | 99.3 | 98.2 | 97.7 | 47.0 |
| ARC-Challenge (accuracy) | 97.7 | **98.3** | 97.8 | 94.5 | 93.7 | 28.6 |
| WinoGrande (accuracy) | 93.5 | **97.5** | 92.0 | 73.6 | 73.2 | 50.5 |
| HellaSwag (accuracy) | 98.2 | **98.6** | 94.5 | 83.3 | 81.9 | 33.1 |
| GSM8K (accuracy) | **80.8** | 67.3 | 79.9 | 50.3 | 48.7 | 21.6 |
| ChessBench (accuracy) | **24.7** | 23.0 | 17.2 | 14.2 | 11.2 | 7.7 |
| MuSR (accuracy) | 83.5 | **86.0** | 66.1 | 61.2 | 57.9 | 43.2 |
| SATA-Bench (case exact accuracy) | 33.8 | **36.7** | 26.4 | 27.5 | 26.7 | 0.3 |
| BRIGHT (nDCG@10) | 45.9 | 39.3 | **47.5** | 42.9 | 38.5 | 19.9 |
| Amazon ESCI (macro-F1) | **57.5** | 57.4 | 55.2 | 53.4 | 49.2 | 24.4 |
| ACOS (per-review F1) | **33.3** | 25.9 | 29.5 | 24.5 | 18.3 | 3.5 |
| FinEntity (macro-F1) | 96.2 | **97.1** | 87.0 | 89.0 | 88.4 | 61.0 |
| VAST (macro-F1) | 59.5 | 49.6 | **64.6** | 55.7 | 55.4 | 40.5 |
| NLI4CT (macro-F1) | 82.9 | 78.6 | **84.1** | 78.4 | 74.9 | 47.7 |
| CRUXEval (accuracy) | **86.7** | 86.1 | 73.0 | 64.7 | 51.2 | 40.2 |
| CLadder (accuracy) | 94.0 | **97.7** | 72.6 | 67.8 | 62.0 | 52.9 |
| ForecastBench (Brier, lower is better) | 13.9 | **10.6** | 17.4 | 29.6 | 17.6 | 41.1 |
| Habermas Machine (accuracy) | 68.7 | **71.8** | 45.9 | 45.0 | 39.4 | 33.4 |
| PhishNChips (accuracy) | 79.6 | 75.0 | 62.5 | **85.4** | 50.7 | 50.1 |
| MMLU-Pro (accuracy) | 65.9 | 65.3 | **82.7** | 56.9 | 51.1 | 13.6 |
| BBH (accuracy) | 73.7 | 68.9 | **92.9** | 70.7 | 65.2 | 34.1 |
| RAGTruth (hallucination F1) | **79.4** | 35.6 | 76.5 | 70.4 | 46.2 | 48.8 |
| HoVer (accuracy) | 65.2 | 61.2 | **72.9** | 70.9 | 58.8 | 55.8 |
| When2Call MCQ (accuracy) | 72.4 | 65.6 | **81.0** | 75.4 | 49.6 | 11.9 |
| New Yorker (accuracy) | 69.5 | 66.1 | **70.1** | 63.6 | 58.1 | 27.1 |
| Median latency (ms) | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 | **5.8** |
| p95 latency (ms) | 238.6 | **122.4** | 536.0 | 211.2 | 187.9 | 222.5 |

### Workflow evals
Decision accuracy on four end-to-end business workflows from [Typesafe Evals](https://evals.typesafe.ai/), scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.

| Workflow | Metric | Clef | Clef-flash | Jev |
|---|---|---:|---:|---:|
| Invoice processing | Exact actions | **64.7** | 57.1 | 61.8 |
| Invoice processing | Primary action | **86.2** | 73.3 | 83.1 |
| Customer service | Exact actions | 76.3 | **77.0** | 76.0 |
| Security incidents | Exact actions | **62.9** | 61.7 | 61.7 |
| Agent trace observability | Primary action | 68.5 | 69.8 | **71.6** |

</details>

## Source

- Model: [Cloudflare/clef](https://huggingface.co/Cloudflare/clef)
- Base model: [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
- License: [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0)
- Checksums: [SHA256SUMS](SHA256SUMS)