Instructions to use aufklarer/Clef-flash-9B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aufklarer/Clef-flash-9B-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Clef-flash-9B-MLX-4bit aufklarer/Clef-flash-9B-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Clef-flash 9B — MLX 4-bit
Clef-flash is a model from Cloudflare that helps your app choose what to do next. Give it some text and a list of options, and it returns a score for each option. For example, “Please turn the kitchen lights on” can map to “Turn lights on.” Your app can then use that choice to perform the action.
This version runs locally on Apple Silicon with MLX and supports text input.
| Property | Value |
|---|---|
| Backbone | Qwen3.5, 9B class, 32 layers, hidden size 4096 |
| Format | MLX safetensors |
| Quantization | 4-bit affine, groups of 64; original joint head retained |
| Weight size | 5.04 GB backbone + 243.5 MB joint head |
| Input | Text state and schema; no images, video, or audio |
| Output | Choice probabilities, probability of true (noul), expected score index |
| Context | Reference encoder limit 16,384 tokens; Swift default 4,096 |
| License | Apache-2.0 |
The scores show how strongly the model favors each option; a high score does not guarantee a correct answer. Send one request at a time to each loaded model.
Files
| File | Size | Purpose |
|---|---|---|
model.safetensors |
5,038,160,953 B | Quantized backbone, including separate output embedding |
joint_head.safetensors |
243,538,016 B | Joint schema decision head |
config.json |
3,028 B | Root model configuration and quantization |
joint_head_config.json |
119 B | Head dimensions |
tokenizer.json |
19,989,339 B | Tokenizer vocabulary and rules |
tokenizer_config.json |
1,186 B | Tokenizer settings |
clef_mlx.py |
9,954 B | Pinned upstream Python inference reference |
LICENSE |
11,544 B | Upstream Apache-2.0 license and notices |
UPSTREAM_README.md |
4,321 B | Original converter's card and attribution |
export.json |
Generated | Source revision and per-file checksums |
example-request.json |
Small JSON | Three-field Swift CLI example |
validation-summary.json |
Generated | Bundle checksums and native runtime smoke-test result |
Development measurements
Native MLX Swift, release build, compiled Metal shaders, Apple M5 Pro, 48 GB unified memory. One 303-token request with three fields: light action, factual-question check, and urgency. Model load is excluded; encoding through materialized probabilities is included.
| Measurement | Result | Scope |
|---|---|---|
| Warm median | 0.918 s | Five sequential repeats of the same request |
| Warm mean | 0.921 s | Same five repeats |
| Warm range | 0.916–0.930 s | Same five repeats |
| First standalone inference | 4.02 s | One first run, including warm-up effects |
| Reference token IDs | 303/303 identical | Python MLX reference |
| Maximum probability difference | About 0.00292 | Same three-field reference example |
This is a fidelity and timing smoke test, not a broad accuracy benchmark. No memory benchmark, p95 estimate, real robot test, or comparison with hosted inference is claimed. These measurements describe the native Swift implementation; other runtimes and requests can differ. Swift support is a development preview and was not yet released at the time of these measurements (2026-10-02).
Swift
Use the Clef library from speech-swift
on Apple Silicon. Clef support is currently a development preview; it requires a
checkout containing the Clef library and clef-decide product.
Download this model's files into a local directory, then load it:
import Foundation
import Clef
let model = try Clef.load(
from: URL(fileURLWithPath: "/path/to/Clef-flash-9B-MLX-4bit")
)
let result = try model.decide(
state: "Please turn the kitchen lights on.",
questions: [
ClefQuestion(
id: "action",
instructions: "Which action was requested?",
kind: .choice([
"on": "Turn lights on",
"off": "Turn lights off",
"other": "Other request"
])
)
]
)
for decision in result.decisions {
print(decision.id, decision.selectedOption, decision.confidence)
print(decision.probabilities)
}
Load the model once, then reuse it for each request. If the text and options are too long, the library returns an error rather than cutting them short.
Command line
swift build -c release --product clef-decide --disable-sandbox
scripts/build_mlx_metallib.sh release
.build/release/clef-decide /path/to/downloaded/model /path/to/downloaded/model/example-request.json
Sources
Links
- Downloads last month
- -
4-bit