Clef-flash 9B — MLX 4-bit

Clef-flash is a model from Cloudflare that helps your app choose what to do next. Give it some text and a list of options, and it returns a score for each option. For example, “Please turn the kitchen lights on” can map to “Turn lights on.” Your app can then use that choice to perform the action.

This version runs locally on Apple Silicon with MLX and supports text input.

Property Value
Backbone Qwen3.5, 9B class, 32 layers, hidden size 4096
Format MLX safetensors
Quantization 4-bit affine, groups of 64; original joint head retained
Weight size 5.04 GB backbone + 243.5 MB joint head
Input Text state and schema; no images, video, or audio
Output Choice probabilities, probability of true (noul), expected score index
Context Reference encoder limit 16,384 tokens; Swift default 4,096
License Apache-2.0

The scores show how strongly the model favors each option; a high score does not guarantee a correct answer. Send one request at a time to each loaded model.

Files

File Size Purpose
model.safetensors 5,038,160,953 B Quantized backbone, including separate output embedding
joint_head.safetensors 243,538,016 B Joint schema decision head
config.json 3,028 B Root model configuration and quantization
joint_head_config.json 119 B Head dimensions
tokenizer.json 19,989,339 B Tokenizer vocabulary and rules
tokenizer_config.json 1,186 B Tokenizer settings
clef_mlx.py 9,954 B Pinned upstream Python inference reference
LICENSE 11,544 B Upstream Apache-2.0 license and notices
UPSTREAM_README.md 4,321 B Original converter's card and attribution
export.json Generated Source revision and per-file checksums
example-request.json Small JSON Three-field Swift CLI example
validation-summary.json Generated Bundle checksums and native runtime smoke-test result

Development measurements

Native MLX Swift, release build, compiled Metal shaders, Apple M5 Pro, 48 GB unified memory. One 303-token request with three fields: light action, factual-question check, and urgency. Model load is excluded; encoding through materialized probabilities is included.

Measurement Result Scope
Warm median 0.918 s Five sequential repeats of the same request
Warm mean 0.921 s Same five repeats
Warm range 0.916–0.930 s Same five repeats
First standalone inference 4.02 s One first run, including warm-up effects
Reference token IDs 303/303 identical Python MLX reference
Maximum probability difference About 0.00292 Same three-field reference example

This is a fidelity and timing smoke test, not a broad accuracy benchmark. No memory benchmark, p95 estimate, real robot test, or comparison with hosted inference is claimed. These measurements describe the native Swift implementation; other runtimes and requests can differ. Swift support is a development preview and was not yet released at the time of these measurements (2026-10-02).

Swift

Use the Clef library from speech-swift on Apple Silicon. Clef support is currently a development preview; it requires a checkout containing the Clef library and clef-decide product. Download this model's files into a local directory, then load it:

import Foundation
import Clef

let model = try Clef.load(
    from: URL(fileURLWithPath: "/path/to/Clef-flash-9B-MLX-4bit")
)

let result = try model.decide(
    state: "Please turn the kitchen lights on.",
    questions: [
        ClefQuestion(
            id: "action",
            instructions: "Which action was requested?",
            kind: .choice([
                "on": "Turn lights on",
                "off": "Turn lights off",
                "other": "Other request"
            ])
        )
    ]
)

for decision in result.decisions {
    print(decision.id, decision.selectedOption, decision.confidence)
    print(decision.probabilities)
}

Load the model once, then reuse it for each request. If the text and options are too long, the library returns an error rather than cutting them short.

Command line

swift build -c release --product clef-decide --disable-sandbox
scripts/build_mlx_metallib.sh release
.build/release/clef-decide /path/to/downloaded/model /path/to/downloaded/model/example-request.json

Sources

Links

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aufklarer/Clef-flash-9B-MLX-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(26)
this model

Collection including aufklarer/Clef-flash-9B-MLX-4bit