clef-flash (Core ML)

Cloudflare's clef-flash, a 9B decision model post-trained from Qwen3.5-9B, converted to Core ML for Apple silicon. Text only, 8-bit weights, runs on the Mac GPU.

Same contract as Clef and the Jev / SystemOne API: a state (text or JSON) and typed questions (noul yes/no, choice, ordered score) go in; one probability for every allowed option of every question comes out of a single forward pass. Nothing is generated.

Runtime: ClefFlashManager in FluidUse (Swift, macOS 15+), with a SwiftUI support-ticket triage demo (ClefFlashDemo).

Packages

File Role Size Precision
part00 โ€ฆ part07.mlpackage Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions L256 L512 L1024 L2048 per sequence bucket, weights shared 6.5 GB total 8-bit weights (per channel), fp16 compute
Head.mlpackage Clef's joint schema head, fixed shape: โ‰ค 16 questions, โ‰ค 96 options; same four functions 465 MB fp32
embeddings.f16 input token embeddings (host-side gather) 2.0 GB fp16
output_embeddings.f16 output embeddings, the head's lexical option rows (clef-flash does not tie them) 2.0 GB fp16
tokenizer.json, config.json tokenizer; shapes, buckets and token ids for the host

Each decoder part: hidden [1, L, 4096], cos / sin [L, 64] (fp16) โ†’ out [1, L, 4096]. Text-only records use positions 0..L-1 on all three M-RoPE axes. The host renders the record exactly like Clef's encode_record, gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference host. Images and video are not converted.

Quality

Reference: the same decoder in fp32 (matches Hugging Face's Qwen3_5TextModel to 5e-5 relative) streamed four layers at a time, plus Cloudflare's own JointSchemaHead. 605 records / 611 questions (README samples, ARC-Easy and ARC-Challenge test, first 300 each):

This model (8-bit Core ML) fp32 reference Cloudflare published
answers that differ from fp32 1 / 611 (a near-tie: fp32 top-2 within 0.005) โ€” โ€”
median probability difference 0.0004 โ€” โ€”
ARC-Easy 100.0% 100.0% 99.5%
ARC-Challenge 99.0% 98.67% 98.3%

ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within about a point.

Speed (M5 Pro, 24 GB, GPU)

Median
one decoder part (4 layers), 512-token bucket 62 ms
one record, 3 questions, ~370 tokens (L512), Swift runtime 0.49 s (p95 0.57 s)
load (compiled), 8 parts + head ~40 s

The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply. The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the ANE, so this is a GPU model.

Usage (Swift)

import FluidUse

let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512)
let result = try await manager.answer(
    state: "Our checkout is down and customers can't pay.",
    questions: [
        ("team", .choice(instructions: "Which team should handle this ticket?",
                         criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])),
        ("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])),
    ])
print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds)

License and credits

Apache-2.0, same as Cloudflare/clef-flash (LICENSE is Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision. Conversion by Fluid Inference.

Downloads last month
88
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for FluidInference/clef-flash-coreml

Finetuned
Qwen/Qwen3.5-9B
Quantized
(45)
this model