--- license: apache-2.0 base_model: Cloudflare/clef-flash base_model_relation: quantized pipeline_tag: text-classification library_name: coreml tags: - coreml - apple - decision-model - typed-output - structured-output - classification - clef - systemone - qwen3.5 - fluidinference --- # clef-flash (Core ML) Cloudflare's [clef-flash](https://huggingface.co/Cloudflare/clef-flash), a 9B decision model post-trained from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), converted to Core ML for Apple silicon. **Text only**, 8-bit weights, runs on the Mac GPU. Same contract as Clef and the Jev / SystemOne API: a `state` (text or JSON) and typed questions (`noul` yes/no, `choice`, ordered `score`) go in; **one probability for every allowed option of every question** comes out of a single forward pass. Nothing is generated. Runtime: `ClefFlashManager` in [FluidUse](https://github.com/FluidInference/FluidUse) (Swift, macOS 15+), with a SwiftUI support-ticket triage demo (`ClefFlashDemo`). ## Packages | File | Role | Size | Precision | |---|---|---:|---| | `part00` … `part07.mlpackage` | Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions `L256` `L512` `L1024` `L2048` per sequence bucket, weights shared | 6.5 GB total | 8-bit weights (per channel), fp16 compute | | `Head.mlpackage` | Clef's joint schema head, fixed shape: ≤ 16 questions, ≤ 96 options; same four functions | 465 MB | fp32 | | `embeddings.f16` | input token embeddings (host-side gather) | 2.0 GB | fp16 | | `output_embeddings.f16` | output embeddings, the head's lexical option rows (clef-flash does not tie them) | 2.0 GB | fp16 | | `tokenizer.json`, `config.json` | tokenizer; shapes, buckets and token ids for the host | | | Each decoder part: `hidden [1, L, 4096]`, `cos` / `sin [L, 64]` (fp16) → `out [1, L, 4096]`. Text-only records use positions `0..L-1` on all three M-RoPE axes. The host renders the record exactly like Clef's `encode_record`, gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference host. Images and video are not converted. ## Quality Reference: the same decoder in fp32 (matches Hugging Face's `Qwen3_5TextModel` to 5e-5 relative) streamed four layers at a time, plus Cloudflare's own `JointSchemaHead`. 605 records / 611 questions (README samples, ARC-Easy and ARC-Challenge test, first 300 each): | | This model (8-bit Core ML) | fp32 reference | Cloudflare published | |---|---:|---:|---:| | answers that differ from fp32 | 1 / 611 (a near-tie: fp32 top-2 within 0.005) | — | — | | median probability difference | 0.0004 | — | — | | ARC-Easy | 100.0% | 100.0% | 99.5% | | ARC-Challenge | 99.0% | 98.67% | 98.3% | ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within about a point. ## Speed (M5 Pro, 24 GB, GPU) | | Median | |---|---:| | one decoder part (4 layers), 512-token bucket | 62 ms | | **one record, 3 questions, ~370 tokens (L512)**, Swift runtime | **0.49 s** (p95 0.57 s) | | load (compiled), 8 parts + head | ~40 s | The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply. The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the ANE, so this is a GPU model. ## Usage (Swift) ```swift import FluidUse let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512) let result = try await manager.answer( state: "Our checkout is down and customers can't pay.", questions: [ ("team", .choice(instructions: "Which team should handle this ticket?", criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])), ("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])), ]) print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds) ``` ## License and credits Apache-2.0, same as [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) (`LICENSE` is Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision. Conversion by [Fluid Inference](https://fluidinference.com).