clef-flash-coreml / README.md
alexwengg's picture
clef-flash Core ML: 8-bit text decoder (8 parts x 4 buckets) + joint schema head
c1b8fda verified
|
Raw History Blame Contribute Delete
4.31 kB
---
license: apache-2.0
base_model: Cloudflare/clef-flash
base_model_relation: quantized
pipeline_tag: text-classification
library_name: coreml
tags:
- coreml
- apple
- decision-model
- typed-output
- structured-output
- classification
- clef
- systemone
- qwen3.5
- fluidinference
---
# clef-flash (Core ML)
Cloudflare's [clef-flash](https://huggingface.co/Cloudflare/clef-flash), a 9B decision model post-trained from
[Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), converted to Core ML for Apple silicon. **Text only**, 8-bit
weights, runs on the Mac GPU.
Same contract as Clef and the Jev / SystemOne API: a `state` (text or JSON) and typed questions (`noul` yes/no,
`choice`, ordered `score`) go in; **one probability for every allowed option of every question** comes out of a
single forward pass. Nothing is generated.
Runtime: `ClefFlashManager` in [FluidUse](https://github.com/FluidInference/FluidUse) (Swift, macOS 15+), with a
SwiftUI support-ticket triage demo (`ClefFlashDemo`).
## Packages
| File | Role | Size | Precision |
|---|---|---:|---|
| `part00` … `part07.mlpackage` | Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions `L256` `L512` `L1024` `L2048` per sequence bucket, weights shared | 6.5 GB total | 8-bit weights (per channel), fp16 compute |
| `Head.mlpackage` | Clef's joint schema head, fixed shape: ≀ 16 questions, ≀ 96 options; same four functions | 465 MB | fp32 |
| `embeddings.f16` | input token embeddings (host-side gather) | 2.0 GB | fp16 |
| `output_embeddings.f16` | output embeddings, the head's lexical option rows (clef-flash does not tie them) | 2.0 GB | fp16 |
| `tokenizer.json`, `config.json` | tokenizer; shapes, buckets and token ids for the host | | |
Each decoder part: `hidden [1, L, 4096]`, `cos` / `sin [L, 64]` (fp16) β†’ `out [1, L, 4096]`. Text-only records use
positions `0..L-1` on all three M-RoPE axes. The host renders the record exactly like Clef's `encode_record`,
gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference
host. Images and video are not converted.
## Quality
Reference: the same decoder in fp32 (matches Hugging Face's `Qwen3_5TextModel` to 5e-5 relative) streamed four
layers at a time, plus Cloudflare's own `JointSchemaHead`. 605 records / 611 questions (README samples, ARC-Easy and
ARC-Challenge test, first 300 each):
| | This model (8-bit Core ML) | fp32 reference | Cloudflare published |
|---|---:|---:|---:|
| answers that differ from fp32 | 1 / 611 (a near-tie: fp32 top-2 within 0.005) | β€” | β€” |
| median probability difference | 0.0004 | β€” | β€” |
| ARC-Easy | 100.0% | 100.0% | 99.5% |
| ARC-Challenge | 99.0% | 98.67% | 98.3% |
ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within
about a point.
## Speed (M5 Pro, 24 GB, GPU)
| | Median |
|---|---:|
| one decoder part (4 layers), 512-token bucket | 62 ms |
| **one record, 3 questions, ~370 tokens (L512)**, Swift runtime | **0.49 s** (p95 0.57 s) |
| load (compiled), 8 parts + head | ~40 s |
The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply.
The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the
ANE, so this is a GPU model.
## Usage (Swift)
```swift
import FluidUse
let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512)
let result = try await manager.answer(
state: "Our checkout is down and customers can't pay.",
questions: [
("team", .choice(instructions: "Which team should handle this ticket?",
criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])),
("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])),
])
print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds)
```
## License and credits
Apache-2.0, same as [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) (`LICENSE` is
Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision.
Conversion by [Fluid Inference](https://fluidinference.com).