File size: 4,312 Bytes
c1b8fda
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
---
license: apache-2.0
base_model: Cloudflare/clef-flash
base_model_relation: quantized
pipeline_tag: text-classification
library_name: coreml
tags:
- coreml
- apple
- decision-model
- typed-output
- structured-output
- classification
- clef
- systemone
- qwen3.5
- fluidinference
---

# clef-flash (Core ML)

Cloudflare's [clef-flash](https://huggingface.co/Cloudflare/clef-flash), a 9B decision model post-trained from
[Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), converted to Core ML for Apple silicon. **Text only**, 8-bit
weights, runs on the Mac GPU.

Same contract as Clef and the Jev / SystemOne API: a `state` (text or JSON) and typed questions (`noul` yes/no,
`choice`, ordered `score`) go in; **one probability for every allowed option of every question** comes out of a
single forward pass. Nothing is generated.

Runtime: `ClefFlashManager` in [FluidUse](https://github.com/FluidInference/FluidUse) (Swift, macOS 15+), with a
SwiftUI support-ticket triage demo (`ClefFlashDemo`).

## Packages

| File | Role | Size | Precision |
|---|---|---:|---|
| `part00` … `part07.mlpackage` | Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions `L256` `L512` `L1024` `L2048` per sequence bucket, weights shared | 6.5 GB total | 8-bit weights (per channel), fp16 compute |
| `Head.mlpackage` | Clef's joint schema head, fixed shape: ≤ 16 questions, ≤ 96 options; same four functions | 465 MB | fp32 |
| `embeddings.f16` | input token embeddings (host-side gather) | 2.0 GB | fp16 |
| `output_embeddings.f16` | output embeddings, the head's lexical option rows (clef-flash does not tie them) | 2.0 GB | fp16 |
| `tokenizer.json`, `config.json` | tokenizer; shapes, buckets and token ids for the host | | |

Each decoder part: `hidden [1, L, 4096]`, `cos` / `sin [L, 64]` (fp16) → `out [1, L, 4096]`. Text-only records use
positions `0..L-1` on all three M-RoPE axes. The host renders the record exactly like Clef's `encode_record`,
gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference
host. Images and video are not converted.

## Quality

Reference: the same decoder in fp32 (matches Hugging Face's `Qwen3_5TextModel` to 5e-5 relative) streamed four
layers at a time, plus Cloudflare's own `JointSchemaHead`. 605 records / 611 questions (README samples, ARC-Easy and
ARC-Challenge test, first 300 each):

| | This model (8-bit Core ML) | fp32 reference | Cloudflare published |
|---|---:|---:|---:|
| answers that differ from fp32 | 1 / 611 (a near-tie: fp32 top-2 within 0.005) | — | — |
| median probability difference | 0.0004 | — | — |
| ARC-Easy | 100.0% | 100.0% | 99.5% |
| ARC-Challenge | 99.0% | 98.67% | 98.3% |

ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within
about a point.

## Speed (M5 Pro, 24 GB, GPU)

| | Median |
|---|---:|
| one decoder part (4 layers), 512-token bucket | 62 ms |
| **one record, 3 questions, ~370 tokens (L512)**, Swift runtime | **0.49 s** (p95 0.57 s) |
| load (compiled), 8 parts + head | ~40 s |

The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply.
The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the
ANE, so this is a GPU model.

## Usage (Swift)

```swift
import FluidUse

let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512)
let result = try await manager.answer(
    state: "Our checkout is down and customers can't pay.",
    questions: [
        ("team", .choice(instructions: "Which team should handle this ticket?",
                         criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])),
        ("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])),
    ])
print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds)
```

## License and credits

Apache-2.0, same as [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) (`LICENSE` is
Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision.
Conversion by [Fluid Inference](https://fluidinference.com).