|
Download README.md from FluidInference/clef-flash-coreml: direct link, hf CLI and curl.
- Browser
- Download file 4.31 kB
-
https://huggingface.co/FluidInference/clef-flash-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/clef-flash-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/clef-flash-coreml/resolve/main/README.md
4.31 kB
| license: apache-2.0 | |
| base_model: Cloudflare/clef-flash | |
| base_model_relation: quantized | |
| pipeline_tag: text-classification | |
| library_name: coreml | |
| tags: | |
| - coreml | |
| - apple | |
| - decision-model | |
| - typed-output | |
| - structured-output | |
| - classification | |
| - clef | |
| - systemone | |
| - qwen3.5 | |
| - fluidinference | |
| # clef-flash (Core ML) | |
| Cloudflare's [clef-flash](https://huggingface.co/Cloudflare/clef-flash), a 9B decision model post-trained from | |
| [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), converted to Core ML for Apple silicon. **Text only**, 8-bit | |
| weights, runs on the Mac GPU. | |
| Same contract as Clef and the Jev / SystemOne API: a `state` (text or JSON) and typed questions (`noul` yes/no, | |
| `choice`, ordered `score`) go in; **one probability for every allowed option of every question** comes out of a | |
| single forward pass. Nothing is generated. | |
| Runtime: `ClefFlashManager` in [FluidUse](https://github.com/FluidInference/FluidUse) (Swift, macOS 15+), with a | |
| SwiftUI support-ticket triage demo (`ClefFlashDemo`). | |
| ## Packages | |
| | File | Role | Size | Precision | | |
| |---|---|---:|---| | |
| | `part00` β¦ `part07.mlpackage` | Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions `L256` `L512` `L1024` `L2048` per sequence bucket, weights shared | 6.5 GB total | 8-bit weights (per channel), fp16 compute | | |
| | `Head.mlpackage` | Clef's joint schema head, fixed shape: β€ 16 questions, β€ 96 options; same four functions | 465 MB | fp32 | | |
| | `embeddings.f16` | input token embeddings (host-side gather) | 2.0 GB | fp16 | | |
| | `output_embeddings.f16` | output embeddings, the head's lexical option rows (clef-flash does not tie them) | 2.0 GB | fp16 | | |
| | `tokenizer.json`, `config.json` | tokenizer; shapes, buckets and token ids for the host | | | | |
| Each decoder part: `hidden [1, L, 4096]`, `cos` / `sin [L, 64]` (fp16) β `out [1, L, 4096]`. Text-only records use | |
| positions `0..L-1` on all three M-RoPE axes. The host renders the record exactly like Clef's `encode_record`, | |
| gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference | |
| host. Images and video are not converted. | |
| ## Quality | |
| Reference: the same decoder in fp32 (matches Hugging Face's `Qwen3_5TextModel` to 5e-5 relative) streamed four | |
| layers at a time, plus Cloudflare's own `JointSchemaHead`. 605 records / 611 questions (README samples, ARC-Easy and | |
| ARC-Challenge test, first 300 each): | |
| | | This model (8-bit Core ML) | fp32 reference | Cloudflare published | | |
| |---|---:|---:|---:| | |
| | answers that differ from fp32 | 1 / 611 (a near-tie: fp32 top-2 within 0.005) | β | β | | |
| | median probability difference | 0.0004 | β | β | | |
| | ARC-Easy | 100.0% | 100.0% | 99.5% | | |
| | ARC-Challenge | 99.0% | 98.67% | 98.3% | | |
| ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within | |
| about a point. | |
| ## Speed (M5 Pro, 24 GB, GPU) | |
| | | Median | | |
| |---|---:| | |
| | one decoder part (4 layers), 512-token bucket | 62 ms | | |
| | **one record, 3 questions, ~370 tokens (L512)**, Swift runtime | **0.49 s** (p95 0.57 s) | | |
| | load (compiled), 8 parts + head | ~40 s | | |
| The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply. | |
| The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the | |
| ANE, so this is a GPU model. | |
| ## Usage (Swift) | |
| ```swift | |
| import FluidUse | |
| let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512) | |
| let result = try await manager.answer( | |
| state: "Our checkout is down and customers can't pay.", | |
| questions: [ | |
| ("team", .choice(instructions: "Which team should handle this ticket?", | |
| criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])), | |
| ("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])), | |
| ]) | |
| print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds) | |
| ``` | |
| ## License and credits | |
| Apache-2.0, same as [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) (`LICENSE` is | |
| Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision. | |
| Conversion by [Fluid Inference](https://fluidinference.com). | |