alexwengg's picture
clef-text-0.6b Core ML: fp16 decoder (L256-L2048) + head, GPU / Neural Engine
3c61a60 verified
|
Raw History Blame Contribute Delete
3.99 kB
metadata
license: apache-2.0
base_model: FluidInference/clef-text-0.6b
base_model_relation: quantized
pipeline_tag: text-classification
library_name: coreml
tags:
  - coreml
  - apple
  - neural-engine
  - decision-model
  - typed-output
  - classification
  - clef
  - systemone
  - distillation
  - qwen3
  - fluidinference

clef-text-0.6b (Core ML)

A 0.70B-parameter text decision model for Apple silicon, distilled from Cloudflare's clef-flash (9B) into Qwen3-0.6B plus Clef's joint schema head. Same contract as Clef and the Jev / SystemOne API: a state and typed questions (noul, choice, score) in, a probability for every option of every question out, one forward pass.

Runs on the GPU or 100% on the Neural Engine (every decoder op placed on the ANE at all four buckets). PyTorch weights: FluidInference/clef-text-0.6b. Swift runtime: ClefTextManager in FluidUse.

Packages

File Role Size
Decoder.mlpackage Qwen3-0.6B decoder + final norm, fp16; functions L256 L512 L1024 L2048 (shared weights) 848 MB
Head.mlpackage joint schema head, fp16, ≤ 16 questions / ≤ 96 options, same four functions 197 MB
embeddings.f16 token embeddings (tied: also the head's lexical option rows) 297 MB
tokenizer.json, config.json tokenizer; shapes and ids for the host

Decoder inputs: hidden [1, L, 1024], cos / sin [L, 128], additive causal mask [1, 1, L, L] (all fp16) → states [1, L, 1024].

Quality

Held-out test splits (official test / validation sets, 5,124 questions), gold accuracy:

Task this model clef-flash 9B (teacher)
DBpedia-14 entity type 99.7 100.0
AG News topic 90.7 92.0
SST-2 sentiment 90.0 94.0
BANKING77 intent (77-way) 86.4 94.5
TweetEval offensive 85.0 82.3
Yelp stars (5-way) 69.0 70.7
Emotion (6-way) 68.0 58.0
BoolQ 82.3 90.7
MNLI 72.3 84.3
ARC-Easy 80.0 100.0
CommonsenseQA 65.0 94.3
ARC-Challenge 60.7 97.3
all 79.1 88.2

Good at routing and classification (within 0–4 points of the teacher on topic, entity, sentiment, rating, toxicity, emotion); not a knowledge model — science / commonsense QA and NLI stay well below the 9B. Teacher agreement 86.4%, calibration ECE 0.046 (dev).

Core ML on the Neural Engine vs the PyTorch student on the same 5,118 test questions: 15 different answers (0.3%, all near-ties: PyTorch top-2 within 0.11), median probability drift 0.013, gold accuracy equal within ±0.7 per task.

Speed (M5 Pro)

Decoder at 512 tokens: 35 ms on the GPU (.cpuAndGPU), 64 ms on the Neural Engine (.cpuAndNeuralEngine), 200 ms CPU-only. On this Mac the GPU is ~2× faster; the Neural Engine leaves the GPU free for other work.

Bucket Neural Engine, decoder + head
256 tokens 29 ms
512 tokens 69 ms
1,024 tokens 196 ms
2,048 tokens 591 ms

Swift runtime, 3-question support ticket (370 tokens): **48 ms on the GPU** (1,000 tickets in 57 s in the FluidUse demo), ~77 ms on the Neural Engine — vs ~0.5 s for clef-flash Core ML (9B, GPU).

Training

24,208 records from 12 public datasets (train splits; test splits held out) labelled by clef-flash (8-bit Core ML, 1 flip / 611 vs fp32), KL (T = 2) + 0.3 gold cross-entropy, LoRA r64 on the backbone + full head, 2 epochs. The head starts from clef-flash's head weights wherever the shapes match.

License and credits

Apache-2.0. Teacher: Cloudflare/clef-flash (Apache-2.0); base: Qwen/Qwen3-0.6B (Apache-2.0). By Fluid Inference.