d1-3B β Core AI for iPhone and Mac
LiquidAI/d1-3B converted to Apple Core AI, with the decoder 8-bit palettized to fit an iPhone. d1-3B is a 3B-parameter decision model built on LFM2.5-VL-3B: give it a state β text, JSON, pictures, or pictures with text β and named questions, and it returns typed answers (a yes/no probability, a pick among named options, or a level) in one forward pass, with no generated tokens.
These are the exact files the D1-3B iPhone app runs, including System One Arcade's Keep It Clean demo: a live camera feed where every frame near a raised middle finger is pixelated before it airs.
| iPhone 17 Pro | Mac (M3 Max) | |
|---|---|---|
| Fidelity vs the original model (golden set) | 738/740 same answer (the 2 others are near-ties) | 740/742 |
| Keep It Clean (860 HaGRIDv2 photos, threshold 0.35) | 184/200 caught, 24/660 false alarms β same as the original | same |
| A camera frame (one question about a photo) | 0.42 s | 0.25 s |
The same .aimodel files run on both: Core AI specializes them for the device's GPU on first load and caches the
result.
Files
| File | Size | What |
|---|---|---|
decoder.aimodel |
2.3 GB | the LFM2 hybrid decoder (30 layers), 8-bit k-means palettized Linear layers, fp16 math; entrypoints answer, prefill, branch over one copy of the weights |
vision.aimodel |
815 MB | SigLIP2 so400m vision tower + projector, fp16, one image crop per call |
embed.f16 |
500 MB | the (tied) token embedding table, 128,000 Γ 2,048 fp16 β looked up on the host, and used for the readout |
tokenizer.json |
17 MB | the original model's tokenizer |
meta.json |
image_token_id, hidden size, the graphs' float input type |
|
LICENSE |
LFM Open License v1.0 |
On the iPhone the app needs the com.apple.developer.kernel.increased-memory-limit entitlement; with it, about
6.4 GB stays available after loading. First load specializes the graphs (β 15 s); later loads come from Core AI's
cache (β 3β5 s).
How the graphs are used
| Graph | Inputs β outputs |
|---|---|
answer |
embeds (1, L, 2048) fp16, length (1,) int32 β hidden (1, 2048): the final-norm hidden state at position length β 1 (the answer slot) |
prefill |
embeds (1, P, 2048), length (1,) β k_cache, v_cache (8, 1, 8, P, 64) for the 8 attention layers (after q/k norms and RoPE) and conv_tail (22, 1, 2048, 2), the last two values of BΒ·x for the 22 conv layers |
branch |
embeds (B, T, 2048), lengths (B,), prefix_len (1,), the three caches β hidden (B, 2048) at each row's last real token; positions continue at prefix_len + t |
vision (main) |
patches (1, S, 768), row_w/col_w (1, S, 16), valid (1, S), S = 64β¦1,024 patch slots in 2Γ2-block order β tokens (1, S/4, 2048) |
All lengths are dynamic: right-pad to a bucket (the runtime uses powers of two, 128β¦8,192 tokens on the iPhone;
batch 1/4/16). Padding is exact β causal attention and convolution never let a pad reach a real token, cache keys
past prefix_len are masked, and reads are one-hot.
One request, end to end:
Prompt (as the original
prompt.py):<|startoftext|><|im_start|>user\n{"<image>" per picture}{state as indent-2 JSON, or the string}\n\n\nQUESTION:\n {question block}<|im_end|>\n<|im_start|>assistant\nA choice lists its options under single-token codes (A, B, β¦); a yes/no ends "Reply with yes or no only."; a level lists 0β¦n. Tokenize with
tokenizer.jsonβ note it uses BPEignore_merges(a pre-token that is a vocabulary entry is kept whole), which some tokenizer libraries do not implement.Pictures. Cap to β€ 1 M pixels (Pillow bicubic), then LFM2-VL preprocessing: 2β10 tiles of 512 px plus a thumbnail for large images (bicubic, uint8, clipped), 16 px patches normalized to [β1, 1], 64β256 tokens per crop. Each
<image>expands to<|image_start|>, per tile<|img_row_R_col_C|>+ 256<image>,<|img_thumbnail|>+ the thumbnail's<image>tokens,<|image_end|>. Runvisiononce per crop and put its first S/4 rows at the<image>positions.Decoder. One question:
answerover the whole prompt. Several:prefillthe shared part (BOS, pictures, state) once, then each question's suffix throughbranch.Readout. Dot
hiddenwith theembed.f16rows of each option's single-token forms (yes/Yes/YES, a digit, an option code and its space-prefixed form); each option scores its best form; softmax over the options. Only those few logits are needed β the full-vocabulary log-softmax cancels.
On the iPhone GPU, keep the number of distinct prefill/branch shapes per loaded model small: after the 4th distinct
branch shape in one process, calls returned wrong (finite) values; reloading the model clears it. Prompts past
8,192 tokens fail on the iPhone GPU.
The reference for every host step is the original model's own code (prompt.py, runner.py and its LFM2-VL
processor in LiquidAI/d1-3B): a port that reproduces its prompt text, token
ids, pixel values and readout exactly gets the accuracy below.
Accuracy
Golden set: the original model's answers on 337 cases (742 questions: Fast Decisions text, images, long contexts, JSON states).
| Build | Where | Same top answer | Max |Ξp| |
|---|---|---|---|
| These files, Swift end to end | iPhone 17 Pro | 738/740 ΒΉ | 0.077 |
| These files, Swift end to end | Mac | 740/742 | 0.077 |
ΒΉ The 16k-token case is refused on the iPhone (8,192 cap). The two differences are near-ties in the original (0.479 vs 0.486; 0.4828 vs 0.4835).
Task accuracy (the port gives the original's answers):
| Task | Result |
|---|---|
| Keep It Clean (HaGRIDv2, 200 middle fingers + 660 other gestures, 512 px, the demo's question), threshold 0.35 | 184/200 caught (92 %), 24/660 false alarms (3.6 %) |
| same, threshold 0.5 | 165/200 (82 %), 5/660 (0.8 %) |
| Fast Decisions, 420 labelled questions | 66.4 % (= the original on the same questions; the model card reports 69.3 % on the dev split) |
The original model card's benchmark scores (Decision Index, image benchmarks) were not re-run on this conversion.
Compression
Only the decoder's 166 Linear layers are compressed (coreai-opt k-means palettization, 8-bit, one look-up table per group of 16 output channels); norms, convolution taps and the embedding table keep full precision. Compared on the golden set, Keep It Clean and Fast Decisions:
| Decoder | Size | Golden | Keep It Clean @0.35 | Verdict |
|---|---|---|---|---|
| 8-bit palettized (this) | 2.3 GB | 740/742 | 184 / 24 | matches the original |
| int8, per-block 32 | 2.4 GB | 738/742 | 182 / 24 | close |
| int4, per-block 32 | 1.3 GB | 716/742 | 182 / 28 | answers drift |
| 4-bit palettized | 1.1 GB | 682/742 | 191 / 24 | drifts |
| 6-bit palettized | 1.7 GB | β | β | breaks the Core AI GPU runtime (runaway memory and disk while specializing) |
Speed (iPhone 17 Pro, GPU, warm, thermal nominal)
| Median | |
|---|---|
| Keep It Clean frame (vision 120 ms + decoder 304 ms) | 424 ms |
| 150 frames back to back (last 30) | 484 ms |
answer at 128 / 512 / 2,048 tokens |
263 / 585 / 2,134 ms |
| vision, one 512 px crop (768 slots) | 143 ms |
Modifications from the original
Per the LFM Open License v1.0, the changes made to the Work:
- Re-authored for Core AI. The decoder was rewritten as three export-friendly graphs over embeddings
(
answer,prefill,branch), returning the answer-slot hidden state instead of vocabulary logits; padding is handled with masks and one-hot reads. The vision tower runs one crop per call over a 2Γ2-block patch layout with separable position-embedding weights. - Compressed. The decoder's Linear weights are 8-bit k-means palettized; computation is fp16. The vision tower and embedding table are fp16 (the original checkpoint is bf16).
- Packaging. The embedding table and tokenizer are shipped as separate files for host-side use.
No weights were retrained or fine-tuned; the differences above are measured against the original.
License
LFM Open License v1.0 (LICENSE), the original model's licence. Commercial use is licensed only to entities with
annual revenue under US$10 million (non-profit research excepted); see the licence for the full terms. Model by
Liquid AI; this conversion is not affiliated with or endorsed by Liquid AI.