d1-3B β€” Core AI for iPhone and Mac

LiquidAI/d1-3B converted to Apple Core AI, with the decoder 8-bit palettized to fit an iPhone. d1-3B is a 3B-parameter decision model built on LFM2.5-VL-3B: give it a state β€” text, JSON, pictures, or pictures with text β€” and named questions, and it returns typed answers (a yes/no probability, a pick among named options, or a level) in one forward pass, with no generated tokens.

These are the exact files the D1-3B iPhone app runs, including System One Arcade's Keep It Clean demo: a live camera feed where every frame near a raised middle finger is pixelated before it airs.

iPhone 17 Pro Mac (M3 Max)
Fidelity vs the original model (golden set) 738/740 same answer (the 2 others are near-ties) 740/742
Keep It Clean (860 HaGRIDv2 photos, threshold 0.35) 184/200 caught, 24/660 false alarms β€” same as the original same
A camera frame (one question about a photo) 0.42 s 0.25 s

The same .aimodel files run on both: Core AI specializes them for the device's GPU on first load and caches the result.

Files

File Size What
decoder.aimodel 2.3 GB the LFM2 hybrid decoder (30 layers), 8-bit k-means palettized Linear layers, fp16 math; entrypoints answer, prefill, branch over one copy of the weights
vision.aimodel 815 MB SigLIP2 so400m vision tower + projector, fp16, one image crop per call
embed.f16 500 MB the (tied) token embedding table, 128,000 Γ— 2,048 fp16 β€” looked up on the host, and used for the readout
tokenizer.json 17 MB the original model's tokenizer
meta.json image_token_id, hidden size, the graphs' float input type
LICENSE LFM Open License v1.0

On the iPhone the app needs the com.apple.developer.kernel.increased-memory-limit entitlement; with it, about 6.4 GB stays available after loading. First load specializes the graphs (β‰ˆ 15 s); later loads come from Core AI's cache (β‰ˆ 3–5 s).

How the graphs are used

Graph Inputs β†’ outputs
answer embeds (1, L, 2048) fp16, length (1,) int32 β†’ hidden (1, 2048): the final-norm hidden state at position length βˆ’ 1 (the answer slot)
prefill embeds (1, P, 2048), length (1,) β†’ k_cache, v_cache (8, 1, 8, P, 64) for the 8 attention layers (after q/k norms and RoPE) and conv_tail (22, 1, 2048, 2), the last two values of BΒ·x for the 22 conv layers
branch embeds (B, T, 2048), lengths (B,), prefix_len (1,), the three caches β†’ hidden (B, 2048) at each row's last real token; positions continue at prefix_len + t
vision (main) patches (1, S, 768), row_w/col_w (1, S, 16), valid (1, S), S = 64…1,024 patch slots in 2Γ—2-block order β†’ tokens (1, S/4, 2048)

All lengths are dynamic: right-pad to a bucket (the runtime uses powers of two, 128…8,192 tokens on the iPhone; batch 1/4/16). Padding is exact β€” causal attention and convolution never let a pad reach a real token, cache keys past prefix_len are masked, and reads are one-hot.

One request, end to end:

  1. Prompt (as the original prompt.py):

    <|startoftext|><|im_start|>user\n{"<image>" per picture}{state as indent-2 JSON, or the string}\n\n\nQUESTION:\n
    {question block}<|im_end|>\n<|im_start|>assistant\n
    

    A choice lists its options under single-token codes (A, B, …); a yes/no ends "Reply with yes or no only."; a level lists 0…n. Tokenize with tokenizer.json β€” note it uses BPE ignore_merges (a pre-token that is a vocabulary entry is kept whole), which some tokenizer libraries do not implement.

  2. Pictures. Cap to ≀ 1 M pixels (Pillow bicubic), then LFM2-VL preprocessing: 2–10 tiles of 512 px plus a thumbnail for large images (bicubic, uint8, clipped), 16 px patches normalized to [βˆ’1, 1], 64–256 tokens per crop. Each <image> expands to <|image_start|>, per tile <|img_row_R_col_C|> + 256 <image>, <|img_thumbnail|> + the thumbnail's <image> tokens, <|image_end|>. Run vision once per crop and put its first S/4 rows at the <image> positions.

  3. Decoder. One question: answer over the whole prompt. Several: prefill the shared part (BOS, pictures, state) once, then each question's suffix through branch.

  4. Readout. Dot hidden with the embed.f16 rows of each option's single-token forms (yes/Yes/YES, a digit, an option code and its space-prefixed form); each option scores its best form; softmax over the options. Only those few logits are needed β€” the full-vocabulary log-softmax cancels.

On the iPhone GPU, keep the number of distinct prefill/branch shapes per loaded model small: after the 4th distinct branch shape in one process, calls returned wrong (finite) values; reloading the model clears it. Prompts past 8,192 tokens fail on the iPhone GPU.

The reference for every host step is the original model's own code (prompt.py, runner.py and its LFM2-VL processor in LiquidAI/d1-3B): a port that reproduces its prompt text, token ids, pixel values and readout exactly gets the accuracy below.

Accuracy

Golden set: the original model's answers on 337 cases (742 questions: Fast Decisions text, images, long contexts, JSON states).

Build Where Same top answer Max |Ξ”p|
These files, Swift end to end iPhone 17 Pro 738/740 ΒΉ 0.077
These files, Swift end to end Mac 740/742 0.077

ΒΉ The 16k-token case is refused on the iPhone (8,192 cap). The two differences are near-ties in the original (0.479 vs 0.486; 0.4828 vs 0.4835).

Task accuracy (the port gives the original's answers):

Task Result
Keep It Clean (HaGRIDv2, 200 middle fingers + 660 other gestures, 512 px, the demo's question), threshold 0.35 184/200 caught (92 %), 24/660 false alarms (3.6 %)
same, threshold 0.5 165/200 (82 %), 5/660 (0.8 %)
Fast Decisions, 420 labelled questions 66.4 % (= the original on the same questions; the model card reports 69.3 % on the dev split)

The original model card's benchmark scores (Decision Index, image benchmarks) were not re-run on this conversion.

Compression

Only the decoder's 166 Linear layers are compressed (coreai-opt k-means palettization, 8-bit, one look-up table per group of 16 output channels); norms, convolution taps and the embedding table keep full precision. Compared on the golden set, Keep It Clean and Fast Decisions:

Decoder Size Golden Keep It Clean @0.35 Verdict
8-bit palettized (this) 2.3 GB 740/742 184 / 24 matches the original
int8, per-block 32 2.4 GB 738/742 182 / 24 close
int4, per-block 32 1.3 GB 716/742 182 / 28 answers drift
4-bit palettized 1.1 GB 682/742 191 / 24 drifts
6-bit palettized 1.7 GB β€” β€” breaks the Core AI GPU runtime (runaway memory and disk while specializing)

Speed (iPhone 17 Pro, GPU, warm, thermal nominal)

Median
Keep It Clean frame (vision 120 ms + decoder 304 ms) 424 ms
150 frames back to back (last 30) 484 ms
answer at 128 / 512 / 2,048 tokens 263 / 585 / 2,134 ms
vision, one 512 px crop (768 slots) 143 ms

Modifications from the original

Per the LFM Open License v1.0, the changes made to the Work:

  • Re-authored for Core AI. The decoder was rewritten as three export-friendly graphs over embeddings (answer, prefill, branch), returning the answer-slot hidden state instead of vocabulary logits; padding is handled with masks and one-hot reads. The vision tower runs one crop per call over a 2Γ—2-block patch layout with separable position-embedding weights.
  • Compressed. The decoder's Linear weights are 8-bit k-means palettized; computation is fp16. The vision tower and embedding table are fp16 (the original checkpoint is bf16).
  • Packaging. The embedding table and tokenizer are shipped as separate files for host-side use.

No weights were retrained or fine-tuned; the differences above are measured against the original.

License

LFM Open License v1.0 (LICENSE), the original model's licence. Commercial use is licensed only to entities with annual revenue under US$10 million (non-profit research excepted); see the licence for the full terms. Model by Liquid AI; this conversion is not affiliated with or endorsed by Liquid AI.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for smdesai/d1-3B-CoreAI

Finetuned
LiquidAI/d1-3B
Quantized
(17)
this model