D1A-E4B v0.6 · Core ML (text, ≤ 512 tokens)

D1A-E4B v0.6 for the Apple Neural Engine. D1A is a small decision model: a document and typed questions go in, and a calibrated probability for every option comes out, in one pass, with no generated text (jonpol01/d1a). This repository is the text model converted to Core ML; the conversion is described in jonpol01/d1a#224.

Status

  • Correct at ≤ 512 tokens. On 28 real requests (42 questions), no clear-cut answer changes against the fp32 PyTorch rebuild of the same model, on an M1 Max and an M4.
  • Speed (median request): M4 Neural Engine 1.8 s (MLX on the GPU: 0.63 s); M1 Max 2.9 s (MLX: 0.81 s).
  • Use it where the GPU should stay free or power matters (iPhone, iPad). On a Mac, the MLX build is faster.
  • Not included yet: photo, voice and video (their encoders are separate models, still to convert), and inputs longer than 512 tokens (2,048-token programs give wrong answers on the M4 Neural Engine).

Run it

pip install coremltools numpy torch transformers safetensors "d1a @ git+https://github.com/jonpol01/d1a"
hf download JohnP1/d1a-e4b-coreml --revision v0.6 --local-dir d1a-e4b-coreml
python d1a-e4b-coreml/reference/run.py d1a-e4b-coreml request.json

request.json is a System One request, for example:

{"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
 "questions": {"team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"returns": null, "shipping": null, "billing": null}},
               "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}}}

The first run compiles the chunks into coreml/.compiled/ (a few minutes); later runs load them.

Files

Path What
coreml/L512_<a>-<b>.mlpackage 7 chunks of the 42-layer text model, int8 weights (per block of 64). Chunk 18-24 outputs the shared keys and values that chunks 24-42 read.
embeddings.safetensors the token and per-layer embedding tables (MLX affine-quantized, from JohnP1/d1a-e4b-mlx-q8@v0.6), read row by row on the host
head.safetensors, d1a_config.json the pointer head, the temperature and the pad and delimiter ids
config.json, tokenizer.json, tokenizer_config.json the Gemma 4 model config and tokenizer
io.json each chunk's input names
reference/ Python reference: run.py (one request), workers.py (one process per chunk), d1a_coreml.py (host-side embeddings and masks)

How a request runs

  1. Encode it with D1A's layout (d1a.core.encoding.encode) and right-pad it to 512 tokens with the pad token.
  2. Look up the token embeddings (scaled by √hidden) and the per-layer embeddings (scaled by √256) on the host.
  3. Build the block-causal masks (full and sliding), clamped to −1e4.
  4. Run the chunks in order, each in its own process. Several large Neural Engine programs in one process can crash the runtime.
  5. Read the hidden states at each question's <decide> and </opt> positions, score them with the head, divide by the temperature, and take the softmax.

macOS 15 / iOS 18 or later.

D1A builds compared

Measured on 2026-10-11 on the same 28 real requests (42 questions, all ≤ 512 tokens), on an Apple M1 Max and a Mac mini M4.

E2B · MLX E2B · Core ML E4B · MLX E4B · Core ML
Version v0.6 v0.6 v0.6 v0.6
Repository JohnP1/d1a-e2b-mlx-q8 JohnP1/d1a-e2b-coreml JohnP1/d1a-e4b-mlx-q8 JohnP1/d1a-e4b-coreml
Runs on Mac GPU Neural Engine (iPhone, iPad, M4-class Macs) Mac GPU Neural Engine
Text yes yes yes yes
Photo, voice, video (16 frames, no sound) yes not yet (next) yes not yet
Input length long documents ≤ 512 tokens long documents ≤ 512 tokens
Questions yes/no, choice, score, with a calibrated probability for each option same same same
Languages questions in English; documents in English or Japanese same same same
PR labels, 487 newest PRs: type / severity 87.5% / 77.2%; 0 of 30 P0/P1 found same model 90.6% / 79.1%; 13 of 30 P0/P1 same model
Size 3.8 GB + 1.0 GB media 3.5 GB (1.8 GB chunks + 1.6 GB embeddings) 6.6 GB + 1.0 GB media 6.1 GB (3.9 GB chunks + 2.1 GB embeddings)
Median request, M4 (p95) 0.20 s (0.26) 0.96 s (0.97) 0.63 s (0.89) 1.80 s (1.91)
Median request, M1 Max (p95) 0.17 s (0.23) 9.7 s (10.0) ² 0.81 s (1.33) 2.93 s (2.95)
Load, M4 / M1 Max 2 s / 9 s 40 s / 247 s 3 s / – 52 s / 106 s
Answer changes vs fp32 PyTorch ¹ 0 of 37 clear-cut (M4: 1 near-tie of 42) 0 of 37 clear-cut (M4: 1 near-tie of 42) 0 of 38 clear-cut 0 of 38 clear-cut
Photo / voice / video, median (M1 Max) 0.83 s over the 14 demo requests – – –

¹ Each build is compared with the fp32 PyTorch rebuild of the same v0.6 model. "Clear-cut" means the reference's top answer leads by more than 0.05; a near-tie can flip either way. ² On the M1 Max's Neural Engine, the E2B chunks fall back to the CPU: each chunk takes the same time with or without the Neural Engine. The M4 runs the same files on its Neural Engine.

License

Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). D1A is built on Kev by Jared Palmer (Apache-2.0).

Downloads last month
114
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JohnP1/d1a-e4b-coreml

Quantized
(31)
this model

Collection including JohnP1/d1a-e4b-coreml