D1A-E4B v0.6 · Core ML (text, ≤ 512 tokens)
D1A-E4B v0.6 for the Apple Neural Engine. D1A is a small decision model: a document and typed questions go in, and a calibrated probability for every option comes out, in one pass, with no generated text (jonpol01/d1a). This repository is the text model converted to Core ML; the conversion is described in jonpol01/d1a#224.
Status
- Correct at ≤ 512 tokens. On 28 real requests (42 questions), no clear-cut answer changes against the fp32 PyTorch rebuild of the same model, on an M1 Max and an M4.
- Speed (median request): M4 Neural Engine 1.8 s (MLX on the GPU: 0.63 s); M1 Max 2.9 s (MLX: 0.81 s).
- Use it where the GPU should stay free or power matters (iPhone, iPad). On a Mac, the MLX build is faster.
- Not included yet: photo, voice and video (their encoders are separate models, still to convert), and inputs longer than 512 tokens (2,048-token programs give wrong answers on the M4 Neural Engine).
Run it
pip install coremltools numpy torch transformers safetensors "d1a @ git+https://github.com/jonpol01/d1a"
hf download JohnP1/d1a-e4b-coreml --revision v0.6 --local-dir d1a-e4b-coreml
python d1a-e4b-coreml/reference/run.py d1a-e4b-coreml request.json
request.json is a System One request, for example:
{"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"returns": null, "shipping": null, "billing": null}},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"}}}
The first run compiles the chunks into coreml/.compiled/ (a few minutes); later runs load them.
Files
| Path | What |
|---|---|
coreml/L512_<a>-<b>.mlpackage |
7 chunks of the 42-layer text model, int8 weights (per block of 64). Chunk 18-24 outputs the shared keys and values that chunks 24-42 read. |
embeddings.safetensors |
the token and per-layer embedding tables (MLX affine-quantized, from JohnP1/d1a-e4b-mlx-q8@v0.6), read row by row on the host |
head.safetensors, d1a_config.json |
the pointer head, the temperature and the pad and delimiter ids |
config.json, tokenizer.json, tokenizer_config.json |
the Gemma 4 model config and tokenizer |
io.json |
each chunk's input names |
reference/ |
Python reference: run.py (one request), workers.py (one process per chunk), d1a_coreml.py (host-side embeddings and masks) |
How a request runs
- Encode it with D1A's layout (
d1a.core.encoding.encode) and right-pad it to 512 tokens with the pad token. - Look up the token embeddings (scaled by √hidden) and the per-layer embeddings (scaled by √256) on the host.
- Build the block-causal masks (full and sliding), clamped to −1e4.
- Run the chunks in order, each in its own process. Several large Neural Engine programs in one process can crash the runtime.
- Read the hidden states at each question's
<decide>and</opt>positions, score them with the head, divide by the temperature, and take the softmax.
macOS 15 / iOS 18 or later.
D1A builds compared
Measured on 2026-10-11 on the same 28 real requests (42 questions, all ≤ 512 tokens), on an Apple M1 Max and a Mac mini M4.
| E2B · MLX | E2B · Core ML | E4B · MLX | E4B · Core ML | |
|---|---|---|---|---|
| Version | v0.6 | v0.6 | v0.6 | v0.6 |
| Repository | JohnP1/d1a-e2b-mlx-q8 | JohnP1/d1a-e2b-coreml | JohnP1/d1a-e4b-mlx-q8 | JohnP1/d1a-e4b-coreml |
| Runs on | Mac GPU | Neural Engine (iPhone, iPad, M4-class Macs) | Mac GPU | Neural Engine |
| Text | yes | yes | yes | yes |
| Photo, voice, video (16 frames, no sound) | yes | not yet (next) | yes | not yet |
| Input length | long documents | ≤ 512 tokens | long documents | ≤ 512 tokens |
| Questions | yes/no, choice, score, with a calibrated probability for each option | same | same | same |
| Languages | questions in English; documents in English or Japanese | same | same | same |
| PR labels, 487 newest PRs: type / severity | 87.5% / 77.2%; 0 of 30 P0/P1 found | same model | 90.6% / 79.1%; 13 of 30 P0/P1 | same model |
| Size | 3.8 GB + 1.0 GB media | 3.5 GB (1.8 GB chunks + 1.6 GB embeddings) | 6.6 GB + 1.0 GB media | 6.1 GB (3.9 GB chunks + 2.1 GB embeddings) |
| Median request, M4 (p95) | 0.20 s (0.26) | 0.96 s (0.97) | 0.63 s (0.89) | 1.80 s (1.91) |
| Median request, M1 Max (p95) | 0.17 s (0.23) | 9.7 s (10.0) ² | 0.81 s (1.33) | 2.93 s (2.95) |
| Load, M4 / M1 Max | 2 s / 9 s | 40 s / 247 s | 3 s / – | 52 s / 106 s |
| Answer changes vs fp32 PyTorch ¹ | 0 of 37 clear-cut (M4: 1 near-tie of 42) | 0 of 37 clear-cut (M4: 1 near-tie of 42) | 0 of 38 clear-cut | 0 of 38 clear-cut |
| Photo / voice / video, median (M1 Max) | 0.83 s over the 14 demo requests | – | – | – |
¹ Each build is compared with the fp32 PyTorch rebuild of the same v0.6 model. "Clear-cut" means the reference's top answer leads by more than 0.05; a near-tie can flip either way. ² On the M1 Max's Neural Engine, the E2B chunks fall back to the CPU: each chunk takes the same time with or without the Neural Engine. The M4 runs the same files on its Neural Engine.
License
Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). D1A is built on Kev by Jared Palmer (Apache-2.0).
- Downloads last month
- 114
Model tree for JohnP1/d1a-e4b-coreml
Base model
google/gemma-4-E4B