File size: 5,353 Bytes
7bf8323
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
---
license: apache-2.0
base_model: vllm-sr/Decision-2.0-Eos-0.8B
tags:
- coreml
- decision-making
- on-device
- system-one
---

# Decision-2.0-Eos-0.8B · Core ML

[vllm-sr/Decision-2.0-Eos-0.8B](https://huggingface.co/vllm-sr/Decision-2.0-Eos-0.8B) (vLLM Semantic Router, Apache-2.0)
converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions
about one input, with a probability for every option and no text generation.

This port answers **every question of a request in one model call**. The text all the questions share (the context) is
read once, and each question then continues from it as if it were alone. Eos uses a Qwen3.5 hybrid backbone (Gated
DeltaNet + attention), so "continues from it" covers the DeltaNet recurrent state and conv history as well as the
attention keys. See [How it works](#how-it-works).

| | |
|---|---|
| Source | `vllm-sr/Decision-2.0-Eos-0.8B` @ `3594047d` (Qwen3.5-0.8B-Base backbone + shared decision head) |
| Package | `Decision2EosPacked.mlpackage`, fp16, 1.4 GB, five functions sharing one copy of the weights |
| Functions | `S128_C256_N32`, `S256_C512_N64`, `S512_C768_N96`, `S512_C1024_N128`, `S1024_C1024_N128` (shared-prefix tokens / packed question tokens / option slots) |
| Requires | macOS 15+ / iOS 18+; run on the GPU (`CPU_AND_GPU`) |

Kai 0.6B (Qwen3 backbone) is at [FluidInference/decision-2.0-kai-coreml](https://huggingface.co/FluidInference/decision-2.0-kai-coreml),
with the same Python runtime.

## Quickstart

```bash
pip install coremltools tokenizers numpy
hf download FluidInference/decision-2.0-eos-coreml --local-dir decision-2.0-eos-coreml
cd decision-2.0-eos-coreml && python example.py
```

`decision2_coreml.py` mirrors the upstream `system_one()` output (`choice`/`noul`/`score`, `probabilities`,
`confidence`, `legend`). It picks the smallest function that fits a request. A request that fits none is split into
chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds.

## Fidelity

Against the upstream runtime (fp32, Transformers 5.18, MPS) on the **typed-decisions TEST** split
(`LocalLLaMA/typed-decisions`, 400 requests, 2,000 decisions):

| | Upstream | Core ML fp16 |
|---|---:|---:|
| Choice accuracy (600) | 0.430 | 0.433 |
| Yes/No accuracy (600) | 0.598 | 0.598 |
| Score accuracy (800) | 0.379 | 0.376 |
| Answers that differ from upstream | — | 4 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.003) |
| Max probability difference | — | 0.0077 |

Tokenization (via `tokenizers`) matches the upstream encoder on all 2,000 questions. The fp32 PyTorch version of the
graph (`conversion/eos_graph.py`, `EosChunked`) matches upstream to 2e-6. The accuracies are only for the comparison
with upstream: they are not the model card's benchmark protocol.

## Speed (MacBook Pro M5 Pro, 24 GB, macOS 27)

| Request | Upstream PyTorch, MPS | Core ML (GPU) |
|---|---:|---:|
| Card quickstart, 3 questions | 306 ms | **34 ms** |
| typed-decisions request, 5 questions (median) | 1,595 ms | **110 ms** |
| One email, 64 questions | 15,647 ms | **874 ms** (13 calls) |

Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: S128_C256 34 ms,
S256_C512 64 ms, S512_C768 109 ms, S512_C1024 131 ms, S1024_C1024 176 ms. On a Mac the upstream runtime has none of its
GPU kernels (flash-linear-attention, causal-conv1d), so the speedup is specific to this machine. The Neural Engine was
not used: on Kev-0.8B, which has the same backbone, it was about 24× slower than the GPU.

## How it works

One call takes the shared prefix (right-padded to S) and the questions' suffixes packed end to end (C tokens):

- **Attention layers:** one mask. Prefix tokens are causal; a suffix token sees the real prefix tokens and earlier
  tokens of its own question.
- **Gated DeltaNet layers:** the prefix runs the chunked delta rule (chunk 64; padded positions get β = g = 0, exact
  no-ops), which yields its final recurrent state S₀. The packed questions also run in chunks of 64. A token whose
  question began in an earlier chunk continues from the carried state; a token whose question began in this chunk
  starts from S₀ (`cont`). Decays and the WY inverse are masked to the question within each chunk (`seg_chunks`), and
  only the question of a chunk's last token carries its state forward (`last_seg`). A question's first three
  convolution lags read the prefix's last three inputs (`lag_keep`, `lag_tail`).
- **Readout:** Decision 2.0's head (bilinear + MLP over LayerNorm'd hidden states) at each option's last token and the
  question's final token (`cand_idx`, `query_idx`).

So every question sees exactly its own upstream row (`Context … Question … Options … Decision:`), and the delta-rule
cost stays linear in packed tokens. Full input list: `coreml_config.json`.

## Conversion

`conversion/`: `eos_graph.py` (Decision 2.0 graph on top of `kev_stages.py` / `qwen35_export.py`, the Qwen3.5 export
from [Kev-0.8B Core ML](https://huggingface.co/FluidInference/kev-0.8b-coreml)), `convert_eos.py S cP N` (torch 2.7.0,
coremltools 9.0), `combine.py`, parity and timing scripts.

## License

Apache-2.0, as the source model. `LICENSE`, `tokenizer*.json` and `decision_config.json` are copied unchanged from the
source repository.