alexwengg's picture
Decision-2.0-Eos-0.8B Core ML: shared prefix + chunked packed questions, fp16 package, runtime, parity reports
7bf8323 verified
|
Raw History Blame Contribute Delete
5.35 kB
---
license: apache-2.0
base_model: vllm-sr/Decision-2.0-Eos-0.8B
tags:
- coreml
- decision-making
- on-device
- system-one
---
# Decision-2.0-Eos-0.8B · Core ML
[vllm-sr/Decision-2.0-Eos-0.8B](https://huggingface.co/vllm-sr/Decision-2.0-Eos-0.8B) (vLLM Semantic Router, Apache-2.0)
converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions
about one input, with a probability for every option and no text generation.
This port answers **every question of a request in one model call**. The text all the questions share (the context) is
read once, and each question then continues from it as if it were alone. Eos uses a Qwen3.5 hybrid backbone (Gated
DeltaNet + attention), so "continues from it" covers the DeltaNet recurrent state and conv history as well as the
attention keys. See [How it works](#how-it-works).
| | |
|---|---|
| Source | `vllm-sr/Decision-2.0-Eos-0.8B` @ `3594047d` (Qwen3.5-0.8B-Base backbone + shared decision head) |
| Package | `Decision2EosPacked.mlpackage`, fp16, 1.4 GB, five functions sharing one copy of the weights |
| Functions | `S128_C256_N32`, `S256_C512_N64`, `S512_C768_N96`, `S512_C1024_N128`, `S1024_C1024_N128` (shared-prefix tokens / packed question tokens / option slots) |
| Requires | macOS 15+ / iOS 18+; run on the GPU (`CPU_AND_GPU`) |
Kai 0.6B (Qwen3 backbone) is at [FluidInference/decision-2.0-kai-coreml](https://huggingface.co/FluidInference/decision-2.0-kai-coreml),
with the same Python runtime.
## Quickstart
```bash
pip install coremltools tokenizers numpy
hf download FluidInference/decision-2.0-eos-coreml --local-dir decision-2.0-eos-coreml
cd decision-2.0-eos-coreml && python example.py
```
`decision2_coreml.py` mirrors the upstream `system_one()` output (`choice`/`noul`/`score`, `probabilities`,
`confidence`, `legend`). It picks the smallest function that fits a request. A request that fits none is split into
chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds.
## Fidelity
Against the upstream runtime (fp32, Transformers 5.18, MPS) on the **typed-decisions TEST** split
(`LocalLLaMA/typed-decisions`, 400 requests, 2,000 decisions):
| | Upstream | Core ML fp16 |
|---|---:|---:|
| Choice accuracy (600) | 0.430 | 0.433 |
| Yes/No accuracy (600) | 0.598 | 0.598 |
| Score accuracy (800) | 0.379 | 0.376 |
| Answers that differ from upstream | — | 4 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.003) |
| Max probability difference | — | 0.0077 |
Tokenization (via `tokenizers`) matches the upstream encoder on all 2,000 questions. The fp32 PyTorch version of the
graph (`conversion/eos_graph.py`, `EosChunked`) matches upstream to 2e-6. The accuracies are only for the comparison
with upstream: they are not the model card's benchmark protocol.
## Speed (MacBook Pro M5 Pro, 24 GB, macOS 27)
| Request | Upstream PyTorch, MPS | Core ML (GPU) |
|---|---:|---:|
| Card quickstart, 3 questions | 306 ms | **34 ms** |
| typed-decisions request, 5 questions (median) | 1,595 ms | **110 ms** |
| One email, 64 questions | 15,647 ms | **874 ms** (13 calls) |
Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: S128_C256 34 ms,
S256_C512 64 ms, S512_C768 109 ms, S512_C1024 131 ms, S1024_C1024 176 ms. On a Mac the upstream runtime has none of its
GPU kernels (flash-linear-attention, causal-conv1d), so the speedup is specific to this machine. The Neural Engine was
not used: on Kev-0.8B, which has the same backbone, it was about 24× slower than the GPU.
## How it works
One call takes the shared prefix (right-padded to S) and the questions' suffixes packed end to end (C tokens):
- **Attention layers:** one mask. Prefix tokens are causal; a suffix token sees the real prefix tokens and earlier
tokens of its own question.
- **Gated DeltaNet layers:** the prefix runs the chunked delta rule (chunk 64; padded positions get β = g = 0, exact
no-ops), which yields its final recurrent state S₀. The packed questions also run in chunks of 64. A token whose
question began in an earlier chunk continues from the carried state; a token whose question began in this chunk
starts from S₀ (`cont`). Decays and the WY inverse are masked to the question within each chunk (`seg_chunks`), and
only the question of a chunk's last token carries its state forward (`last_seg`). A question's first three
convolution lags read the prefix's last three inputs (`lag_keep`, `lag_tail`).
- **Readout:** Decision 2.0's head (bilinear + MLP over LayerNorm'd hidden states) at each option's last token and the
question's final token (`cand_idx`, `query_idx`).
So every question sees exactly its own upstream row (`Context … Question … Options … Decision:`), and the delta-rule
cost stays linear in packed tokens. Full input list: `coreml_config.json`.
## Conversion
`conversion/`: `eos_graph.py` (Decision 2.0 graph on top of `kev_stages.py` / `qwen35_export.py`, the Qwen3.5 export
from [Kev-0.8B Core ML](https://huggingface.co/FluidInference/kev-0.8b-coreml)), `convert_eos.py S cP N` (torch 2.7.0,
coremltools 9.0), `combine.py`, parity and timing scripts.
## License
Apache-2.0, as the source model. `LICENSE`, `tokenizer*.json` and `decision_config.json` are copied unchanged from the
source repository.