--- license: apache-2.0 base_model: vllm-sr/Decision-2.0-Eos-0.8B tags: - coreml - decision-making - on-device - system-one --- # Decision-2.0-Eos-0.8B · Core ML [vllm-sr/Decision-2.0-Eos-0.8B](https://huggingface.co/vllm-sr/Decision-2.0-Eos-0.8B) (vLLM Semantic Router, Apache-2.0) converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions about one input, with a probability for every option and no text generation. This port answers **every question of a request in one model call**. The text all the questions share (the context) is read once, and each question then continues from it as if it were alone. Eos uses a Qwen3.5 hybrid backbone (Gated DeltaNet + attention), so "continues from it" covers the DeltaNet recurrent state and conv history as well as the attention keys. See [How it works](#how-it-works). | | | |---|---| | Source | `vllm-sr/Decision-2.0-Eos-0.8B` @ `3594047d` (Qwen3.5-0.8B-Base backbone + shared decision head) | | Package | `Decision2EosPacked.mlpackage`, fp16, 1.4 GB, five functions sharing one copy of the weights | | Functions | `S128_C256_N32`, `S256_C512_N64`, `S512_C768_N96`, `S512_C1024_N128`, `S1024_C1024_N128` (shared-prefix tokens / packed question tokens / option slots) | | Requires | macOS 15+ / iOS 18+; run on the GPU (`CPU_AND_GPU`) | Kai 0.6B (Qwen3 backbone) is at [FluidInference/decision-2.0-kai-coreml](https://huggingface.co/FluidInference/decision-2.0-kai-coreml), with the same Python runtime. ## Quickstart ```bash pip install coremltools tokenizers numpy hf download FluidInference/decision-2.0-eos-coreml --local-dir decision-2.0-eos-coreml cd decision-2.0-eos-coreml && python example.py ``` `decision2_coreml.py` mirrors the upstream `system_one()` output (`choice`/`noul`/`score`, `probabilities`, `confidence`, `legend`). It picks the smallest function that fits a request. A request that fits none is split into chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds. ## Fidelity Against the upstream runtime (fp32, Transformers 5.18, MPS) on the **typed-decisions TEST** split (`LocalLLaMA/typed-decisions`, 400 requests, 2,000 decisions): | | Upstream | Core ML fp16 | |---|---:|---:| | Choice accuracy (600) | 0.430 | 0.433 | | Yes/No accuracy (600) | 0.598 | 0.598 | | Score accuracy (800) | 0.379 | 0.376 | | Answers that differ from upstream | — | 4 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.003) | | Max probability difference | — | 0.0077 | Tokenization (via `tokenizers`) matches the upstream encoder on all 2,000 questions. The fp32 PyTorch version of the graph (`conversion/eos_graph.py`, `EosChunked`) matches upstream to 2e-6. The accuracies are only for the comparison with upstream: they are not the model card's benchmark protocol. ## Speed (MacBook Pro M5 Pro, 24 GB, macOS 27) | Request | Upstream PyTorch, MPS | Core ML (GPU) | |---|---:|---:| | Card quickstart, 3 questions | 306 ms | **34 ms** | | typed-decisions request, 5 questions (median) | 1,595 ms | **110 ms** | | One email, 64 questions | 15,647 ms | **874 ms** (13 calls) | Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: S128_C256 34 ms, S256_C512 64 ms, S512_C768 109 ms, S512_C1024 131 ms, S1024_C1024 176 ms. On a Mac the upstream runtime has none of its GPU kernels (flash-linear-attention, causal-conv1d), so the speedup is specific to this machine. The Neural Engine was not used: on Kev-0.8B, which has the same backbone, it was about 24× slower than the GPU. ## How it works One call takes the shared prefix (right-padded to S) and the questions' suffixes packed end to end (C tokens): - **Attention layers:** one mask. Prefix tokens are causal; a suffix token sees the real prefix tokens and earlier tokens of its own question. - **Gated DeltaNet layers:** the prefix runs the chunked delta rule (chunk 64; padded positions get β = g = 0, exact no-ops), which yields its final recurrent state S₀. The packed questions also run in chunks of 64. A token whose question began in an earlier chunk continues from the carried state; a token whose question began in this chunk starts from S₀ (`cont`). Decays and the WY inverse are masked to the question within each chunk (`seg_chunks`), and only the question of a chunk's last token carries its state forward (`last_seg`). A question's first three convolution lags read the prefix's last three inputs (`lag_keep`, `lag_tail`). - **Readout:** Decision 2.0's head (bilinear + MLP over LayerNorm'd hidden states) at each option's last token and the question's final token (`cand_idx`, `query_idx`). So every question sees exactly its own upstream row (`Context … Question … Options … Decision:`), and the delta-rule cost stays linear in packed tokens. Full input list: `coreml_config.json`. ## Conversion `conversion/`: `eos_graph.py` (Decision 2.0 graph on top of `kev_stages.py` / `qwen35_export.py`, the Qwen3.5 export from [Kev-0.8B Core ML](https://huggingface.co/FluidInference/kev-0.8b-coreml)), `convert_eos.py S cP N` (torch 2.7.0, coremltools 9.0), `combine.py`, parity and timing scripts. ## License Apache-2.0, as the source model. `LICENSE`, `tokenizer*.json` and `decision_config.json` are copied unchanged from the source repository.