|
Download README.md from FluidInference/decision-2.0-eos-coreml: direct link, hf CLI and curl.
- Browser
- Download file 5.35 kB
-
https://huggingface.co/FluidInference/decision-2.0-eos-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/decision-2.0-eos-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/decision-2.0-eos-coreml/resolve/main/README.md
5.35 kB
| license: apache-2.0 | |
| base_model: vllm-sr/Decision-2.0-Eos-0.8B | |
| tags: | |
| - coreml | |
| - decision-making | |
| - on-device | |
| - system-one | |
| # Decision-2.0-Eos-0.8B · Core ML | |
| [vllm-sr/Decision-2.0-Eos-0.8B](https://huggingface.co/vllm-sr/Decision-2.0-Eos-0.8B) (vLLM Semantic Router, Apache-2.0) | |
| converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions | |
| about one input, with a probability for every option and no text generation. | |
| This port answers **every question of a request in one model call**. The text all the questions share (the context) is | |
| read once, and each question then continues from it as if it were alone. Eos uses a Qwen3.5 hybrid backbone (Gated | |
| DeltaNet + attention), so "continues from it" covers the DeltaNet recurrent state and conv history as well as the | |
| attention keys. See [How it works](#how-it-works). | |
| | | | | |
| |---|---| | |
| | Source | `vllm-sr/Decision-2.0-Eos-0.8B` @ `3594047d` (Qwen3.5-0.8B-Base backbone + shared decision head) | | |
| | Package | `Decision2EosPacked.mlpackage`, fp16, 1.4 GB, five functions sharing one copy of the weights | | |
| | Functions | `S128_C256_N32`, `S256_C512_N64`, `S512_C768_N96`, `S512_C1024_N128`, `S1024_C1024_N128` (shared-prefix tokens / packed question tokens / option slots) | | |
| | Requires | macOS 15+ / iOS 18+; run on the GPU (`CPU_AND_GPU`) | | |
| Kai 0.6B (Qwen3 backbone) is at [FluidInference/decision-2.0-kai-coreml](https://huggingface.co/FluidInference/decision-2.0-kai-coreml), | |
| with the same Python runtime. | |
| ## Quickstart | |
| ```bash | |
| pip install coremltools tokenizers numpy | |
| hf download FluidInference/decision-2.0-eos-coreml --local-dir decision-2.0-eos-coreml | |
| cd decision-2.0-eos-coreml && python example.py | |
| ``` | |
| `decision2_coreml.py` mirrors the upstream `system_one()` output (`choice`/`noul`/`score`, `probabilities`, | |
| `confidence`, `legend`). It picks the smallest function that fits a request. A request that fits none is split into | |
| chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds. | |
| ## Fidelity | |
| Against the upstream runtime (fp32, Transformers 5.18, MPS) on the **typed-decisions TEST** split | |
| (`LocalLLaMA/typed-decisions`, 400 requests, 2,000 decisions): | |
| | | Upstream | Core ML fp16 | | |
| |---|---:|---:| | |
| | Choice accuracy (600) | 0.430 | 0.433 | | |
| | Yes/No accuracy (600) | 0.598 | 0.598 | | |
| | Score accuracy (800) | 0.379 | 0.376 | | |
| | Answers that differ from upstream | — | 4 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.003) | | |
| | Max probability difference | — | 0.0077 | | |
| Tokenization (via `tokenizers`) matches the upstream encoder on all 2,000 questions. The fp32 PyTorch version of the | |
| graph (`conversion/eos_graph.py`, `EosChunked`) matches upstream to 2e-6. The accuracies are only for the comparison | |
| with upstream: they are not the model card's benchmark protocol. | |
| ## Speed (MacBook Pro M5 Pro, 24 GB, macOS 27) | |
| | Request | Upstream PyTorch, MPS | Core ML (GPU) | | |
| |---|---:|---:| | |
| | Card quickstart, 3 questions | 306 ms | **34 ms** | | |
| | typed-decisions request, 5 questions (median) | 1,595 ms | **110 ms** | | |
| | One email, 64 questions | 15,647 ms | **874 ms** (13 calls) | | |
| Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: S128_C256 34 ms, | |
| S256_C512 64 ms, S512_C768 109 ms, S512_C1024 131 ms, S1024_C1024 176 ms. On a Mac the upstream runtime has none of its | |
| GPU kernels (flash-linear-attention, causal-conv1d), so the speedup is specific to this machine. The Neural Engine was | |
| not used: on Kev-0.8B, which has the same backbone, it was about 24× slower than the GPU. | |
| ## How it works | |
| One call takes the shared prefix (right-padded to S) and the questions' suffixes packed end to end (C tokens): | |
| - **Attention layers:** one mask. Prefix tokens are causal; a suffix token sees the real prefix tokens and earlier | |
| tokens of its own question. | |
| - **Gated DeltaNet layers:** the prefix runs the chunked delta rule (chunk 64; padded positions get β = g = 0, exact | |
| no-ops), which yields its final recurrent state S₀. The packed questions also run in chunks of 64. A token whose | |
| question began in an earlier chunk continues from the carried state; a token whose question began in this chunk | |
| starts from S₀ (`cont`). Decays and the WY inverse are masked to the question within each chunk (`seg_chunks`), and | |
| only the question of a chunk's last token carries its state forward (`last_seg`). A question's first three | |
| convolution lags read the prefix's last three inputs (`lag_keep`, `lag_tail`). | |
| - **Readout:** Decision 2.0's head (bilinear + MLP over LayerNorm'd hidden states) at each option's last token and the | |
| question's final token (`cand_idx`, `query_idx`). | |
| So every question sees exactly its own upstream row (`Context … Question … Options … Decision:`), and the delta-rule | |
| cost stays linear in packed tokens. Full input list: `coreml_config.json`. | |
| ## Conversion | |
| `conversion/`: `eos_graph.py` (Decision 2.0 graph on top of `kev_stages.py` / `qwen35_export.py`, the Qwen3.5 export | |
| from [Kev-0.8B Core ML](https://huggingface.co/FluidInference/kev-0.8b-coreml)), `convert_eos.py S cP N` (torch 2.7.0, | |
| coremltools 9.0), `combine.py`, parity and timing scripts. | |
| ## License | |
| Apache-2.0, as the source model. `LICENSE`, `tokenizer*.json` and `decision_config.json` are copied unchanged from the | |
| source repository. | |