Decision-2.0-Eos-0.8B · Core ML

vllm-sr/Decision-2.0-Eos-0.8B (vLLM Semantic Router, Apache-2.0) converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions about one input, with a probability for every option and no text generation.

This port answers every question of a request in one model call. The text all the questions share (the context) is read once, and each question then continues from it as if it were alone. Eos uses a Qwen3.5 hybrid backbone (Gated DeltaNet + attention), so "continues from it" covers the DeltaNet recurrent state and conv history as well as the attention keys. See How it works.

Source vllm-sr/Decision-2.0-Eos-0.8B @ 3594047d (Qwen3.5-0.8B-Base backbone + shared decision head)
Package Decision2EosPacked.mlpackage, fp16, 1.4 GB, five functions sharing one copy of the weights
Functions S128_C256_N32, S256_C512_N64, S512_C768_N96, S512_C1024_N128, S1024_C1024_N128 (shared-prefix tokens / packed question tokens / option slots)
Requires macOS 15+ / iOS 18+; run on the GPU (CPU_AND_GPU)

Kai 0.6B (Qwen3 backbone) is at FluidInference/decision-2.0-kai-coreml, with the same Python runtime.

Quickstart

pip install coremltools tokenizers numpy
hf download FluidInference/decision-2.0-eos-coreml --local-dir decision-2.0-eos-coreml
cd decision-2.0-eos-coreml && python example.py

decision2_coreml.py mirrors the upstream system_one() output (choice/noul/score, probabilities, confidence, legend). It picks the smallest function that fits a request. A request that fits none is split into chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds.

Fidelity

Against the upstream runtime (fp32, Transformers 5.18, MPS) on the typed-decisions TEST split (LocalLLaMA/typed-decisions, 400 requests, 2,000 decisions):

Upstream Core ML fp16
Choice accuracy (600) 0.430 0.433
Yes/No accuracy (600) 0.598 0.598
Score accuracy (800) 0.379 0.376
Answers that differ from upstream — 4 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.003)
Max probability difference — 0.0077

Tokenization (via tokenizers) matches the upstream encoder on all 2,000 questions. The fp32 PyTorch version of the graph (conversion/eos_graph.py, EosChunked) matches upstream to 2e-6. The accuracies are only for the comparison with upstream: they are not the model card's benchmark protocol.

Speed (MacBook Pro M5 Pro, 24 GB, macOS 27)

Request Upstream PyTorch, MPS Core ML (GPU)
Card quickstart, 3 questions 306 ms 34 ms
typed-decisions request, 5 questions (median) 1,595 ms 110 ms
One email, 64 questions 15,647 ms 874 ms (13 calls)

Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: S128_C256 34 ms, S256_C512 64 ms, S512_C768 109 ms, S512_C1024 131 ms, S1024_C1024 176 ms. On a Mac the upstream runtime has none of its GPU kernels (flash-linear-attention, causal-conv1d), so the speedup is specific to this machine. The Neural Engine was not used: on Kev-0.8B, which has the same backbone, it was about 24× slower than the GPU.

How it works

One call takes the shared prefix (right-padded to S) and the questions' suffixes packed end to end (C tokens):

  • Attention layers: one mask. Prefix tokens are causal; a suffix token sees the real prefix tokens and earlier tokens of its own question.
  • Gated DeltaNet layers: the prefix runs the chunked delta rule (chunk 64; padded positions get β = g = 0, exact no-ops), which yields its final recurrent state S₀. The packed questions also run in chunks of 64. A token whose question began in an earlier chunk continues from the carried state; a token whose question began in this chunk starts from S₀ (cont). Decays and the WY inverse are masked to the question within each chunk (seg_chunks), and only the question of a chunk's last token carries its state forward (last_seg). A question's first three convolution lags read the prefix's last three inputs (lag_keep, lag_tail).
  • Readout: Decision 2.0's head (bilinear + MLP over LayerNorm'd hidden states) at each option's last token and the question's final token (cand_idx, query_idx).

So every question sees exactly its own upstream row (Context … Question … Options … Decision:), and the delta-rule cost stays linear in packed tokens. Full input list: coreml_config.json.

Conversion

conversion/: eos_graph.py (Decision 2.0 graph on top of kev_stages.py / qwen35_export.py, the Qwen3.5 export from Kev-0.8B Core ML), convert_eos.py S cP N (torch 2.7.0, coremltools 9.0), combine.py, parity and timing scripts.

License

Apache-2.0, as the source model. LICENSE, tokenizer*.json and decision_config.json are copied unchanged from the source repository.

Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/decision-2.0-eos-coreml

Quantized
(2)
this model