Decision-2.0-Kai-0.6B · Core ML
vllm-sr/Decision-2.0-Kai-0.6B (vLLM Semantic Router, Apache-2.0) converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions about one input, with a probability for every option and no text generation.
This port answers every question of a request in one model call. The text all the questions share (the context) is processed once, and each question's own tokens then attend to that shared prefix and to themselves. Every question therefore sees exactly the prompt it would see alone, as in the upstream runtime.
| Source | vllm-sr/Decision-2.0-Kai-0.6B @ cd49ea38 (Qwen3-0.6B-Base backbone + shared decision head) |
| Package | Decision2KaiPacked.mlpackage, fp16, 1.1 GB, four functions sharing one copy of the weights |
| Functions | L256_N32, L512_N64, L1024_N128, L2048_N256 (packed tokens / option slots per call) |
| Requires | macOS 15+ / iOS 18+ (multifunction package); run on the GPU (CPU_AND_GPU) |
Quickstart
pip install coremltools tokenizers numpy
hf download FluidInference/decision-2.0-kai-coreml --local-dir decision-2.0-kai-coreml
cd decision-2.0-kai-coreml && python example.py
from decision2_coreml import Decision2CoreML
model = Decision2CoreML(".")
model.system_one(state="...", questions={"route": {"type": "choice", "instructions": "...", "criteria": {...}}})
decision2_coreml.py mirrors the upstream system_one() output (choice/noul/score, probabilities,
confidence, legend), including the packaged five-level Score offsets (score_bias.json). It picks the smallest
function that fits a request. A request too long for one call is split into chunks, each repeating the shared prefix,
on whichever function needs the fewest milliseconds.
Fidelity
Against the upstream runtime (fp32, Transformers 5.18) on the typed-decisions TEST split
(LocalLLaMA/typed-decisions, 400 requests, 2,000 decisions):
| Upstream | Core ML fp16 | |
|---|---|---|
| Choice accuracy (600) | 0.423 | 0.422 |
| Yes/No accuracy (600) | 0.628 | 0.628 |
| Score accuracy (800) | 0.398 | 0.398 |
| Answers that differ from upstream | — | 5 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.009) |
| Max probability difference | — | 0.0082 |
Tokenization (via tokenizers) matches the upstream encoder on all 2,000 questions. The accuracies are only for the
comparison with upstream: they are not the model card's benchmark protocol.
Speed (MacBook Pro M5 Pro, 24 GB, macOS 27)
| Request | Upstream PyTorch, CPU | Upstream PyTorch, MPS | Core ML (GPU) |
|---|---|---|---|
| Card quickstart, 3 questions | 263 ms | 89 ms | 28 ms |
| typed-decisions request, 5 questions (median) | — | 418 ms | 61 ms |
| One email, 64 questions | 9,514 ms | 3,650 ms | 442 ms (7 calls) |
Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: L256 14 ms, L512 28 ms, L1024 58 ms, L2048 142 ms. The upstream runtime on a Mac has none of its GPU fast paths (HIP graphs, fused kernels), so treat the speedup as specific to this machine. The Neural Engine was slower (72 ms vs 28 ms on the quickstart set).
Graph inputs
See coreml_config.json. input_ids, position_ids [1, L] int32; mask [1, 1, L, L] fp16 additive (0 allowed,
−1e4 blocked); cand_idx, query_idx [N] int32 (each option's last token, and its question's final token). Output
logits [N] fp32, one raw logit per option. Unused slots are ignored.
Conversion
Decision-2.0-Eos-0.8B (Qwen3.5 hybrid) is at FluidInference/decision-2.0-eos-coreml, with the same runtime.
conversion/: kai_graph.py (traceable Qwen3 backbone + decision head, loads the source safetensors),
convert.py L N (torch 2.7.0, coremltools 9.0), combine.py (multifunction package), pack.py + typed_eval.py +
typed_ref.py (parity against the upstream runtime; pack.py imports the source repo's vendored encoder from kai/).
License
Apache-2.0, as the source model. LICENSE, tokenizer*.json, score_bias.json and decision_config.json are copied
unchanged from the source repository.
- Downloads last month
- 9