Decision-2.0-Kai-0.6B · Core ML

vllm-sr/Decision-2.0-Kai-0.6B (vLLM Semantic Router, Apache-2.0) converted to Core ML for on-device use on Apple silicon. Decision 2.0 answers Choice, Yes/No and Score questions about one input, with a probability for every option and no text generation.

This port answers every question of a request in one model call. The text all the questions share (the context) is processed once, and each question's own tokens then attend to that shared prefix and to themselves. Every question therefore sees exactly the prompt it would see alone, as in the upstream runtime.

Source vllm-sr/Decision-2.0-Kai-0.6B @ cd49ea38 (Qwen3-0.6B-Base backbone + shared decision head)
Package Decision2KaiPacked.mlpackage, fp16, 1.1 GB, four functions sharing one copy of the weights
Functions L256_N32, L512_N64, L1024_N128, L2048_N256 (packed tokens / option slots per call)
Requires macOS 15+ / iOS 18+ (multifunction package); run on the GPU (CPU_AND_GPU)

Quickstart

pip install coremltools tokenizers numpy
hf download FluidInference/decision-2.0-kai-coreml --local-dir decision-2.0-kai-coreml
cd decision-2.0-kai-coreml && python example.py
from decision2_coreml import Decision2CoreML

model = Decision2CoreML(".")
model.system_one(state="...", questions={"route": {"type": "choice", "instructions": "...", "criteria": {...}}})

decision2_coreml.py mirrors the upstream system_one() output (choice/noul/score, probabilities, confidence, legend), including the packaged five-level Score offsets (score_bias.json). It picks the smallest function that fits a request. A request too long for one call is split into chunks, each repeating the shared prefix, on whichever function needs the fewest milliseconds.

Fidelity

Against the upstream runtime (fp32, Transformers 5.18) on the typed-decisions TEST split (LocalLLaMA/typed-decisions, 400 requests, 2,000 decisions):

Upstream Core ML fp16
Choice accuracy (600) 0.423 0.422
Yes/No accuracy (600) 0.628 0.628
Score accuracy (800) 0.398 0.398
Answers that differ from upstream — 5 / 2,000 (all near-ties: upstream top-2 margin ≤ 0.009)
Max probability difference — 0.0082

Tokenization (via tokenizers) matches the upstream encoder on all 2,000 questions. The accuracies are only for the comparison with upstream: they are not the model card's benchmark protocol.

Speed (MacBook Pro M5 Pro, 24 GB, macOS 27)

Request Upstream PyTorch, CPU Upstream PyTorch, MPS Core ML (GPU)
Card quickstart, 3 questions 263 ms 89 ms 28 ms
typed-decisions request, 5 questions (median) — 418 ms 61 ms
One email, 64 questions 9,514 ms 3,650 ms 442 ms (7 calls)

Times are full requests (tokenize, pack, predict, answer) and exclude model loading. Per call: L256 14 ms, L512 28 ms, L1024 58 ms, L2048 142 ms. The upstream runtime on a Mac has none of its GPU fast paths (HIP graphs, fused kernels), so treat the speedup as specific to this machine. The Neural Engine was slower (72 ms vs 28 ms on the quickstart set).

Graph inputs

See coreml_config.json. input_ids, position_ids [1, L] int32; mask [1, 1, L, L] fp16 additive (0 allowed, −1e4 blocked); cand_idx, query_idx [N] int32 (each option's last token, and its question's final token). Output logits [N] fp32, one raw logit per option. Unused slots are ignored.

Conversion

Decision-2.0-Eos-0.8B (Qwen3.5 hybrid) is at FluidInference/decision-2.0-eos-coreml, with the same runtime.

conversion/: kai_graph.py (traceable Qwen3 backbone + decision head, loads the source safetensors), convert.py L N (torch 2.7.0, coremltools 9.0), combine.py (multifunction package), pack.py + typed_eval.py + typed_ref.py (parity against the upstream runtime; pack.py imports the source repo's vendored encoder from kai/).

License

Apache-2.0, as the source model. LICENSE, tokenizer*.json, score_bias.json and decision_config.json are copied unchanged from the source repository.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/decision-2.0-kai-coreml

Quantized
(2)
this model