--- license: apache-2.0 base_model: internlm/Intern-Decision-0.8B library_name: coreml pipeline_tag: text-classification tags: - coreml - apple-silicon - decision - qwen3.5 --- # Intern-Decision-0.8B for Core ML Core ML export of [internlm/Intern-Decision-0.8B](https://huggingface.co/internlm/Intern-Decision-0.8B) (Shanghai AI Laboratory, Apache-2.0, snapshot `85a0cc5a`): a typed-decision model fine-tuned from Qwen3.5-0.8B that answers a set of `choice` / `score` / `noul` (yes/no) questions about a JSON state in one prefill pass. The text path is exported here; the vision tower is not (text-only requests). ## Files | Path | What | | --- | --- | | `L320_F8/DecisionRow_fp16.mlpackage` | Requests up to 320 tokens and 8 fields (the model card's three-field shape is 319 tokens). | | `L512_F8/DecisionRow_fp16.mlpackage` | Up to 512 tokens, 8 fields. | | `L512_F8/DecisionRow_w8.mlpackage` | Same, int8 per-channel weights (480 MB; 2 of 240 answers differ from the reference). | | `L1024_F16/DecisionRow_fp16.mlpackage` | Up to 1,024 tokens, 16 fields. | | `*/config.json` | Bucket dimensions, marker / pad / answer-symbol ids, temperature, system prompt. | | `embeddings.f16` | Token embeddings (fp16, 248,320 × 1,024), gathered on the host. | | `tokenizer.json` | The checkpoint's tokenizer (Qwen3.5 plus the `` token). | Inputs: `hidden` [1, L, 1024] (embedding rows of the right-padded prompt), `cos` / `sin` [L, 64] (RoPE tables for positions 0…L−1), `field_onehot` [F, L] (row *i* selects the token before field *i*'s ``). Output: `logits` [F, 62] over the answer symbols; take the first *n* entries for a field with *n* options, softmax, divide log-probs by the temperature. fp16, GPU (`cpuAndGPU`), iOS 17 / macOS 14. Swift runtime: `InternDecisionManager` in [FluidUse](https://github.com/FluidInference/FluidUse); conversion scripts in [mobius](https://github.com/FluidInference/mobius) (`models/computer-use/intern-decision-0.8b/coreml`). ## What the package computes Intern-Decision renders the request as one chat prompt: a fixed system prompt, a user turn with the state as JSON and the decision schema (one line per field, its options mapped to the answer symbols `A`–`Z`, `a`–`z`, `0`–`9`), and an assistant turn that is a JSON skeleton with one `` token per field. The logits at the position immediately before each marker, restricted to that field's first *n* symbols, are its answer; the published temperature (2.7478 for 0.8B) rescales the restricted softmax without changing the argmax. Nothing is generated. `DecisionRow` (`decision_export.py`) is one fixed-length request: the Qwen3.5 decoder from `qwen35_export.py` (the Kev-0.8B / Cua-S1-4B export), the final norm at the pre-marker positions selected with a one-hot map per field, and the 62 tied-embedding rows of the answer symbols. Token embeddings are gathered on the host (`embeddings.f16`); the host applies the restricted softmax and temperature. Prompt compilation and the chat template come from the checkpoint's own `inference.py` and tokenizer, so the token stream is the reference's. | Package | Tokens | Fields | Fits | | --- | ---: | ---: | --- | | `L256_F4` | 256 | 4 | Jevbench easy/original (~220 tokens), AG News (p95 299) | | `L320_F8` | 320 | 8 | the model card's three-field request shape | | `L384_F8`, `L512_F8` | 384 / 512 | 8 | WildJailBreak (p95 536), 62% of ToolACE | | `L768_F16`, `L1024_F16` | 768 / 1,024 | 16 | Typed Decision (5 fields, p50 852, max 1,223), ToolACE (max 886) | The fixed prompt (system prompt, headings, skeleton) is about 250 tokens, so the smallest three-field request is 319 tokens; 40% of Jevbench-Hard exceeds 1,024 tokens (max 4,074). ## Fidelity Reference: the checkpoint's `DecisionEngine` in fp32 on the Apple GPU (MPS), probabilities after temperature scaling, on the bundled accuracy suites from the [Intern-Decision repo](https://github.com/InternLM/Intern-Decision) (`benchmarks/accuracy-v1`, shuffled with seed 0). The fp32 PyTorch wrapper matches the reference to 3.6e-6 (50 Typed Decision fields). | Package | Suites | Decisions | Top answer differs | Max \|Δp\| | | --- | --- | ---: | ---: | ---: | | `L512_F8` fp16 | Jevbench (3), ToolACE, AG News, WildJailBreak | 240 | 0 | 0.006 | | `L1024_F16` fp16 | Typed Decision (5 fields), Jevbench-Hard, ToolACE | 280 | 0 | 0.007 | | `L320_F8` fp16 | Jevbench easy/original, AG News, WildJailBreak | 100 | 0 | 0.006 | | `L256_F4` fp16 | Jevbench easy/original, AG News, WildJailBreak | 100 | 0 | 0.006 | | `L512_F8` int8 (per-channel) | same as `L512_F8` fp16 | 240 | 2 | 0.047 | Reference and Core ML accuracy against the suite labels are identical on every subset (reports in the mobius directory). int4 per-block compression needs an iOS 18 deployment target and was not built. ## Latency One request, 319 input tokens, three fields (choice, yes/no, score), the model card's RTX 4090 shape. M5 Pro (24 GB), warmed, `bench.py`; Core ML on `CPU_AND_GPU`, times include the host embedding gather and RoPE tables. | Runtime | p50 | p95 | | --- | ---: | ---: | | Core ML fp16 `L320_F8` | 57 ms | 59 ms | | Core ML fp16 `L384_F8` | 66 ms | 67 ms | | Core ML fp16 `L512_F8` | 88 ms | 96 ms | | Core ML fp16 `L768_F16` | 133 ms | 145 ms | | Core ML fp16 `L1024_F16` | 182 ms | 190 ms | | PyTorch MPS bf16 (checkpoint `inference.py`) | 150 ms | 170 ms | | PyTorch MPS fp32 | 177 ms | 191 ms | The pass is compute-bound and scales with the bucket, not the request, so ship the smallest bucket that fits. int8 weights leave GPU time unchanged (88 ms at `L512_F8`) and halve the package (955 → 480 MB). `ComputeUnit.ALL` matches `CPU_AND_GPU`: the Gated DeltaNet backbone does not run on the Neural Engine (see the Kev-0.8B notes). The model card reports 34 ms for this request on an RTX 4090.