--- license: cc-by-nc-4.0 base_model: answerdotai/ModernBERT-base library_name: coreml pipeline_tag: text-classification language: - en tags: - typed-decisions - coreml - apple-neural-engine - modernbert - non-commercial --- # decision-modernbert-base (Core ML) A small typed-decision model for Apple devices: give it a **state** (text or JSON), a **question** (`choice`, `noul` yes/no, or `score`) and its **options**, and it returns one calibrated probability per option in a few milliseconds, on device. Fine-tuned from [ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) (149.6M parameters) and compiled to fixed-shape fp16 Core ML programs. **License: CC BY-NC 4.0 — non-commercial use only.** The fine-tuning corpus includes sources licensed for non-commercial research only (see [Training data](#training-data)), so these weights are released for research and personal use. They are not part of FluidInference's Apache-2.0 model set. The small Python runtime in `dmodel_mac/` may be used under Apache-2.0. ## Results Apple M5 Pro, macOS 27.0. | | | |---|---| | [Decision Index 0.2](https://github.com/apolinario/decision-index) | **20.12** (raw 39.74); every one of 151,034 scored requests answered | | Areas (knowledge / language / retrieval / tools / arts) | 10.1 / 20.6 / 36.6 / 23.4 / 9.8 | | Median latency per question (end to end, warm) | 5.6 ms (p95 34 ms) | | Calibration error (ECE, held-out dev, temperature 1.265) | 0.012 | | Neural Engine placement | 891 of 896 ops | | Core ML vs PyTorch fp32 (500 questions) | same answer on 499; max probability difference 0.015 | The Index score is a self-run with the public decision-index kit (`19ad28ec`) on the hash-verified 0.2 suite; it is not an official leaderboard entry. For reference, [laya](https://huggingface.co/convaiinnovations/laya) (421M), which inspired this model, is listed at 5.51. Strongest benchmarks (chance-corrected skill): BANKING77 81, CLINC150 81, GSM8K 72, HoVer 57, ContractNLI 51, BFCL 50. It is at chance on 11 of 40, mostly reasoning- and knowledge-heavy sets (ANLI, NLI4CT, CLadder, HLE, CRUXEval, SATA-Bench, ForecastBench). Good at routing, intent and tool selection; not a general reasoner. ## Usage ```bash pip install coremltools tokenizers numpy hf download FluidInference/decision-modernbert-base-coreml --local-dir decision-modernbert-base cd decision-modernbert-base ``` ```python from dmodel_mac.engine import CoreMLDecisionEngine engine = CoreMLDecisionEngine(config="engine.json") # run from the repo folder; paths are relative response, raw = engine( "I was charged twice for one order and do not recognise the second charge.", {"route": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "Charges, refunds, invoices", "shipping": "Delivery and tracking", "account": "Login and profile"}}}) print(response["answers"]["route"]) # choice 'billing', p ≈ 0.87 ``` `engine(state, questions)` follows the Decision Index engine contract: several questions per request, every option gets a probability, and `noul` answers return `{"noul": p_yes}`. ## How it works - One sequence per window: `[CLS] [SEP] ([MASK]