|
Download README.md from FluidInference/decision-modernbert-base-coreml: direct link, hf CLI and curl.
- Browser
- Download file 5.81 kB
-
https://huggingface.co/FluidInference/decision-modernbert-base-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/decision-modernbert-base-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/decision-modernbert-base-coreml/resolve/main/README.md
5.81 kB
| license: cc-by-nc-4.0 | |
| base_model: answerdotai/ModernBERT-base | |
| library_name: coreml | |
| pipeline_tag: text-classification | |
| language: | |
| - en | |
| tags: | |
| - typed-decisions | |
| - coreml | |
| - apple-neural-engine | |
| - modernbert | |
| - non-commercial | |
| # decision-modernbert-base (Core ML) | |
| A small typed-decision model for Apple devices: give it a **state** (text or JSON), a **question** (`choice`, | |
| `noul` yes/no, or `score`) and its **options**, and it returns one calibrated probability per option in a few | |
| milliseconds, on device. Fine-tuned from [ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) | |
| (149.6M parameters) and compiled to fixed-shape fp16 Core ML programs. | |
| **License: CC BY-NC 4.0 — non-commercial use only.** The fine-tuning corpus includes sources licensed for | |
| non-commercial research only (see [Training data](#training-data)), so these weights are released for research | |
| and personal use. They are not part of FluidInference's Apache-2.0 model set. The small Python runtime in | |
| `dmodel_mac/` may be used under Apache-2.0. | |
| ## Results | |
| Apple M5 Pro, macOS 27.0. | |
| | | | | |
| |---|---| | |
| | [Decision Index 0.2](https://github.com/apolinario/decision-index) | **20.12** (raw 39.74); every one of 151,034 scored requests answered | | |
| | Areas (knowledge / language / retrieval / tools / arts) | 10.1 / 20.6 / 36.6 / 23.4 / 9.8 | | |
| | Median latency per question (end to end, warm) | 5.6 ms (p95 34 ms) | | |
| | Calibration error (ECE, held-out dev, temperature 1.265) | 0.012 | | |
| | Neural Engine placement | 891 of 896 ops | | |
| | Core ML vs PyTorch fp32 (500 questions) | same answer on 499; max probability difference 0.015 | | |
| The Index score is a self-run with the public decision-index kit (`19ad28ec`) on the hash-verified 0.2 suite; it | |
| is not an official leaderboard entry. For reference, [laya](https://huggingface.co/convaiinnovations/laya) | |
| (421M), which inspired this model, is listed at 5.51. | |
| Strongest benchmarks (chance-corrected skill): BANKING77 81, CLINC150 81, GSM8K 72, HoVer 57, ContractNLI 51, | |
| BFCL 50. It is at chance on 11 of 40, mostly reasoning- and knowledge-heavy sets (ANLI, NLI4CT, CLadder, HLE, | |
| CRUXEval, SATA-Bench, ForecastBench). Good at routing, intent and tool selection; not a general reasoner. | |
| ## Usage | |
| ```bash | |
| pip install coremltools tokenizers numpy | |
| hf download FluidInference/decision-modernbert-base-coreml --local-dir decision-modernbert-base | |
| cd decision-modernbert-base | |
| ``` | |
| ```python | |
| from dmodel_mac.engine import CoreMLDecisionEngine | |
| engine = CoreMLDecisionEngine(config="engine.json") # run from the repo folder; paths are relative | |
| response, raw = engine( | |
| "I was charged twice for one order and do not recognise the second charge.", | |
| {"route": {"type": "choice", "instructions": "Which team should handle this?", | |
| "criteria": {"billing": "Charges, refunds, invoices", | |
| "shipping": "Delivery and tracking", | |
| "account": "Login and profile"}}}) | |
| print(response["answers"]["route"]) # choice 'billing', p ≈ 0.87 | |
| ``` | |
| `engine(state, questions)` follows the Decision Index engine contract: several questions per request, every | |
| option gets a probability, and `noul` answers return `{"noul": p_yes}`. | |
| ## How it works | |
| - One sequence per window: `[CLS] <type> <question> [SEP] ([MASK] <option>)* [SEP] <state slice> [SEP]`. | |
| Each option is scored at its own `[MASK]` marker. | |
| - **Nothing is truncated.** Long questions keep their head and tail, long option lists are split across windows, | |
| and the state is scanned in 25%-overlapping slices; an option's score is the log-mean-exp over every window it | |
| appears in (`dmodel_mac/render.py`). Up to 255 options per question. | |
| - Three Core ML buckets (128, 256, 512 tokens × 64 option slots). `engine.json` runs 128 on CPU+Neural Engine and | |
| 256/512 on all units, the fastest measured placement; softmax uses the calibrated temperature. | |
| | File | | | |
| |---|---| | |
| | `coreml/dmodel_base_L{128,256,512}_K64.mlpackage` | inputs `input_ids`, `attention_mask` int32 `[1,L]`, `marker_map` fp16 `[1,64,L]` → `logits` `[1,64]` | | |
| | `engine.json` | buckets, compute units, temperature | | |
| | `tokenizer.json`, `config.json` | ModernBERT-base tokenizer and config | | |
| | `dmodel_mac/` | reference renderer and engine (Python) | | |
| ## Training data | |
| About 123k rows across 52 tasks, deduplicated and decontaminated against all 155,390 Decision Index 0.2 rows | |
| (exact lines and 13-gram shingles). Gold labels only; no teacher model. Sources include BANKING77, CLINC150, | |
| MASSIVE, MultiNLI, ANLI, FEVER, HoVer, BoolQ, CommonsenseQA, SciQ, QASC, ARC, MMLU, MedMCQA, GSM8K, MATH, | |
| WinoGrande, HellaSwag, ContractNLI, NLI4CT, VAST, iSarcasmEval, Humicroedit, New Yorker captions, ACOS, | |
| Amazon ESCI, MS MARCO, RAGTruth, Enron spam, Twitter financial sentiment, Lichess puzzles, Glaive function | |
| calling, ToolACE, When2Call, HelpSteer 2/3 and UltraFeedback, plus programmatically verified rule tasks. | |
| Several of these (for example ANLI and SciQ under CC BY-NC, MS MARCO's research-only terms) restrict commercial | |
| use, hence this model's license. Upstream train splits of some Index families (ANLI, WinoGrande, HellaSwag, | |
| ContractNLI, VAST, NLI4CT, iSarcasmEval, Humicroedit, RAGTruth, HoVer, ESCI, New Yorker, When2Call, BANKING77, | |
| CLINC150) are in the corpus; their test rows are not. | |
| Full fine-tune on Apple MPS: 2 epochs, AdamW 5e-5, checkpoint chosen by dev macro NLL (dev macro accuracy | |
| 67.6%), one temperature fitted on a separate calibration split. | |
| ## Limitations | |
| - English only. Weak at multi-step reasoning, knowledge-heavy questions and fine detail checks. | |
| - Long inputs cost several windows: p95 latency is about 6× the median. | |
| - Probabilities are calibrated on the training distribution; recalibrate on your own data before relying on them | |
| for new kinds of questions. | |