MIMIR-1

Decisions your agents can act on. MIMIR is an English-only, non-generative decision model: give it a context, a question and the options, and it returns a typed answer with calibrated probabilities, the parts of the context it relied on, and a certified signal for when to act and when to escalate. It never generates text, so there is nothing to parse and nothing to hallucinate.

  • 419M parameters โ€” a ModernBERT-Large encoder plus a decision head.
  • Certified, not vibes: selective thresholds with finite-sample guarantees ship in the box, so the model can tell you when its answer is safe to act on โ€” and when it is not.
  • Built for the decision layer under agents: routing requests, gating tool calls, verifying claims.
Repository Mythologic/MIMIR-1 (this card)
Package mimir-decisions on PyPI โ€” pip install "mimir-decisions[local]"
Code and docs abderahmane-ai/mimir
Licence Mythologic Community License โ€” free for research, personal use and organisations under $1M annual revenue; each release becomes Apache-2.0 two years after publication; full text

Intended use and out of scope

Use it for structured decisions where the caller must be able to act or escalate on the model's own terms: routing a ticket to a team, gating a refund tool call, verifying a claim against evidence passages, ranking vendors against a requirement, rating an incident on an ordered scale, estimating a bounded number with an interval.

Do not use it for anything that needs generated prose (it cannot produce text), records that are not in English, or as the only reviewer of a decision whose options were themselves wrong. Where the model cannot back an answer it returns ABSTAINED or DEFERRED โ€” route those to a person or a stronger check rather than reading them as a vote against an option.

Installation

pip install "mimir-decisions[local]"        # CPU engine (certified configuration)
pip install "mimir-decisions[local-gpu]"    # CUDA engine
pip install mimir-decisions                 # data models and HTTP client only

Python 3.11 or later. The distribution is mimir-decisions; the import and the command line are both mimir.

Quickstart

from mimir import Mimir

model = Mimir.from_pretrained("Mythologic/MIMIR-1")  # downloads and verifies on first use

result = model.choose(
    "My card was charged twice for the same order.",
    "Which team should handle this?",
    options={"billing": "Billing: payments, refunds", "security": "Security: account access"},
)
print(result.status)          # DECIDED
print(result.answer)          # billing
print(result.probabilities)   # calibrated probability of each option
print(result.certificate)     # the certified threshold the decision was checked against

Three outcomes, not two

Every call ends DECIDED, ABSTAINED or DEFERRED.

  • DECIDED โ€” an option is chosen and certified to act on at your risk level.
  • ABSTAINED โ€” none of the listed options is supported; this is an answer, not a failure.
  • DEFERRED โ€” the certificate does not cover this call; result.deferral.reason says why: below_threshold (confidence missed the certified threshold), out_of_distribution (the input sits outside what the thresholds were certified on) or no_certified_threshold (nothing is certified for this decision type at this risk level).

A decision layer that cannot refuse is a random generator with calibrated-looking outputs. The refusal is the product.

What it decides

You ask You get
choose โ€” pick one option, or none an option id, or None
multi-choice โ€” pick every option that applies the option ids that apply
yes_no โ€” answer a yes/no question True or False
verify โ€” judge a claim against evidence supported, contradicted or not_enough_information
rank โ€” order candidates, best first candidate ids, best first
rate โ€” rate on your ordered scale a level id
estimate โ€” estimate a number in [low, high] a number with an interval

Every result carries status, confidence, relevant_context (the context parts behind the answer, most relevant first), certificate, deferral and latency_ms. Choice, yes/no and verify results add abstain_probability and a conformal prediction_set.

A context is a string, a list of passages, tables with typed numbers and dates, or a JSON state. decide_many batches many decisions, and every method has an async form. The answer space is defined at request time, so new option sets need no retraining.

Benchmarks

Head-to-head against Laya (convaiinnovations/laya) and GLiNER2.5-Decide (fastino/GLiNER2.5-Decide), each run by us on identical records with paired bootstrap 95% intervals. Exact-match accuracy at full coverage:

Task n MIMIR Laya GLiNER MIMIR โˆ’ Laya
Banking77 3,076 0.883 0.335 0.706 +0.547 [+0.530, +0.564]
MASSIVE en 2,974 0.866 0.443 0.643 +0.423 [+0.403, +0.443]
typed-decisions 2,000 0.725 0.361 0.487 +0.363 [+0.332, +0.394]
prompt-injections 116 0.957 0.681 0.638 +0.276 [+0.190, +0.362]
XNLI en 4,995 0.828 0.676 0.398 +0.152 [+0.137, +0.167]
BoolQ 3,270 0.840 0.777 0.727 +0.064 [+0.048, +0.080]
AG News 7,600 0.921 0.926 0.735 โˆ’0.005 [โˆ’0.011, +0.000]
SST-5 2,210 0.362 0.360 0.446 +0.001 [โˆ’0.029, +0.033]
fast-decisions 2,900 0.432 0.527 0.620 โˆ’0.050 [โˆ’0.071, โˆ’0.030]
DAIR Emotion 2,000 0.080 0.589 0.562 โˆ’0.508 [โˆ’0.530, โˆ’0.485]

MIMIR leads decisively on six tasks and ties Laya on AG News and SST-5; GLiNER leads on its own fast-decisions showcase. The DAIR row is abstention, not error: the model answers "none of the listed options" on 1,792 of 2,000 records (mean abstain mass 0.656), because none of the six listed emotions is supported at its calibrated abstention level. Forced to pick among the six it scores 0.595, above Laya's 0.589. If your pipeline needs a label on every row, read the argmax over the options; if it needs honesty, read the abstention.

The certificate

The shipped fp32 policy carries selective thresholds proven on 95,550 held-out records at four risk levels, at confidence 0.95. Read the table as a contract: at each risk level, on the calls the model chooses to take, its observed error stays under the risk.

Binary, categorical and ranking (taken / evaluated, observed error on taken calls):

Risk level Binary Categorical Ranking
0.5% 2,909 of 19,319 (15.1%) ยท 0.17% err 15,171 of 54,728 (27.7%) ยท 0.36% err defers
1% 6,781 of 19,319 (35.1%) ยท 0.69% err 19,947 of 54,728 (36.4%) ยท 0.83% err 1,441 of 3,863 (37.3%) ยท 0.35% err
2% 8,903 of 19,319 (46.1%) ยท 1.63% err 25,917 of 54,728 (47.4%) ยท 1.76% err 2,339 of 3,863 (60.5%) ยท 1.28% err
5% 11,173 of 19,319 (57.8%) ยท 4.45% err 35,761 of 54,728 (65.3%) ยท 4.68% err 3,422 of 3,863 (88.6%) ยท 1.52% err

Ordinal defers at 0.5%, 1% and 2% risk; at 5% it takes 192 of 14,160 records (1.04% error). Multilabel and continuous defer at every risk level in this release โ€” their first testable thresholds fail on held-out data, and shipping them would be shipping an unproven promise. The package's default risk level is 1%.

Thresholds were fit on the PyTorch model's outputs and transferred to the ONNX graph, whose outputs match the model on every decision over 312 verification records (max relative readout gap 6.5e-5, zero decision differences).

On your own data and hardware. The certified configuration is CPU with ONNX Runtime 1.30.0. On other runtimes or hardware, either run the shipped 200-record equivalence set (the first load does this automatically and requires every decision to match) or calibrate your own policy from your own labels:

from mimir import Mimir
from mimir.core.labels import LabelledDecision
from mimir.runtime.calibrate import calibrate

model = Mimir.from_pretrained("Mythologic/MIMIR-1")
labelled: list[LabelledDecision] = [...]  # your context/decision/label triples
policy = calibrate(model, labelled, risk=0.01, confidence=0.95)

mimir doctor --verify runs the same equivalence check on demand.

Integrity

Before any model file is read, the package verifies the manifest's Sigstore signature against the abderahmane-ai/mimir release workflow, checks every file's SHA-256 against the manifest, and checks the ONNX graph against its operator allowlist. No pickle is used anywhere. The manifest and its signature cover every file below, including the licence.

Load paths

The release root is the mimir-decisions contract (ONNX graphs plus certified policy); the same weights ship as plain transformers under a subfolder. Only what you load is downloaded.

Path Format Best at
Mythologic/MIMIR-1 (repo root) ONNX graphs + certified policy, via mimir-decisions production: routing, gating, verification
Mythologic/MIMIR-1, subfolder="transformers" same weights for plain transformers research, custom runtimes

Package mode (recommended)

mimir-decisions handles chunking, typed readouts, evidence, the distance gate and the certificate in one call, and additionally serves them over HTTP and MCP:

mimir serve  # HTTP, configured decision endpoints
mimir mcp    # MCP server: configured decision tools by default

Both serve the certified fp32 graph with the shipped policy, so the certificate above applies to what they return. Adapters for eight agent frameworks, worked examples and the full API reference live in the package repository.

Transformers mode

With no mimir-decisions install (tested transformers==5.17):

from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained(
    "Mythologic/MIMIR-1", subfolder="transformers", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("Mythologic/MIMIR-1", subfolder="transformers")
outputs = model(chunk_ids, chunk_mask, ...)  # the 21 graph inputs, in order
outputs.utilities  # [4 depths, records, candidates]: every depth's readouts

Load plus forward on pre-encoded batches only: text handling, calibration, the server and MCP stay in mimir-decisions. The staged modelling code holds the inference-only head forward โ€” no loss, no training machinery โ€” the same computation the shipped model.onnx already carries.

Speed

About 1.7 decisions per second on an Apple M4 CPU, flat from 1 to 32 concurrent clients โ€” the batching absorbs concurrency rather than queueing it. GPU throughput is not measured in this release.

Architecture

  • Backbone: a ModernBERT-Large encoder plus a decision head trained from scratch: question-conditioned state memory, a candidate set processor, and a shared reasoning cell applied once per iteration. 419M parameters total.
  • Typed readouts: choice, multi-choice, yes/no, verification, ranking, rating and estimation, each producing its own typed result.
  • Budget: 512-token chunks capped at 256 question tokens; tested to 150 options, 49 levels and 26,406 context tokens. Larger inputs are refused.
  • Refusal built in: an out-of-distribution distance gate, conformal prediction sets, selective thresholds with finite-sample certificates, and an evidence mass over the context behind every decision.

Training

Training data and procedure are private. No optimiser state, scheduler state or RNG ships with the release; training cannot be resumed from these files.

Limitations

  • English only. Non-English input is out of distribution.
  • Certified coverage is selective by design. Ordinal certifies only at 5% risk; multilabel and continuous defer at every level. Under the shipped policy those calls come back DEFERRED โ€” run them uncertified only if your pipeline accepts an uncertified answer, or calibrate your own policy.
  • Abstention changes accuracy accounting. On tasks whose listed options frequently fail to apply, the model abstains instead of guessing (see DAIR Emotion above). Read ABSTAINED as an answer, not a wrong one.
  • The fp16 graph ships without a policy. There is no CUDA certificate in this release; certify it yourself or run the equivalence set first.
  • One measured speed point. The 1.7 decisions/s figure is M4 CPU only.

FAQ

Can I use a GPU? Yes, with the fp16 graph โ€” it ships without a policy, so calibrate your own or run the equivalence set first. There is no CUDA certificate in this release.

What do I do with ABSTAINED and DEFERRED? Route them to a person or a stronger check. Pipelines that treat every call as an answer turn refusals into someone else's silent default.

Why does ordinal only certify at 5% risk? Its first testable thresholds fail on held-out data at lower risks; at 5% the certified slice is small (192 records) with a 1.04% observed error. Multilabel and continuous have no certifiable threshold at any level in this release.

How do I reproduce the numbers? The release carries a 200-record equivalence set: replay the graph on any CPU and every shipped decision must match. mimir doctor --verify runs it on demand.

Files

file description
onnx/model.onnx (+ _data) the public fp32 graph, CPU-certified
onnx/model_fp16.onnx (+ _data) the fp16 graph, uncertified
transformers/model.safetensors the same weights for plain transformers
transformers/config.json encoder and head shapes, auto_map, architectures
transformers/modeling_mimir.py, transformers/configuration_mimir.py the architecture, transformers and torch only
tokenizer.json the 50,368-entry vocabulary
config.json variants, limits, layout and graph contract for mimir-decisions
policy/fp32.json (+ .npz) the certified selective thresholds at four risk levels
equivalence/ 200 records with inputs, so any CPU replays the graph
LICENSE the Mythologic Community Model License
manifest.json SHA-256 of every file above
manifest.json.sigstore the Sigstore bundle signing the manifest

Links

Licence

Mythologic Community License (see LICENSE): free for research, personal use and organisations under $1M annual revenue; commercial use above that needs a commercial licence; each release becomes available under the Apache License 2.0 two years after its publication. Commercial licensing, on-premises and indemnity terms: abderahmane.ainouche.ai@gmail.com.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Mythologic/MIMIR-1

Quantized
(2)
this model

Evaluation results