MIMIR-1
Decisions your agents can act on. MIMIR is an English-only, non-generative decision model: give it a context, a question and the options, and it returns a typed answer with calibrated probabilities, the parts of the context it relied on, and a certified signal for when to act and when to escalate. It never generates text, so there is nothing to parse and nothing to hallucinate.
- 419M parameters โ a ModernBERT-Large encoder plus a decision head.
- Certified, not vibes: selective thresholds with finite-sample guarantees ship in the box, so the model can tell you when its answer is safe to act on โ and when it is not.
- Built for the decision layer under agents: routing requests, gating tool calls, verifying claims.
| Repository | Mythologic/MIMIR-1 (this card) |
| Package | mimir-decisions on PyPI โ pip install "mimir-decisions[local]" |
| Code and docs | abderahmane-ai/mimir |
| Licence | Mythologic Community License โ free for research, personal use and organisations under $1M annual revenue; each release becomes Apache-2.0 two years after publication; full text |
Intended use and out of scope
Use it for structured decisions where the caller must be able to act or escalate on the model's own terms: routing a ticket to a team, gating a refund tool call, verifying a claim against evidence passages, ranking vendors against a requirement, rating an incident on an ordered scale, estimating a bounded number with an interval.
Do not use it for anything that needs generated prose (it cannot produce text),
records that are not in English, or as the only reviewer of a decision whose options
were themselves wrong. Where the model cannot back an answer it returns ABSTAINED or
DEFERRED โ route those to a person or a stronger check rather than reading them as a
vote against an option.
Installation
pip install "mimir-decisions[local]" # CPU engine (certified configuration)
pip install "mimir-decisions[local-gpu]" # CUDA engine
pip install mimir-decisions # data models and HTTP client only
Python 3.11 or later. The distribution is mimir-decisions; the import and the command
line are both mimir.
Quickstart
from mimir import Mimir
model = Mimir.from_pretrained("Mythologic/MIMIR-1") # downloads and verifies on first use
result = model.choose(
"My card was charged twice for the same order.",
"Which team should handle this?",
options={"billing": "Billing: payments, refunds", "security": "Security: account access"},
)
print(result.status) # DECIDED
print(result.answer) # billing
print(result.probabilities) # calibrated probability of each option
print(result.certificate) # the certified threshold the decision was checked against
Three outcomes, not two
Every call ends DECIDED, ABSTAINED or DEFERRED.
DECIDEDโ an option is chosen and certified to act on at your risk level.ABSTAINEDโ none of the listed options is supported; this is an answer, not a failure.DEFERREDโ the certificate does not cover this call;result.deferral.reasonsays why:below_threshold(confidence missed the certified threshold),out_of_distribution(the input sits outside what the thresholds were certified on) orno_certified_threshold(nothing is certified for this decision type at this risk level).
A decision layer that cannot refuse is a random generator with calibrated-looking outputs. The refusal is the product.
What it decides
| You ask | You get |
|---|---|
choose โ pick one option, or none |
an option id, or None |
| multi-choice โ pick every option that applies | the option ids that apply |
yes_no โ answer a yes/no question |
True or False |
verify โ judge a claim against evidence |
supported, contradicted or not_enough_information |
rank โ order candidates, best first |
candidate ids, best first |
rate โ rate on your ordered scale |
a level id |
estimate โ estimate a number in [low, high] |
a number with an interval |
Every result carries status, confidence, relevant_context (the context parts
behind the answer, most relevant first), certificate, deferral and latency_ms.
Choice, yes/no and verify results add abstain_probability and a conformal
prediction_set.
A context is a string, a list of passages, tables with typed numbers and dates, or a
JSON state. decide_many batches many decisions, and every method has an async form.
The answer space is defined at request time, so new option sets need no retraining.
Benchmarks
Head-to-head against Laya (convaiinnovations/laya) and GLiNER2.5-Decide
(fastino/GLiNER2.5-Decide), each run by us on identical records with paired bootstrap
95% intervals. Exact-match accuracy at full coverage:
| Task | n | MIMIR | Laya | GLiNER | MIMIR โ Laya |
|---|---|---|---|---|---|
| Banking77 | 3,076 | 0.883 | 0.335 | 0.706 | +0.547 [+0.530, +0.564] |
| MASSIVE en | 2,974 | 0.866 | 0.443 | 0.643 | +0.423 [+0.403, +0.443] |
| typed-decisions | 2,000 | 0.725 | 0.361 | 0.487 | +0.363 [+0.332, +0.394] |
| prompt-injections | 116 | 0.957 | 0.681 | 0.638 | +0.276 [+0.190, +0.362] |
| XNLI en | 4,995 | 0.828 | 0.676 | 0.398 | +0.152 [+0.137, +0.167] |
| BoolQ | 3,270 | 0.840 | 0.777 | 0.727 | +0.064 [+0.048, +0.080] |
| AG News | 7,600 | 0.921 | 0.926 | 0.735 | โ0.005 [โ0.011, +0.000] |
| SST-5 | 2,210 | 0.362 | 0.360 | 0.446 | +0.001 [โ0.029, +0.033] |
| fast-decisions | 2,900 | 0.432 | 0.527 | 0.620 | โ0.050 [โ0.071, โ0.030] |
| DAIR Emotion | 2,000 | 0.080 | 0.589 | 0.562 | โ0.508 [โ0.530, โ0.485] |
MIMIR leads decisively on six tasks and ties Laya on AG News and SST-5; GLiNER leads on its own fast-decisions showcase. The DAIR row is abstention, not error: the model answers "none of the listed options" on 1,792 of 2,000 records (mean abstain mass 0.656), because none of the six listed emotions is supported at its calibrated abstention level. Forced to pick among the six it scores 0.595, above Laya's 0.589. If your pipeline needs a label on every row, read the argmax over the options; if it needs honesty, read the abstention.
The certificate
The shipped fp32 policy carries selective thresholds proven on 95,550 held-out records at four risk levels, at confidence 0.95. Read the table as a contract: at each risk level, on the calls the model chooses to take, its observed error stays under the risk.
Binary, categorical and ranking (taken / evaluated, observed error on taken calls):
| Risk level | Binary | Categorical | Ranking |
|---|---|---|---|
| 0.5% | 2,909 of 19,319 (15.1%) ยท 0.17% err | 15,171 of 54,728 (27.7%) ยท 0.36% err | defers |
| 1% | 6,781 of 19,319 (35.1%) ยท 0.69% err | 19,947 of 54,728 (36.4%) ยท 0.83% err | 1,441 of 3,863 (37.3%) ยท 0.35% err |
| 2% | 8,903 of 19,319 (46.1%) ยท 1.63% err | 25,917 of 54,728 (47.4%) ยท 1.76% err | 2,339 of 3,863 (60.5%) ยท 1.28% err |
| 5% | 11,173 of 19,319 (57.8%) ยท 4.45% err | 35,761 of 54,728 (65.3%) ยท 4.68% err | 3,422 of 3,863 (88.6%) ยท 1.52% err |
Ordinal defers at 0.5%, 1% and 2% risk; at 5% it takes 192 of 14,160 records (1.04% error). Multilabel and continuous defer at every risk level in this release โ their first testable thresholds fail on held-out data, and shipping them would be shipping an unproven promise. The package's default risk level is 1%.
Thresholds were fit on the PyTorch model's outputs and transferred to the ONNX graph, whose outputs match the model on every decision over 312 verification records (max relative readout gap 6.5e-5, zero decision differences).
On your own data and hardware. The certified configuration is CPU with ONNX Runtime 1.30.0. On other runtimes or hardware, either run the shipped 200-record equivalence set (the first load does this automatically and requires every decision to match) or calibrate your own policy from your own labels:
from mimir import Mimir
from mimir.core.labels import LabelledDecision
from mimir.runtime.calibrate import calibrate
model = Mimir.from_pretrained("Mythologic/MIMIR-1")
labelled: list[LabelledDecision] = [...] # your context/decision/label triples
policy = calibrate(model, labelled, risk=0.01, confidence=0.95)
mimir doctor --verify runs the same equivalence check on demand.
Integrity
Before any model file is read, the package verifies the manifest's Sigstore signature
against the abderahmane-ai/mimir release workflow, checks every file's SHA-256 against
the manifest, and checks the ONNX graph against its operator allowlist. No pickle is
used anywhere. The manifest and its signature cover every file below, including the
licence.
Load paths
The release root is the mimir-decisions contract (ONNX graphs plus certified policy);
the same weights ship as plain transformers under a subfolder. Only what you load is
downloaded.
| Path | Format | Best at |
|---|---|---|
Mythologic/MIMIR-1 (repo root) |
ONNX graphs + certified policy, via mimir-decisions |
production: routing, gating, verification |
Mythologic/MIMIR-1, subfolder="transformers" |
same weights for plain transformers |
research, custom runtimes |
Package mode (recommended)
mimir-decisions handles chunking, typed readouts, evidence, the distance gate and the
certificate in one call, and additionally serves them over HTTP and MCP:
mimir serve # HTTP, configured decision endpoints
mimir mcp # MCP server: configured decision tools by default
Both serve the certified fp32 graph with the shipped policy, so the certificate above applies to what they return. Adapters for eight agent frameworks, worked examples and the full API reference live in the package repository.
Transformers mode
With no mimir-decisions install (tested transformers==5.17):
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained(
"Mythologic/MIMIR-1", subfolder="transformers", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("Mythologic/MIMIR-1", subfolder="transformers")
outputs = model(chunk_ids, chunk_mask, ...) # the 21 graph inputs, in order
outputs.utilities # [4 depths, records, candidates]: every depth's readouts
Load plus forward on pre-encoded batches only: text handling, calibration, the server
and MCP stay in mimir-decisions. The staged modelling code holds the inference-only
head forward โ no loss, no training machinery โ the same computation the shipped
model.onnx already carries.
Speed
About 1.7 decisions per second on an Apple M4 CPU, flat from 1 to 32 concurrent clients โ the batching absorbs concurrency rather than queueing it. GPU throughput is not measured in this release.
Architecture
- Backbone: a ModernBERT-Large encoder plus a decision head trained from scratch: question-conditioned state memory, a candidate set processor, and a shared reasoning cell applied once per iteration. 419M parameters total.
- Typed readouts: choice, multi-choice, yes/no, verification, ranking, rating and estimation, each producing its own typed result.
- Budget: 512-token chunks capped at 256 question tokens; tested to 150 options, 49 levels and 26,406 context tokens. Larger inputs are refused.
- Refusal built in: an out-of-distribution distance gate, conformal prediction sets, selective thresholds with finite-sample certificates, and an evidence mass over the context behind every decision.
Training
Training data and procedure are private. No optimiser state, scheduler state or RNG ships with the release; training cannot be resumed from these files.
Limitations
- English only. Non-English input is out of distribution.
- Certified coverage is selective by design. Ordinal certifies only at 5% risk;
multilabel and continuous defer at every level. Under the shipped policy those calls
come back
DEFERREDโ run them uncertified only if your pipeline accepts an uncertified answer, or calibrate your own policy. - Abstention changes accuracy accounting. On tasks whose listed options frequently
fail to apply, the model abstains instead of guessing (see DAIR Emotion above). Read
ABSTAINEDas an answer, not a wrong one. - The fp16 graph ships without a policy. There is no CUDA certificate in this release; certify it yourself or run the equivalence set first.
- One measured speed point. The 1.7 decisions/s figure is M4 CPU only.
FAQ
Can I use a GPU? Yes, with the fp16 graph โ it ships without a policy, so calibrate your own or run the equivalence set first. There is no CUDA certificate in this release.
What do I do with ABSTAINED and DEFERRED? Route them to a person or a stronger check. Pipelines that treat every call as an answer turn refusals into someone else's silent default.
Why does ordinal only certify at 5% risk? Its first testable thresholds fail on held-out data at lower risks; at 5% the certified slice is small (192 records) with a 1.04% observed error. Multilabel and continuous have no certifiable threshold at any level in this release.
How do I reproduce the numbers?
The release carries a 200-record equivalence set: replay the graph on any CPU and every
shipped decision must match. mimir doctor --verify runs it on demand.
Files
| file | description |
|---|---|
onnx/model.onnx (+ _data) |
the public fp32 graph, CPU-certified |
onnx/model_fp16.onnx (+ _data) |
the fp16 graph, uncertified |
transformers/model.safetensors |
the same weights for plain transformers |
transformers/config.json |
encoder and head shapes, auto_map, architectures |
transformers/modeling_mimir.py, transformers/configuration_mimir.py |
the architecture, transformers and torch only |
tokenizer.json |
the 50,368-entry vocabulary |
config.json |
variants, limits, layout and graph contract for mimir-decisions |
policy/fp32.json (+ .npz) |
the certified selective thresholds at four risk levels |
equivalence/ |
200 records with inputs, so any CPU replays the graph |
LICENSE |
the Mythologic Community Model License |
manifest.json |
SHA-256 of every file above |
manifest.json.sigstore |
the Sigstore bundle signing the manifest |
Links
- Package:
mimir-decisionson PyPI,from mimir import Mimir - Code: abderahmane-ai/mimir on GitHub, with the documentation site at abderahmane-ai.github.io/mimir
- MCP Registry:
io.github.abderahmane-ai/mimir - Container:
ghcr.io/abderahmane-ai/mimir
Licence
Mythologic Community License (see LICENSE): free for research, personal use
and organisations under $1M annual revenue; commercial use above that needs a
commercial licence; each release becomes available under the Apache License 2.0 two
years after its publication. Commercial licensing, on-premises and indemnity terms:
abderahmane.ainouche.ai@gmail.com.
- Downloads last month
- -
Model tree for Mythologic/MIMIR-1
Base model
answerdotai/ModernBERT-Large-InstructEvaluation results
- Exact-match accuracy on Banking77 testtest set self-reported0.883
- Exact-match accuracy on MASSIVE English testtest set self-reported0.866
- Exact-match accuracy on typed-decisions testtest set self-reported0.725
- Exact-match accuracy on prompt-injectionstest set self-reported0.957
- Exact-match accuracy on XNLI English testtest set self-reported0.828
- Exact-match accuracy on BoolQ testtest set self-reported0.840
- Exact-match accuracy on AG News testtest set self-reported0.921