Mimir (impacte/mimir-laya-router)

Mimir β€” the Norse keeper of the Well of Knowledge, whose counsel even Odin sought. This model is the homelab's counsel: given a task and the fleet's model cards, it decides which model should answer.

Mimir is convaiinnovations/laya (421M, ModernBERT-large + typed decision head) fine-tuned into a model router. It takes a user request plus a list of candidate models (name + capability card) and returns which model should handle it, with calibrated probabilities β€” in a single non-autoregressive forward pass (~25 ms).

Base convaiinnovations/laya (Apache-2.0)
Architecture ModernBERT-large encoder + 2-layer decision head + option-marker scorer (421M)
Question type choice over the model lineup (plus a none option)
Training RLCD β€” GRPO-style policy gradients against a strictly proper scoring rule + soft cross-entropy
License Apache-2.0 (inherited from Laya)

How it works

Each candidate model is rendered as a choice option whose text is its capability card; a none option catches vague or out-of-scope requests. Laya scores every option at its own [MASK] marker and softmaxes over the lineup, so the answer space is defined at request time β€” add or remove models without retraining.

import laya

router = laya.load("impacte/mimir-laya-router")

res = router.predict(
    "Refactor this Rust module and add unit tests for the parser.",
    {"route": {
        "type": "choice",
        "instructions": "Which agent should handle this request? Judge only by each agent's card.",
        "criteria": {
            "qwen3.8-27b": "27B general model, strongest for code, math and multi-step reasoning",
            "minicpm-2b": "2B tiny model, fastest; quick facts and short answers",
            "none": "no agent fits; ask a clarifying question",
        },
    }},
)
print(res["answers"]["route"]["choice"])        # -> qwen3.8-27b
print(res["answers"]["route"]["probabilities"]) # -> {'qwen3.8-27b': 0.97, ...}

A FastAPI sidecar (POST /route) is included in the training repo.

Results

Held-out accuracy 0.9688 (95% CI 0.922–1.000), none F1 0.957, ECE 0.0262, Brier 0.0762 (n=64).

Baselines (same held-out rows)

router accuracy none F1 latency p50 (ms)
random 0.1860 0.000 β€”
always_none 0.1938 0.325 β€”
always_first 0.2016 0.000 β€”
bm25 0.2016 0.279 0.05
base_laya 0.2403 0.336 23.41
llm_router 0.6822 0.585 151.33

Stronger LLM baselines (reasoning on / few-shot / 27B)

variant accuracy none F1 latency p50 (ms)
gemma4:e4b__think0_fs0 0.7364 0.681 144
gemma4:e4b__think1_fs0 0.7364 0.550 6001
gemma4:e4b__think1_fs1 0.7752 0.700 4972
qwen3.8-27b__think0_fs1 0.7752 0.800 955

The strongest LLM-router configuration reaches 0.775; Mimir is ~19 points ahead at ~40–200Γ— lower latency.

Ablation: RLCD vs cross-entropy only (same held-out split)

training objective accuracy honest ECE Brier NLL
RLCD (this model) 0.9688 0.0262 0.0762 0.2022
CE only 0.9375 0.0623 0.1035 0.2209

The RLCD term improves both accuracy and calibration over plain cross-entropy (single seed, n=64 β€” treat the accuracy delta as indicative).

Calibration (disjoint split)

setting temperature ECE NLL mean max prob
raw (T=1) 1.000 0.0459 0.6359 0.9854
fit on test (leaky) 4.160 0.0165 0.2004 0.9542
fit on disjoint calib (honest) 4.471 0.0262 0.2022 0.9426

Split: 65 calib / 64 test. The earlier card reported ECE 0.0139 with the temperature fit on the same test set; the honest held-out ECE is 0.0262.

Accuracy vs number of candidate models

K options accuracy mean max prob
2 0.9612 0.993
3 0.9302 0.992
5 0.8760 0.985
8 0.6667 0.970
12 0.4651 0.963
20 0.1705 0.924
40 0.0698 0.440

The original test set only had 4–7 options; accuracy collapses beyond ~5–8 candidates. Keep the lineup small or route in two stages (provider β†’ model).

Robustness

  • card paraphrase (templated): 0.9225 (Ξ” -0.039)
  • card paraphrase (LLM): 0.8217 (Ξ” -0.140)
  • real vs fictional names: 0.9565 vs 0.9623
  • request buried after long context: 0.3721 (vs 0.9612 request-only)
  • option-order shuffle: 0.9302 (Ξ” -0.031)

Training data

Requests were generated by a local gemma4:e4b (1,770 seeds across 14 categories: code, math, heavy reasoning, data analysis, long context, chat, summarization, extraction, translation, creative, light reasoning, quick facts, classification, ambiguous). Each request is expanded into a lineup where exactly one agent's card covers the request category (the gold label), with ~80% fictional model names (forcing the model to read the card, not memorise names) and ~20% the real homelab fleet. Ambiguous requests route to none. 12,000 training rows; the test set uses disjoint seeds.

Limitations

  • English only β€” the base is the English Laya checkpoint.
  • Card quality matters β€” routing is only as good as the capability descriptions you supply.
  • Scales poorly with the number of candidates. Accuracy is ~0.96 at 2 options but falls to ~0.67 at 8, ~0.47 at 12 and ~0.17 at 20 (see the table above). Keep the lineup small (≀5–8) or use a two-stage route (pick a provider first, then a model within it).
  • Sensitive to card phrasing. Rewording the cards with an LLM drops accuracy by ~0.14; the model leans on surface wording, not deep semantics.
  • Only reads the start of the state. The request is truncated to the first ~320 tokens, so a task buried after long context is invisible (accuracy 0.96 β†’ 0.37).
  • none recall is the weak spot; tune the abstain threshold on your own data.
  • Fictional-name training means real deployment names are handled by reading their cards, not by name recognition (verified: real vs fictional names differ by <0.01).

Credits

Citation

@misc{mimir_laya_router2026,
  title  = {Mimir: Laya fine-tuned as a calibrated model router},
  author = {oamazonasgabriel},
  year   = {2026},
  url    = {https://huggingface.co/impacte/mimir-laya-router}
}
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for impacte/mimir-laya-router

Finetuned
(84)
this model