s1-llm-auto-router

A small CPU model that classifies a request to an LLM before you route it. In one encoder pass it answers the seven questions that auto-model-router uses to choose a model:

question type answer
category choice coding, agentic, math, knowledge, long_context, tool_use, design, summarisation, general
difficulty score, 5 levels trivial … frontier
stakes score, 4 levels negligible … high
needs_tools, needs_vision, needs_long_context, follow_up yes/no probability
  • Architecture: jhu-clsp/mmBERT-small (MIT), pruned to 11 layers and distilled from a fine-tuned mmBERT-base teacher, with seven classification heads. The temperature for each head is in meta.json.
  • Format: ONNX with int8 embeddings (model.gq.onnx, 183 MB) and a fast tokenizer. It needs neither PyTorch nor transformers.
  • Latency: about 14 ms p50 and 47 ms p99 per decision (all seven answers) on 4 pinned cores of an AMD EPYC 9555 server, and 26 ms p50 on 4 threads of an older Xeon E5-2680 v4.
  • Hosted API: s1-llm-auto-router on system1models.ai, at USD 0.001 per million input tokens.

Results

Held-out test set, n = 2,175. Gold labels are the LLM-ensemble plurality vote (see Caveats). These are our own measurements on this router task, not an independent benchmark or a general model-quality ranking. The comparison set below is Jev 1.13, Winnow-12B and Weiche-395M fp16; it does not cover all routing models. Our larger mmBERT-base teacher reached 0.775 routing agreement on a partial set of 1,881 items, so it is excluded from this full-set table.

system category flags (mean) all 7 exact routing agreement* p50 latency
s1-llm-auto-router 0.898 0.974 0.819 0.742 14–26 ms CPU
Jev 1.13 (hosted) 0.845 0.935 0.603 0.630 247 ms API
Winnow-12B (GPU) 0.841 0.928 0.608 0.577 671 ms GPU
Weiche-395M fp16 0.784 0.903 0.553 0.622 499 ms CPU

*Routing agreement is the share of requests where auto-model-router's own decision policy picks the same model from the predicted labels as from the gold labels.

On a 100-item subset of the held-out test requests, relabelled blind by Claude Opus independently of the ensemble labelling process, the scores were:

system category routing agreement
s1-llm-auto-router 0.88 0.60
Jev 1.13 0.80 0.57
Winnow-12B 0.75 0.35

At n = 100 these numbers carry roughly ±8 points of uncertainty.

Calibration (expected calibration error, test set): 0.018 for category and 0.008–0.028 for the flags.

Caveats

  • It is a fixed-schema router model. It answers only the seven questions above, with these exact option sets. It is not a general-purpose classifier or typed-decision model, and the text of the questions is not read.
  • The training data is synthetic. It consists of 39k routing requests that LLMs wrote from a persona × stack × task × language grid. The grid covers 22 languages and 29 stacks, including SAP ABAP, PL/I and COBOL. The data contains no real user traffic, so real traffic may be distributed differently.
  • The labels come from LLMs. A pool of five open LLMs labelled requests blind, excluding the model that wrote each request. Each retained request has at least three votes; test requests average four labellers. Training uses their soft vote. Most test gold labels come from the same process, which is why the Opus-labelled set above is reported separately.
  • The training data is disjoint from JevBench. It was deduplicated against all public and sealed JevBench items (exact, containment and embedding near-duplicate checks) before training.
  • Performance varies by language. It is weakest on Greek, Korean and Italian, with category accuracy of 0.73–0.78 at n ≈ 20 each.
  • Context is truncated. Inputs are cut to 512 tokens, and only the first 600 characters of the conversation context are used.
  • It is our own model. It was built by System1 Models, the operator of the hosted API. If we benchmark it ourselves, we label it as our own and disclose that conflict of interest.

Usage

import json, numpy as np, onnxruntime as ort
from tokenizers import Tokenizer
from huggingface_hub import snapshot_download

root = snapshot_download("system1models/s1-llm-auto-router")
meta = json.load(open(f"{root}/meta.json")); T = meta["temperatures"]
tok = Tokenizer.from_file(f"{root}/tokenizer/tokenizer.json")
tok.enable_truncation(512, strategy="only_second")
sess = ort.InferenceSession(f"{root}/model.gq.onnx", providers=["CPUExecutionProvider"])

def softmax(x): e = np.exp(x - x.max()); return e / e.sum()
enc = tok.encode("(new conversation)"[:600], "Refactor this ABAP report to use CDS views")
ids = np.asarray([enc.ids], dtype=np.int64)
cat, dif, stk, flg = sess.run(None, {"input_ids": ids, "attention_mask": np.ones_like(ids)})
CATS = ["coding", "agentic", "math", "knowledge", "long_context", "tool_use", "design", "summarisation", "general"]
p = softmax(cat[0] / T["category"]); print(CATS[p.argmax()], p.max())

Integration with auto-model-router is proposed in PR #5 as classifier backend s1-llm-auto-router. See the hosted API documentation for a complete request example. It is Jev-compatible for this fixed router schema only.

Training

  • Teacher: mmBERT-base, fine-tuned on soft labels.
  • Student: mmBERT-small layers 0, 2, …, 18 and 21, distilled with alpha 0.9 for 12 epochs at learning rate 1.5e-4, on one RTX 4090 for about 1 hour.
  • Data split: 34,671 train, 2,275 dev and 2,175 test examples, split by generation group.
  • Calibration: per-head temperatures fitted on the dev set.

Licence

The released model is licensed under MIT, matching the declared licence of its mmBERT base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for system1models/s1-llm-auto-router

Quantized
(280)
this model