Laya Model Routing (fine-tuned)

Laya (421M) fine-tuned to classify a user query into one of six model-routing classes, so an application can pick the right class of free OpenRouter models (~33 ms per query, single forward pass):

class meaning
fast_cheap greetings, basic lookups, formatting
general_mid everyday reasoning, most chat/Q&A
top_reasoning hard multi-step reasoning, math, planning
code_specialist writing, debugging, reviewing code
security_specialist vulnerabilities, secure coding, privacy, auth
long_context long documents / many turns

Results (held-out set: 100 disjoint queries, judge-labeled)

metric base Laya (zero-shot) this checkpoint
choice_accuracy 0.630 0.820
ECE (lower=better) 0.169 0.125
mean_confidence 0.469 0.926

Majority-class baseline: 0.270 · random: 0.167.

Per-class recall: code_specialist 1.000 · security_specialist 0.944 · long_context 0.833 · fast_cheap 0.852 · general_mid 0.600 · top_reasoning 0.000 (only 2 training rows — never predicted).

Training

  • Data: 200 queries from ai-mitra/llm-router-dataset (shuffled, seed 42), labeled by glm-5.3 as judge via Z.AI
  • Recipe: RLCD (policy gradient + proper scoring rule + cross-entropy), DDP on 2× T4, 12 epochs, calibration temperatures fitted post-training (choice T=4.75)
  • Full pipeline and experiment log: github.com/lunakicks/laya-model-routing

Usage

import laya

agent = laya.load("thechristyjo/laya-model-routing", device="cpu")
questions = {
    "model_class": {
        "type": "choice",
        "instructions": "Which model class should handle this task?",
        "criteria": {
            "fast_cheap": "simple, short, low-stakes queries",
            "general_mid": "everyday reasoning, most chat/Q&A",
            "top_reasoning": "hard multi-step reasoning, math, planning",
            "code_specialist": "writing, debugging, reviewing code",
            "security_specialist": "security, privacy, auth tasks",
            "long_context": "long documents or many turns",
        },
    }
}
print(agent.predict("Write a function to detect a linked-list cycle.", questions)
      ["answers"]["model_class"]["choice"])

Limitations

  • Trained on 200 short one-liner queries; long_context reflects summarization topics, not actual long inputs
  • top_reasoning has 2 training rows — expect it to be under-predicted
  • English-only; judge labels carry the judge's "cheapest competent class" policy
  • Trust the confidence gate only after checking ECE on your own traffic
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support