SeLMRoute-Laya
SeLMRoute-Laya is a CatBoost deployment router that predicts the performance of 20 candidate LLMs from 40 probabilistic semantic features extracted using Laya.
Paper · Source code · Frozen features
What the model does
The pipeline is: query → Laya semantic evidence → 40-dimensional ProbabilityMass vector → 20 predicted candidate scores → candidate ranking.
This checkpoint is a CatBoostRegressor with the MultiRMSE objective. It takes numeric semantic features, not raw query text, and does not generate an answer or call the selected LLM. Higher predicted scores rank first for performance-oriented routing. These regression outputs are not calibrated probabilities and are not guaranteed to fall in [0, 1].
Files and input/output contract
| File | Purpose |
|---|---|
model.cbm |
Original trained CatBoost checkpoint. |
metadata.json |
Exact ordered feature_names, ordered model_names, and training parameters. |
inference.py |
Download-and-run example using the matching frozen features. |
requirements.txt |
Dependencies for this example. |
LICENSE, NOTICE, CITATION.cff |
Original license, notices and citation metadata. |
metadata.json is essential: input column j must correspond to feature_names[j], and output column k to model_names[k]. Use all 40 listed features in that order. The source CSV has 112 columns; passing every column to the router is incorrect. Match the Laya checkpoint to Laya features.
Quick start
python -m pip install -r requirements.txt
python inference.py
The example is reproduced below. It downloads published artifacts from Hugging Face, but does not call JEV or Laya. Both the model repository and the linked dataset repository must have been uploaded first.
import json
import numpy as np
import pandas as pd
from catboost import CatBoostRegressor
from huggingface_hub import hf_hub_download
MODEL_REPO = "Indigma-Innovations/SeLMRoute-Laya"
DATASET_REPO = "Indigma-Innovations/SeLMRoute-Semantic-Features"
metadata_path = hf_hub_download(MODEL_REPO, "metadata.json")
with open(metadata_path, encoding="utf-8") as f:
metadata = json.load(f)
router = CatBoostRegressor()
router.load_model(hf_hub_download(MODEL_REPO, "model.cbm"))
# Frozen features: this example makes no semantic-extraction call.
feature_path = hf_hub_download(
DATASET_REPO,
"data/performance/features_laya.csv",
repo_type="dataset",
)
frame = pd.read_csv(feature_path, dtype={"sample_id": str})
row = frame.iloc[[0]]
# Preserve the checkpoint's exact feature order; exclude all other columns.
X = row[metadata["feature_names"]].to_numpy(dtype=float)
assert X.shape == (1, 40) and np.isfinite(X).all()
scores = np.asarray(router.predict(X), dtype=float).reshape(1, -1)[0]
assert len(scores) == len(metadata["model_names"]) == 20
order = np.argsort(-scores)
print("Sample:", row.iloc[0]["sample_id"])
for idx in order[:5]:
print(metadata["model_names"][int(idx)], float(scores[idx]))
This demonstrates inference on a training-pool sample, not held-out evaluation.
New queries
Extract the same semantic schema with Laya and align the resulting values using metadata.json. Live Laya extraction uses the optional live-laya dependency; its model must be downloaded before local inference. The source project specifies laya==0.3.5. See the live-routing instructions and notebooks/03_live_routing.ipynb in the source repository. The complete source package targets Python 3.13.
Semantic representation
Eight binary probabilities and eight four-level distributions produce 8 + 8 × 4 = 40 inputs. Binary columns end in __noul; ordinal probability columns end in __p__0 through __p__3. The diagnostic task_family probe and auxiliary columns are excluded from this checkpoint's input.
| Probe | Type | Meaning |
|---|---|---|
math_reasoning |
Binary probability | Does correctly solving query require mathematical calculation, symbolic manipulation, or quantitative reasoning beyond copying an explicitly stated value? |
code_reasoning |
Binary probability | Does correctly solving query require writing, modifying, debugging, or reasoning about executable computer code? |
formal_logic |
Binary probability | Does solving query materially require formal logical, combinatorial, rule-based, or constraint-satisfaction reasoning? |
factual_recall |
Binary probability | Can query be answered primarily through factual knowledge or retrieval, with little derivation or multi-step reasoning? |
social_affective |
Binary probability | Does correctness materially depend on interpreting emotion, intention, social context, interpersonal meaning, or conversational affect? |
tool_interaction |
Binary probability | Does completing query require interaction with an external tool, software environment, API, or simulated environment rather than only producing an answer? |
external_knowledge |
Binary probability | Does solving query require factual information that is not explicitly supplied in the query and cannot be derived from the supplied information alone? |
current_information |
Binary probability | Would correctness materially depend on recent, changing, or time-sensitive information? |
domain_specialization |
Four-level distribution | How specialized is the knowledge required to solve query correctly? |
reasoning_depth |
Four-level distribution | How much sequential reasoning is required to reach a correct answer to query? |
constraint_density |
Four-level distribution | How many independent requirements or constraints must simultaneously be satisfied by a correct answer to query? |
context_integration |
Four-level distribution | How much information from different parts of query must be integrated to solve it correctly? |
decomposition_need |
Four-level distribution | To what extent does solving query require decomposing it into distinct intermediate subproblems? |
ambiguity |
Four-level distribution | How underspecified or semantically ambiguous is query? |
exactness |
Four-level distribution | How sensitive is correctness to exact details, precise constraints, exact values, or exact output behavior in query? |
answer_openness |
Four-level distribution | How broad is the set of answers that could reasonably count as correct for query? |
Training and evaluation
This deployment checkpoint was fitted on the full performance-oriented pool: 11,481 queries, 15 datasets, 20 candidate models and 229,620 query-model outcomes. Parameters in the original metadata: 500 iterations, depth 6, learning rate 0.05, random seed 3407 and MultiRMSE loss.
This is a full-data deployment checkpoint, not an evaluation checkpoint. Paper results come from separately trained grouped train/test or grouped out-of-fold experiments. Do not evaluate this checkpoint on its training pool and report that as held-out performance.
The source release reports grouped five-fold OOF AvgAcc of 70.5339% for the Laya ProbabilityMass method. The strongest fixed candidate in the grouped OOF comparison achieved 69.2267%. AvgAcc measures routed task performance, not the percentage of queries where the oracle-best candidate was selected.
Candidate models
The following list is in the exact output order stored in metadata.json:
DeepHermes-3-Llama-3-8B-PreviewDeepSeek-R1-0528-Qwen3-8BDeepSeek-R1-Distill-Qwen-7BFin-R1GLM-Z1-9B-0414Intern-S1-miniLlama-3.1-8B-InstructLlama-3.1-8B-UltraMedicalLlama-3.1-Nemotron-Nano-8B-v1MiMo-7B-RL-0530MiniCPM4.1-8BNVIDIA-Nemotron-Nano-9B-v2OpenThinker3-7BQwen2.5-Coder-7B-InstructQwen3-8Bcogito-v1-preview-llama-8Bgemma-2-9b-itglm-4-9b-chatgranite-3.3-8b-instructinternlm3-8b-instruct
Limitations
The candidate pool is fixed to the 20 outputs above. New candidates, model revisions and deployment distributions require validation and potentially refitting the predictor. Interpretable input evidence does not make every downstream CatBoost decision fully explainable. Semantic extraction can be uncertain or incorrect; it also introduces latency and, for API extraction, cost. This checkpoint does not implement the separate 13-model performance-cost experiment.
Provenance and license
Copied without retraining from models/laya/ at source commit ec42fdbad3d30320679f0a25a7f7df335d379f1d. Apache-2.0 applies to SeLMRoute material; it does not relicense third-party models or benchmarks.
Citation
@article{perifanis2026selmroute,
title = {SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing},
author = {Perifanis, Vasilis and Pavlidis, Nikolaos and Symeonidis, Symeon},
journal = {arXiv preprint arXiv:2609.34736},
year = {2026},
url = {https://arxiv.org/abs/2609.34736}
}