catboost
llm-routing
model-routing
semantic-routing
probabilistic-modeling
interpretable-ai
selmroute
jev

SeLMRoute-JEV

SeLMRoute-JEV is a CatBoost deployment router that predicts the performance of 20 candidate LLMs from 40 probabilistic semantic features extracted using JEV.

Paper · Source code · Frozen features

What the model does

The pipeline is: query → JEV semantic evidence → 40-dimensional ProbabilityMass vector → 20 predicted candidate scores → candidate ranking.

This checkpoint is a CatBoostRegressor with the MultiRMSE objective. It takes numeric semantic features, not raw query text, and does not generate an answer or call the selected LLM. Higher predicted scores rank first for performance-oriented routing. These regression outputs are not calibrated probabilities and are not guaranteed to fall in [0, 1].

Files and input/output contract

File Purpose
model.cbm Original trained CatBoost checkpoint.
metadata.json Exact ordered feature_names, ordered model_names, and training parameters.
inference.py Download-and-run example using the matching frozen features.
requirements.txt Dependencies for this example.
LICENSE, NOTICE, CITATION.cff Original license, notices and citation metadata.

metadata.json is essential: input column j must correspond to feature_names[j], and output column k to model_names[k]. Use all 40 listed features in that order. The source CSV has 112 columns; passing every column to the router is incorrect. Match the JEV checkpoint to JEV features.

Quick start

python -m pip install -r requirements.txt
python inference.py

The example is reproduced below. It downloads published artifacts from Hugging Face, but does not call JEV or Laya. Both the model repository and the linked dataset repository must have been uploaded first.

import json
import numpy as np
import pandas as pd
from catboost import CatBoostRegressor
from huggingface_hub import hf_hub_download

MODEL_REPO = "Indigma-Innovations/SeLMRoute-JEV"
DATASET_REPO = "Indigma-Innovations/SeLMRoute-Semantic-Features"

metadata_path = hf_hub_download(MODEL_REPO, "metadata.json")
with open(metadata_path, encoding="utf-8") as f:
    metadata = json.load(f)

router = CatBoostRegressor()
router.load_model(hf_hub_download(MODEL_REPO, "model.cbm"))

# Frozen features: this example makes no semantic-extraction call.
feature_path = hf_hub_download(
    DATASET_REPO,
    "data/performance/features_jev.csv",
    repo_type="dataset",
)
frame = pd.read_csv(feature_path, dtype={"sample_id": str})
row = frame.iloc[[0]]
# Preserve the checkpoint's exact feature order; exclude all other columns.
X = row[metadata["feature_names"]].to_numpy(dtype=float)
assert X.shape == (1, 40) and np.isfinite(X).all()
scores = np.asarray(router.predict(X), dtype=float).reshape(1, -1)[0]
assert len(scores) == len(metadata["model_names"]) == 20
order = np.argsort(-scores)
print("Sample:", row.iloc[0]["sample_id"])
for idx in order[:5]:
    print(metadata["model_names"][int(idx)], float(scores[idx]))

This demonstrates inference on a training-pool sample, not held-out evaluation.

New queries

Extract the same semantic schema with JEV and align the resulting values using metadata.json. Live JEV extraction requires a TypeSafe API key (TYPESAFE_API_KEY). See the live-routing instructions and notebooks/03_live_routing.ipynb in the source repository. The complete source package targets Python 3.13.

Semantic representation

Eight binary probabilities and eight four-level distributions produce 8 + 8 × 4 = 40 inputs. Binary columns end in __noul; ordinal probability columns end in __p__0 through __p__3. The diagnostic task_family probe and auxiliary columns are excluded from this checkpoint's input.

Probe Type Meaning
math_reasoning Binary probability Does correctly solving query require mathematical calculation, symbolic manipulation, or quantitative reasoning beyond copying an explicitly stated value?
code_reasoning Binary probability Does correctly solving query require writing, modifying, debugging, or reasoning about executable computer code?
formal_logic Binary probability Does solving query materially require formal logical, combinatorial, rule-based, or constraint-satisfaction reasoning?
factual_recall Binary probability Can query be answered primarily through factual knowledge or retrieval, with little derivation or multi-step reasoning?
social_affective Binary probability Does correctness materially depend on interpreting emotion, intention, social context, interpersonal meaning, or conversational affect?
tool_interaction Binary probability Does completing query require interaction with an external tool, software environment, API, or simulated environment rather than only producing an answer?
external_knowledge Binary probability Does solving query require factual information that is not explicitly supplied in the query and cannot be derived from the supplied information alone?
current_information Binary probability Would correctness materially depend on recent, changing, or time-sensitive information?
domain_specialization Four-level distribution How specialized is the knowledge required to solve query correctly?
reasoning_depth Four-level distribution How much sequential reasoning is required to reach a correct answer to query?
constraint_density Four-level distribution How many independent requirements or constraints must simultaneously be satisfied by a correct answer to query?
context_integration Four-level distribution How much information from different parts of query must be integrated to solve it correctly?
decomposition_need Four-level distribution To what extent does solving query require decomposing it into distinct intermediate subproblems?
ambiguity Four-level distribution How underspecified or semantically ambiguous is query?
exactness Four-level distribution How sensitive is correctness to exact details, precise constraints, exact values, or exact output behavior in query?
answer_openness Four-level distribution How broad is the set of answers that could reasonably count as correct for query?

Training and evaluation

This deployment checkpoint was fitted on the full performance-oriented pool: 11,481 queries, 15 datasets, 20 candidate models and 229,620 query-model outcomes. Parameters in the original metadata: 500 iterations, depth 6, learning rate 0.05, random seed 3407 and MultiRMSE loss.

This is a full-data deployment checkpoint, not an evaluation checkpoint. Paper results come from separately trained grouped train/test or grouped out-of-fold experiments. Do not evaluate this checkpoint on its training pool and report that as held-out performance.

The source release reports five-seed AvgAcc of 72.0781% and grouped five-fold OOF AvgAcc of 72.6360% for the JEV ProbabilityMass method. The strongest fixed candidate in the grouped OOF comparison achieved 69.2267%. AvgAcc measures routed task performance, not the percentage of queries where the oracle-best candidate was selected.

Candidate models

The following list is in the exact output order stored in metadata.json:

  1. DeepHermes-3-Llama-3-8B-Preview
  2. DeepSeek-R1-0528-Qwen3-8B
  3. DeepSeek-R1-Distill-Qwen-7B
  4. Fin-R1
  5. GLM-Z1-9B-0414
  6. Intern-S1-mini
  7. Llama-3.1-8B-Instruct
  8. Llama-3.1-8B-UltraMedical
  9. Llama-3.1-Nemotron-Nano-8B-v1
  10. MiMo-7B-RL-0530
  11. MiniCPM4.1-8B
  12. NVIDIA-Nemotron-Nano-9B-v2
  13. OpenThinker3-7B
  14. Qwen2.5-Coder-7B-Instruct
  15. Qwen3-8B
  16. cogito-v1-preview-llama-8B
  17. gemma-2-9b-it
  18. glm-4-9b-chat
  19. granite-3.3-8b-instruct
  20. internlm3-8b-instruct

Limitations

The candidate pool is fixed to the 20 outputs above. New candidates, model revisions and deployment distributions require validation and potentially refitting the predictor. Interpretable input evidence does not make every downstream CatBoost decision fully explainable. Semantic extraction can be uncertain or incorrect; it also introduces latency and, for API extraction, cost. This checkpoint does not implement the separate 13-model performance-cost experiment.

Provenance and license

Copied without retraining from models/jev/ at source commit ec42fdbad3d30320679f0a25a7f7df335d379f1d. Apache-2.0 applies to SeLMRoute material; it does not relicense third-party models or benchmarks.

Citation

@article{perifanis2026selmroute,
  title = {SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing},
  author = {Perifanis, Vasilis and Pavlidis, Nikolaos and Symeonidis, Symeon},
  journal = {arXiv preprint arXiv:2609.34736},
  year = {2026},
  url = {https://arxiv.org/abs/2609.34736}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Indigma/SeLMRoute-JEV

Paper for Indigma/SeLMRoute-JEV