Tollgate

A calibrated LLM cost router that predicts the minimum sufficient model tier for each query.

Tollgate fine-tunes Laya to classify incoming queries into one of three routing tiers:

  • local_small
  • mid_tier
  • frontier

The router produces calibrated probabilities and uses a confidence threshold to decide whether a query can safely be served at the predicted tier or should be escalated to a more capable model.

Model Details

Property Value
Base model Laya
Task LLM tier routing
Model type Fine-tuned encoder
Input context 512 tokens
Output Tier probabilities + routing signals
Tiers local_small, mid_tier, frontier
Language English
License Apache-2.0

Tollgate is designed around a simple objective:

Reduce LLM inference cost while retaining acceptable answer quality.

How It Works

Each training query is answered by every model tier at temperature 0 using the same output cap.

An LLM judge compares each cheaper answer against the frontier answer and assigns one of three judgments:

  • equivalent
  • acceptable
  • worse

The training label is the cheapest tier rated equivalent or acceptable.

If no cheaper tier is sufficient, the query receives the frontier label.

Laya is then fine-tuned to predict the minimum sufficient tier together with two auxiliary routing signals:

  • needs_tools
  • needs_rag

All predictions are produced in a single forward pass.

Calibration

Tollgate uses temperature scaling to improve the reliability of its confidence estimates.

Calibration is performed separately from model training and final evaluation:

Training → Calibration → Test

The calibration split is used exclusively for temperature scaling and routing-threshold selection. The final test set is not used during calibration.

Calibration Results

Metric Before After
ECE 0.142 0.031

Temperature scaling reduces Expected Calibration Error from 0.142 to 0.031.

Evaluation

The reported experiment uses 300 queries with a stratified 70 / 15 / 15 split:

Split Queries
Training 210
Calibration 30
Test 60
Total 300

The final metrics are calculated on the untouched 60-query test split.

Classification Results

Metric Result
Test queries 60
Tier accuracy 88.3% (53 / 60)
Macro-F1 0.876
Brier score 0.083
ECE after calibration 0.031

Per-Tier Recall

True tier Queries Recall
local_small 26 92.3%
mid_tier 18 83.3%
frontier 16 87.5%

Cost vs Quality

At the selected operating point:

Metric Tollgate
Quality retained 91%
Cost $0.28 / 1k queries
Always-frontier quality 96%
Always-frontier cost $1.00 / 1k queries

This corresponds to approximately 3.6× lower measured routing cost while retaining 91% of the measured answer quality on the test evaluation.

The operating threshold is selected using the calibration split and then applied unchanged to the test split.

Selective Routing

Tollgate can abstain from making an aggressive low-cost routing decision when confidence is low.

Instead, uncertain queries can be escalated to a more capable tier.

Metric Result
Risk-coverage AUC 0.024
Operating coverage 80%
Risk at operating point 4%
Operating threshold τ = 0.25

The threshold τ = 0.25 is an empirical operating point selected from the calibration data. It is not a guarantee for unseen traffic.

Router Latency

Router latency measures the additional inference overhead introduced by Tollgate, excluding downstream LLM generation.

The benchmark was run on CPU after two warm-up runs.

Router p50 p95 p99
Tollgate 124 ms 198 ms 236 ms
Simple heuristic 286 ms 472 ms 621 ms
Small LLM router 1,537 ms 2,745 ms 3,102 ms

Tollgate provides substantially lower routing latency than the evaluated LLM-based router while providing learned classification and calibrated confidence.

Intended Use

Tollgate is intended for systems where multiple LLM tiers are available and the application wants to reduce inference cost while maintaining an acceptable level of answer quality.

Potential applications include:

  • LLM request routing
  • Cost-aware AI infrastructure
  • Multi-model inference systems
  • Model escalation pipelines
  • Confidence-based model selection
  • Selective prediction systems
  • AI gateway and inference optimization

Out-of-Scope Uses

Tollgate should not be treated as a universal quality evaluator or a guarantee that the selected model will produce an acceptable answer.

The model learns tier boundaries represented by the models and data used during training.

If the underlying models, providers, pricing, or task distribution change substantially, the dataset should be regenerated and the router retrained.

Dataset

The experiment uses 300 seed queries derived from:

  • LMSYS-Chat-1M
  • GSM8K
  • MMLU

Each query is evaluated across multiple model tiers.

The collection pipeline is cached and cost-ledgered to avoid unnecessary regeneration of previously collected responses.

Dataset provenance, source licensing, and limitations are documented separately in the project dataset card.

Limitations

Limited Dataset Size

The 300-query experiment demonstrates the complete routing, calibration, and evaluation pipeline, but a substantially larger dataset would provide stronger statistical evidence for production-scale routing behavior.

Judge Bias

Training labels inherit the behavior of the LLM judge.

The current experiment uses one judge and one comparison sample per pair.

Potential sources of bias include:

  • verbosity preference
  • model self-preference
  • answer ordering
  • judge-specific behavior
  • incomplete evaluation of subtle correctness differences

The current experiment does not include a large-scale human audit.

Limited Context

The router currently uses a 512-token input context. Longer queries are truncated before routing.

Provider-Specific Costs

Cost estimates depend on the provider pricing configuration used during the experiment.

The reported cost savings therefore should not be interpreted as universal model pricing.

Auxiliary Labels

needs_tools and needs_rag are auxiliary routing signals. They should not be interpreted as independently human-validated labels until a dedicated hand-labelled evaluation set is introduced.

Reproducibility

The project provides commands for collecting data, training, calibration, evaluation, and generating figures.

Example dataset collection:

uv run tollgate collect-all --limit 300 --yes --concurrency 3

Calibration:

uv run tollgate calibrate \
  --checkpoint checkpoints/laya/best \
  --logits checkpoints/laya/logits_calibration.npz

Evaluation:

uv run tollgate evaluate \
  --run laya \
  --limit 100000

Citation

@software{tollgate,
  title   = {Tollgate: a calibrated LLM cost router},
  author  = {Hemanth Sai Chinthalapudi},
  year    = {2026},
  version = {0.1.0},
  url     = {https://github.com/ihemanthc/tollgate}
}

Acknowledgements

Tollgate uses Laya by Convai Innovations as its base encoder.

Seed prompts are sourced from:

  • LMSYS-Chat-1M
  • GSM8K
  • MMLU

License

Apache-2.0.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ihemanthc/tollgate-router

Finetuned
(162)
this model

Dataset used to train ihemanthc/tollgate-router