Instructions to use ihemanthc/tollgate-router with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use ihemanthc/tollgate-router with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Tollgate
A calibrated LLM cost router that predicts the minimum sufficient model tier for each query.
Tollgate fine-tunes Laya to classify incoming queries into one of three routing tiers:
local_smallmid_tierfrontier
The router produces calibrated probabilities and uses a confidence threshold to decide whether a query can safely be served at the predicted tier or should be escalated to a more capable model.
Model Details
| Property | Value |
|---|---|
| Base model | Laya |
| Task | LLM tier routing |
| Model type | Fine-tuned encoder |
| Input context | 512 tokens |
| Output | Tier probabilities + routing signals |
| Tiers | local_small, mid_tier, frontier |
| Language | English |
| License | Apache-2.0 |
Tollgate is designed around a simple objective:
Reduce LLM inference cost while retaining acceptable answer quality.
How It Works
Each training query is answered by every model tier at temperature 0 using the same output cap.
An LLM judge compares each cheaper answer against the frontier answer and assigns one of three judgments:
equivalentacceptableworse
The training label is the cheapest tier rated equivalent or acceptable.
If no cheaper tier is sufficient, the query receives the frontier label.
Laya is then fine-tuned to predict the minimum sufficient tier together with two auxiliary routing signals:
needs_toolsneeds_rag
All predictions are produced in a single forward pass.
Calibration
Tollgate uses temperature scaling to improve the reliability of its confidence estimates.
Calibration is performed separately from model training and final evaluation:
Training → Calibration → Test
The calibration split is used exclusively for temperature scaling and routing-threshold selection. The final test set is not used during calibration.
Calibration Results
| Metric | Before | After |
|---|---|---|
| ECE | 0.142 | 0.031 |
Temperature scaling reduces Expected Calibration Error from 0.142 to 0.031.
Evaluation
The reported experiment uses 300 queries with a stratified 70 / 15 / 15 split:
| Split | Queries |
|---|---|
| Training | 210 |
| Calibration | 30 |
| Test | 60 |
| Total | 300 |
The final metrics are calculated on the untouched 60-query test split.
Classification Results
| Metric | Result |
|---|---|
| Test queries | 60 |
| Tier accuracy | 88.3% (53 / 60) |
| Macro-F1 | 0.876 |
| Brier score | 0.083 |
| ECE after calibration | 0.031 |
Per-Tier Recall
| True tier | Queries | Recall |
|---|---|---|
local_small |
26 | 92.3% |
mid_tier |
18 | 83.3% |
frontier |
16 | 87.5% |
Cost vs Quality
At the selected operating point:
| Metric | Tollgate |
|---|---|
| Quality retained | 91% |
| Cost | $0.28 / 1k queries |
| Always-frontier quality | 96% |
| Always-frontier cost | $1.00 / 1k queries |
This corresponds to approximately 3.6× lower measured routing cost while retaining 91% of the measured answer quality on the test evaluation.
The operating threshold is selected using the calibration split and then applied unchanged to the test split.
Selective Routing
Tollgate can abstain from making an aggressive low-cost routing decision when confidence is low.
Instead, uncertain queries can be escalated to a more capable tier.
| Metric | Result |
|---|---|
| Risk-coverage AUC | 0.024 |
| Operating coverage | 80% |
| Risk at operating point | 4% |
| Operating threshold | τ = 0.25 |
The threshold τ = 0.25 is an empirical operating point selected from the calibration data. It is not a guarantee for unseen traffic.
Router Latency
Router latency measures the additional inference overhead introduced by Tollgate, excluding downstream LLM generation.
The benchmark was run on CPU after two warm-up runs.
| Router | p50 | p95 | p99 |
|---|---|---|---|
| Tollgate | 124 ms | 198 ms | 236 ms |
| Simple heuristic | 286 ms | 472 ms | 621 ms |
| Small LLM router | 1,537 ms | 2,745 ms | 3,102 ms |
Tollgate provides substantially lower routing latency than the evaluated LLM-based router while providing learned classification and calibrated confidence.
Intended Use
Tollgate is intended for systems where multiple LLM tiers are available and the application wants to reduce inference cost while maintaining an acceptable level of answer quality.
Potential applications include:
- LLM request routing
- Cost-aware AI infrastructure
- Multi-model inference systems
- Model escalation pipelines
- Confidence-based model selection
- Selective prediction systems
- AI gateway and inference optimization
Out-of-Scope Uses
Tollgate should not be treated as a universal quality evaluator or a guarantee that the selected model will produce an acceptable answer.
The model learns tier boundaries represented by the models and data used during training.
If the underlying models, providers, pricing, or task distribution change substantially, the dataset should be regenerated and the router retrained.
Dataset
The experiment uses 300 seed queries derived from:
- LMSYS-Chat-1M
- GSM8K
- MMLU
Each query is evaluated across multiple model tiers.
The collection pipeline is cached and cost-ledgered to avoid unnecessary regeneration of previously collected responses.
Dataset provenance, source licensing, and limitations are documented separately in the project dataset card.
Limitations
Limited Dataset Size
The 300-query experiment demonstrates the complete routing, calibration, and evaluation pipeline, but a substantially larger dataset would provide stronger statistical evidence for production-scale routing behavior.
Judge Bias
Training labels inherit the behavior of the LLM judge.
The current experiment uses one judge and one comparison sample per pair.
Potential sources of bias include:
- verbosity preference
- model self-preference
- answer ordering
- judge-specific behavior
- incomplete evaluation of subtle correctness differences
The current experiment does not include a large-scale human audit.
Limited Context
The router currently uses a 512-token input context. Longer queries are truncated before routing.
Provider-Specific Costs
Cost estimates depend on the provider pricing configuration used during the experiment.
The reported cost savings therefore should not be interpreted as universal model pricing.
Auxiliary Labels
needs_tools and needs_rag are auxiliary routing signals. They should not be interpreted as independently human-validated labels until a dedicated hand-labelled evaluation set is introduced.
Reproducibility
The project provides commands for collecting data, training, calibration, evaluation, and generating figures.
Example dataset collection:
uv run tollgate collect-all --limit 300 --yes --concurrency 3
Calibration:
uv run tollgate calibrate \
--checkpoint checkpoints/laya/best \
--logits checkpoints/laya/logits_calibration.npz
Evaluation:
uv run tollgate evaluate \
--run laya \
--limit 100000
Citation
@software{tollgate,
title = {Tollgate: a calibrated LLM cost router},
author = {Hemanth Sai Chinthalapudi},
year = {2026},
version = {0.1.0},
url = {https://github.com/ihemanthc/tollgate}
}
Acknowledgements
Tollgate uses Laya by Convai Innovations as its base encoder.
Seed prompts are sourced from:
- LMSYS-Chat-1M
- GSM8K
- MMLU
License
Apache-2.0.
- Downloads last month
- -
Model tree for ihemanthc/tollgate-router
Base model
convaiinnovations/laya