Instructions to use stanley-nv/toolrouter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Model2Vec
How to use stanley-nv/toolrouter with Model2Vec:
from model2vec import StaticModel model = StaticModel.from_pretrained("stanley-nv/toolrouter") embeddings = model.encode(["It's dangerous to go alone!", "It's a secret to everybody."]) print(embeddings.shape) - Notebooks
- Google Colab
- Kaggle
ToolRouter
A tiny (< 8 MB), CPU-only, sub-millisecond tool router. Give it a user query and a list of tools — each with a name and a description (optionally a few example queries) — and it returns which tool to use, with per-tool probabilities and a confidence, in a strict, structured choice format:
{"state": "will it rain tomorrow in Paris",
"questions": {"tool": {"type": "choice", "instructions": "Pick the single best tool.",
"criteria": {"weather": "Forecasts.", "news": "Headlines.", "other": "No tool needed."}}}}
{"model": "toolrouter-7",
"answers": {"tool": {"type": "choice", "choice": "weather",
"probabilities": {"weather": 0.999996, "news": 4e-06, "other": 0.0}, "confidence": 0.999996}}}
It is open-set: the tools come from the request, so it works for tools it has never seen. The 7 priority tools — wiki, websearch, weather, news, files, shell, other — additionally have a trained "identity" head and are the most accurate.
- Try it: the playground Space lets you edit the tool list and query and see the result — it runs the model in your browser (JavaScript port in
web/, verified against the Python code on 363 cases). - Data:
stanley-nv/tool-routing-data(training, validation and test sets). - Code, training and evaluation: in this repo (
toolrouter/).
Quick start
pip install numpy huggingface_hub # inference needs only numpy (+ huggingface_hub to download)
git clone https://huggingface.co/stanley-nv/toolrouter && cd toolrouter
python -m toolrouter.cli "what's using port 8080"
python examples/demo_tools.py --tools examples/tools.json # route over your own tool descriptions, interactive
from toolrouter import ToolRouter
router = ToolRouter.from_pretrained("stanley-nv/toolrouter") # or ToolRouter("toolrouter.npz")
tools = {
"stock_quotes": "Real-time share prices, tickers and market indices.",
"smart_home": "Control lights, thermostat, locks and other connected devices.",
"calendar": {"description": "Read and create meetings.", "examples": ["move my 3pm to Friday"]}, # examples are optional
"other": "No tool needed: chit-chat, writing, math, advice.",
}
router.choose("how is Google trading today", tools)
# {'type': 'choice', 'choice': 'stock_quotes', 'probabilities': {...}, 'confidence': 0.95}
router.handle(request) # full request -> response, as above
How it works
score(query, tool) = 7·cos(q, t_full) + 3·cos(q, t_desc) + 2·cos(q, t_name) + (.5, .5, .5, .25)·lexical_overlap + identity_head
q,t_*: frozen static word-piece embeddings fromminishlab/potion-base-8M(256-d, MIT), stored int8 in the model file, mean-pooled and normalised;t_full= "name. description". With example queries,cos(q, t_full)becomes the max over description and examples.- The semantic weights are fixed (hand-set from an ablation). Only the identity head is trained: a hashed n-gram bag plus the dense query embedding, one vector per priority tool, applied to the 7 priority tools (and aliases such as
web,terminal). Training is episodic (query + target + random distractor tools with random name/description styles). - Softmax with a calibrated temperature gives
probabilities;confidenceis the probability of the chosen tool. - Files:
toolrouter.npz(7.66 MB, includes the backbone). Pure-python WordPiece tokenizer, numpy inference, no PyTorch needed.
Results
Accuracy; all test data are held out (never used for training). Reproduce with python -m toolrouter.eval.evaluate --data-dir <dataset> --models toolrouter.npz --full.
| Setting | Shipped model | Mean of 3 seeds |
|---|---|---|
7 priority tools, independent test routing/test2 (330 queries) |
0.936 | 0.935 |
| same, reworded tool descriptions | 0.936 | 0.934 |
7 priority tools, routing/test (210 queries) |
0.933 | 0.932 |
| same, reworded tool descriptions | 0.938 | 0.935 |
10 tools never seen in training (tools/test), all offered at once |
0.890 | 0.890 |
| same, random 5-tool slates | 0.933 | 0.937 |
| same, differently worded descriptions / name only | 0.780 / 0.750 | |
| descriptions shifted onto the wrong tools (sanity check) | 0.060 | |
| Predictions with confidence ≥ 0.95 | ≈ 98 % correct (170 of 210 priority queries; 56 of 100 unseen-tool queries) | |
| Latency (CPU, 10 tools) | ≈ 0.10–0.16 ms warm, ~2 ms for the first request |
See docs/REPORT.md for the full development report, ablations and negative results (what did not help: learning the semantic path, fine-tuning on ToolRet, MLP head, MaxSim, IDF pooling, larger/compressed 32M backbone).
Intended use and limitations
- Pre-routing for agents/assistants: pick one of a handful of tools (best with ≤ 10 tools) before calling a bigger model, with a confidence you can threshold (e.g. fall back to an LLM below ~0.8).
- Descriptions matter. Name-only or very short descriptions cost 10–15 points; vague catch-all tools (
other,websearch,wiki) can steal queries from specific tools when many tools are offered (all ~200 tools at once: ~0.5–0.6). - Very short queries (1–2 words, e.g.
git status) with vague or meaningless descriptions in a tiny tool list are the weakest case: the genericothertool can win. With the default tool descriptions such queries route correctly; measured over test queries, 2-tool slates (true tool + one random priority tool) are 98.5–99 % accurate even with meaningless descriptions. - English only, single-turn, one tool per query (no multi-tool plans, no arguments extraction).
wikivswebsearchvsnewsis inherently fuzzy; labelling conventions are in the dataset card. - Evaluation caveat: the test sets were written by Claude (Anthropic) / the author, and the training data were largely written by Claude agents under explicit labelling conventions, so label noise and stylistic overlap are likely; expect lower accuracy on real traffic. Sampling error is about ±3 points. Some design choices were made while looking at the older test sets (
routing/test,tools/test);routing/test2is the cleanest number. - Not evaluated for safety/fairness; do not use it as the sole gate for sensitive actions.
Training and reproducing
git clone https://huggingface.co/datasets/stanley-nv/tool-routing-data ../tool-routing-data
pip install -r requirements-train.txt
TOOLROUTER_FORCE_IPV4=1 python -m toolrouter.training.backbone --out backbone/potion_8m.npz # fetch + convert the frozen backbone
python -m toolrouter.training.train --data-dir ../tool-routing-data --backbone backbone/potion_8m.npz --out toolrouter.npz # ~4 min on a GPU
python -m toolrouter.eval.evaluate --data-dir ../tool-routing-data --models toolrouter.npz --full
python -m unittest discover tests
# browser/JS port: python -m toolrouter.export_web && python tests/make_parity_cases.py --data-dir ../tool-routing-data && node web/test_parity.mjs tests/parity_cases.json web/model.bin
License and attribution
- Code and weights: MIT. The model file embeds an int8-quantised copy of the
potion-base-8Membeddings (MIT, © MinishLab); please keep that attribution. - Training data were generated by Claude (Anthropic) and the author; see the dataset card for provenance and the terms you should review before commercial use. The released training run does not use the ToolRet corpus (an earlier experiment did; it did not help and is excluded).
Citation
@misc{toolrouter2026,
title = {ToolRouter: a tiny open-set tool router with a structured choice interface},
author = {Stanley},
year = {2026},
url = {https://huggingface.co/stanley-nv/toolrouter}
}
Related work: model2vec / potion, Bag of Tricks for Efficient Text Classification, ToolRet.
Model tree for stanley-nv/toolrouter
Base model
minishlab/potion-base-8MDataset used to train stanley-nv/toolrouter
Space using stanley-nv/toolrouter 1
Papers for stanley-nv/toolrouter
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
Bag of Tricks for Efficient Text Classification
Evaluation results
- accuracy on routing/test2 (independentself-reported0.936
- accuracy on tools/testtest set self-reported0.890