ToolRouter

A tiny (< 8 MB), CPU-only, sub-millisecond tool router. Give it a user query and a list of tools — each with a name and a description (optionally a few example queries) — and it returns which tool to use, with per-tool probabilities and a confidence, in a strict, structured choice format:

{"state": "will it rain tomorrow in Paris",
 "questions": {"tool": {"type": "choice", "instructions": "Pick the single best tool.",
   "criteria": {"weather": "Forecasts.", "news": "Headlines.", "other": "No tool needed."}}}}
{"model": "toolrouter-7",
 "answers": {"tool": {"type": "choice", "choice": "weather",
   "probabilities": {"weather": 0.999996, "news": 4e-06, "other": 0.0}, "confidence": 0.999996}}}

It is open-set: the tools come from the request, so it works for tools it has never seen. The 7 priority tools — wiki, websearch, weather, news, files, shell, other — additionally have a trained "identity" head and are the most accurate.

  • Try it: the playground Space lets you edit the tool list and query and see the result — it runs the model in your browser (JavaScript port in web/, verified against the Python code on 363 cases).
  • Data: stanley-nv/tool-routing-data (training, validation and test sets).
  • Code, training and evaluation: in this repo (toolrouter/).

Quick start

pip install numpy huggingface_hub            # inference needs only numpy (+ huggingface_hub to download)
git clone https://huggingface.co/stanley-nv/toolrouter && cd toolrouter
python -m toolrouter.cli "what's using port 8080"
python examples/demo_tools.py --tools examples/tools.json            # route over your own tool descriptions, interactive
from toolrouter import ToolRouter
router = ToolRouter.from_pretrained("stanley-nv/toolrouter")           # or ToolRouter("toolrouter.npz")

tools = {
    "stock_quotes": "Real-time share prices, tickers and market indices.",
    "smart_home":   "Control lights, thermostat, locks and other connected devices.",
    "calendar":     {"description": "Read and create meetings.", "examples": ["move my 3pm to Friday"]},   # examples are optional
    "other":        "No tool needed: chit-chat, writing, math, advice.",
}
router.choose("how is Google trading today", tools)
# {'type': 'choice', 'choice': 'stock_quotes', 'probabilities': {...}, 'confidence': 0.95}
router.handle(request)    # full request -> response, as above

How it works

score(query, tool) = 7·cos(q, t_full) + 3·cos(q, t_desc) + 2·cos(q, t_name) + (.5, .5, .5, .25)·lexical_overlap + identity_head

  • q, t_*: frozen static word-piece embeddings from minishlab/potion-base-8M (256-d, MIT), stored int8 in the model file, mean-pooled and normalised; t_full = "name. description". With example queries, cos(q, t_full) becomes the max over description and examples.
  • The semantic weights are fixed (hand-set from an ablation). Only the identity head is trained: a hashed n-gram bag plus the dense query embedding, one vector per priority tool, applied to the 7 priority tools (and aliases such as web, terminal). Training is episodic (query + target + random distractor tools with random name/description styles).
  • Softmax with a calibrated temperature gives probabilities; confidence is the probability of the chosen tool.
  • Files: toolrouter.npz (7.66 MB, includes the backbone). Pure-python WordPiece tokenizer, numpy inference, no PyTorch needed.

Results

Accuracy; all test data are held out (never used for training). Reproduce with python -m toolrouter.eval.evaluate --data-dir <dataset> --models toolrouter.npz --full.

Setting Shipped model Mean of 3 seeds
7 priority tools, independent test routing/test2 (330 queries) 0.936 0.935
same, reworded tool descriptions 0.936 0.934
7 priority tools, routing/test (210 queries) 0.933 0.932
same, reworded tool descriptions 0.938 0.935
10 tools never seen in training (tools/test), all offered at once 0.890 0.890
same, random 5-tool slates 0.933 0.937
same, differently worded descriptions / name only 0.780 / 0.750
descriptions shifted onto the wrong tools (sanity check) 0.060
Predictions with confidence ≥ 0.95 ≈ 98 % correct (170 of 210 priority queries; 56 of 100 unseen-tool queries)
Latency (CPU, 10 tools) ≈ 0.10–0.16 ms warm, ~2 ms for the first request

See docs/REPORT.md for the full development report, ablations and negative results (what did not help: learning the semantic path, fine-tuning on ToolRet, MLP head, MaxSim, IDF pooling, larger/compressed 32M backbone).

Intended use and limitations

  • Pre-routing for agents/assistants: pick one of a handful of tools (best with ≤ 10 tools) before calling a bigger model, with a confidence you can threshold (e.g. fall back to an LLM below ~0.8).
  • Descriptions matter. Name-only or very short descriptions cost 10–15 points; vague catch-all tools (other, websearch, wiki) can steal queries from specific tools when many tools are offered (all ~200 tools at once: ~0.5–0.6).
  • Very short queries (1–2 words, e.g. git status) with vague or meaningless descriptions in a tiny tool list are the weakest case: the generic other tool can win. With the default tool descriptions such queries route correctly; measured over test queries, 2-tool slates (true tool + one random priority tool) are 98.5–99 % accurate even with meaningless descriptions.
  • English only, single-turn, one tool per query (no multi-tool plans, no arguments extraction). wiki vs websearch vs news is inherently fuzzy; labelling conventions are in the dataset card.
  • Evaluation caveat: the test sets were written by Claude (Anthropic) / the author, and the training data were largely written by Claude agents under explicit labelling conventions, so label noise and stylistic overlap are likely; expect lower accuracy on real traffic. Sampling error is about ±3 points. Some design choices were made while looking at the older test sets (routing/test, tools/test); routing/test2 is the cleanest number.
  • Not evaluated for safety/fairness; do not use it as the sole gate for sensitive actions.

Training and reproducing

git clone https://huggingface.co/datasets/stanley-nv/tool-routing-data ../tool-routing-data
pip install -r requirements-train.txt
TOOLROUTER_FORCE_IPV4=1 python -m toolrouter.training.backbone --out backbone/potion_8m.npz   # fetch + convert the frozen backbone
python -m toolrouter.training.train --data-dir ../tool-routing-data --backbone backbone/potion_8m.npz --out toolrouter.npz   # ~4 min on a GPU
python -m toolrouter.eval.evaluate --data-dir ../tool-routing-data --models toolrouter.npz --full
python -m unittest discover tests
# browser/JS port: python -m toolrouter.export_web && python tests/make_parity_cases.py --data-dir ../tool-routing-data && node web/test_parity.mjs tests/parity_cases.json web/model.bin

License and attribution

  • Code and weights: MIT. The model file embeds an int8-quantised copy of the potion-base-8M embeddings (MIT, © MinishLab); please keep that attribution.
  • Training data were generated by Claude (Anthropic) and the author; see the dataset card for provenance and the terms you should review before commercial use. The released training run does not use the ToolRet corpus (an earlier experiment did; it did not help and is excluded).

Citation

@misc{toolrouter2026,
  title  = {ToolRouter: a tiny open-set tool router with a structured choice interface},
  author = {Stanley},
  year   = {2026},
  url    = {https://huggingface.co/stanley-nv/toolrouter}
}

Related work: model2vec / potion, Bag of Tricks for Efficient Text Classification, ToolRet.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stanley-nv/toolrouter

Finetuned
(8)
this model

Dataset used to train stanley-nv/toolrouter

Space using stanley-nv/toolrouter 1

Papers for stanley-nv/toolrouter

Evaluation results