Instructions to use e2rea1/RouteWeaver-4B-cost with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use e2rea1/RouteWeaver-4B-cost with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="e2rea1/RouteWeaver-4B-cost")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("e2rea1/RouteWeaver-4B-cost", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use e2rea1/RouteWeaver-4B-cost with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "e2rea1/RouteWeaver-4B-cost" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "e2rea1/RouteWeaver-4B-cost", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/e2rea1/RouteWeaver-4B-cost
- SGLang
How to use e2rea1/RouteWeaver-4B-cost with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "e2rea1/RouteWeaver-4B-cost" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "e2rea1/RouteWeaver-4B-cost", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "e2rea1/RouteWeaver-4B-cost" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "e2rea1/RouteWeaver-4B-cost", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use e2rea1/RouteWeaver-4B-cost with Docker Model Runner:
docker model run hf.co/e2rea1/RouteWeaver-4B-cost
RouteWeaver-4B-cost — the accuracy–cost frontier
The cost-aware routing policies from RouteWeaver: Weaving Mode Selection and Execution into Unified LLM Routing.
RouteWeaver-4B is the reported
model: it optimizes task reward only. These three are the same method, the
same data and the same training budget, with one number changed — the weight
α on a second reward component that scores how much the chosen workers cost.
Nothing else differs, and no hand-written policy is involved: raising α is the
only dial.
- 📄 Paper: coming soon
- 💻 Code: https://github.com/LaughKing/RouteWeaver
- 🌐 Project page: https://laughking.github.io/RouteWeaver/
- 🤖 Reported model (α = 0): https://huggingface.co/e2rea1/RouteWeaver-4B
- 📊 Datasets: https://huggingface.co/datasets/e2rea1/RouteWeaver-data
- 🗂️ Everything on the Hub: https://huggingface.co/collections/e2rea1/routeweaver-6ac89596851c1c7650f72aa4
The operating points
| revision | α | Avg accuracy | worker cost per query |
|---|---|---|---|
RouteWeaver-4B |
0 | 61.4% | reference |
alpha-0.1 |
0.1 | 60.4% | −23% |
alpha-0.3 |
0.3 | 58.3% | −61% |
alpha-0.5 |
0.5 | 50.2% | −86% |
Accuracy is the unweighted mean of the five evaluation domains (Math, Code, Knowledge, Reason, Recall). Cost is fixed-reference worker inference cost per query, relative to α = 0. The paper reports the cost reduction and the accuracy drop for each α (1.0, 3.1 and 11.2 points); the accuracies above are those drops applied to the 61.4% of α = 0.
Worker cost falls monotonically in α while accuracy falls gradually at first and then sharply, so α = 0.1 and α = 0.3 are the useful trade-offs and α = 0.5 is the edge of the frontier.
The efficiency reward is S_cost = 1/(1 + C/C_ref), normalized by its own group
statistics and mixed with the task reward at weight α, with C_ref frozen
before training as the P95 trajectory cost of a calibration rollout.
Loading one
Each α is a subfolder, not a separate repository:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"e2rea1/RouteWeaver-4B-cost", subfolder="alpha-0.3", dtype="bfloat16")
tok = AutoTokenizer.from_pretrained(
"e2rea1/RouteWeaver-4B-cost", subfolder="alpha-0.3")
To evaluate one the way the paper does:
git clone https://github.com/LaughKing/RouteWeaver && cd RouteWeaver
bash setup_env.sh && conda activate routeweaver
python data/hf_download.py
python - <<'PY'
from huggingface_hub import snapshot_download
print(snapshot_download("e2rea1/RouteWeaver-4B-cost", allow_patterns="alpha-0.3/*"))
PY
CKPT=<that path>/alpha-0.3 TAG=routeweaver-a03 GPU=0 COST=1 bash scripts/eval.sh paper
COST=1 adds the per-call worker-cost columns, which is how the cost numbers
above are produced.
What these models do
They are not chat models and they do not answer questions. Each writes routing actions that the agent loop in the code repository executes against a pool of six frozen workers, addressed only by anonymous ids paired with capability descriptions — the router never sees model names or prices.
<mode>single</mode>
<route model="worker_2">Who wrote the novel Kokoro?</route>
<observation>Natsume Sōseki.</observation>
<answer>Natsume Sōseki</answer>
See the RouteWeaver-4B card
for the action grammar, the worker pool table, decoding settings and the
limitations — all of which apply here unchanged.
What α changes in the policy's behavior
Raising α does not make the router call cheaper workers inside the same plan; it changes the plan. Cost is dominated by how many worker calls a trajectory makes and how long their replies are, so the efficiency reward pushes the policy toward shorter execution — fewer calls, and single-round where it used to go multi-round or agentic. That is why accuracy degrades slowly at first: the queries that did not need the extra calls lose nothing.
Limitations
Everything on the RouteWeaver-4B card applies, plus:
- The cost numbers are reference prices, not a bill.
C_refand the per-worker prices are frozen configuration, so two runs are comparable; they are not what any provider actually charged. - α = 0.5 is reported for the shape of the frontier, not as a recommended operating point — it gives up 11.2 points of accuracy.
- The trade-off was measured on this worker pool and these prices. A pool with a different price spread would put the frontier somewhere else.
Citation
@article{routeweaver2026,
title = {RouteWeaver: Weaving Mode Selection and Execution into Unified LLM Routing},
author = {Wang, Xiaohan and Zhang, Haozhen and Liu, Qingyuan and Feng, Tao and Wang, Wenya},
journal = {arXiv preprint},
year = {2026}
}
Model tree for e2rea1/RouteWeaver-4B-cost
Base model
Qwen/Qwen3-4B-Instruct-2507