File size: 6,992 Bytes
e7abf41
8ab96a8
69fec11
c9992e6
 
 
 
 
 
 
 
e7abf41
69fec11
0d533c0
69fec11
0d533c0
 
69fec11
0d533c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69fec11
 
 
 
 
0d533c0
 
69fec11
0d533c0
 
 
 
f4df775
 
69fec11
0d533c0
 
 
 
69fec11
0d533c0
69fec11
0d533c0
69fec11
 
 
0d533c0
69fec11
0d533c0
 
 
 
 
69fec11
 
 
 
0d533c0
 
 
 
 
 
 
 
 
69fec11
 
 
 
0d533c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69fec11
 
 
 
0d533c0
69fec11
 
 
0d533c0
 
 
 
 
 
69fec11
0d533c0
 
 
69fec11
0d533c0
69fec11
0d533c0
69fec11
 
0d533c0
 
 
 
 
 
 
 
69fec11
c9992e6
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
---
license: apache-2.0
tags:
  - routefm
  - model-routing
  - llm-routing
  - multimodal-routing
  - in-context-learning
  - pytorch
  - safetensors
  - arxiv:2609.37362
---

# RouteFM: Pretrain Once, Route Anywhere

> **[Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing](https://arxiv.org/abs/2609.37362)**<br>
> Guannan Lai and Han-Jia Ye · Nanjing University

[GitHub](https://github.com/LAMDA-Model-Reuse/RouteFM) ·
[Paper](https://arxiv.org/abs/2609.37362) ·
[PDF](https://arxiv.org/pdf/2609.37362)

![RouteFM overview](assets/intro_new.png)

RouteFM is a pretrained in-context model router. Given an anonymous candidate
pool, behavioral observations of candidate quality and relative cost, and a
new target query, it predicts the target-specific quality and relative cost of
every candidate. A frozen RouteFM adapts to new routing environments through
context alone, without target-domain parameter updates.

On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the
strongest non-RouteFM baseline by **2.23 quality points** with only eight
observations per candidate.

## Model description

RouteFM turns LLM routing from repeated local fitting into global routing
pretraining. Candidate identities, providers, parameter counts, and other
explicit identity features are never exposed to the router. Instead, it:

1. builds a compact capability profile for every anonymous candidate from behavioral context;
2. retrieves complementary evidence from both the profile and the original observations for each target query;
3. compares the current pool jointly with a permutation-equivariant Transformer; and
4. predicts target-specific quality and relative cost.

![RouteFM architecture](assets/method.png)

The released architecture uses 24 capability tokens, hidden dimension 256,
and episodic pretraining with candidate-slot permutation. The training
objective combines point prediction, pairwise ranking, and routing regret.
See the [paper](https://arxiv.org/pdf/2609.37362) for the complete method.

## Released variants

| Variant | Compatible query encoder | Dimension | Query modality | Parameters |
| --- | --- | ---: | --- | ---: |
| RouteFM-Qwen | Qwen3-VL-Embedding-8B | 4096 | Joint text and image | 12,781,059 |
| RouteFM-BGE | BAAI/bge-base-en-v1.5, CLS pooling | 768 | Text only | 11,070,467 |

Each variant contains a canonical `model.safetensors` and `config.json`.
The `legacy/` directory contains the exact PyTorch checkpoint bytes from the
GitHub v1.0.0 release. [`manifest.json`](manifest.json) records sizes,
configurations, immutable revisions, and SHA-256 digests.
The root [`config.json`](config.json) indexes both variants and is the standard
Hugging Face query file used for repository download statistics.

The query encoders are frozen external feature extractors. They are not
included in this repository, and RouteFM is not a fine-tune of either encoder.
Users must obtain the encoders or compatible embedding services separately
and comply with their licenses and access terms.

## Usage

Install the official package directly from GitHub:

```bash
python -m pip install git+https://github.com/LAMDA-Model-Reuse/RouteFM.git
```

The package downloads the correct safetensors checkpoint from the immutable
Hub artifact revision, validates its SHA-256 digest, and stores it in the
standard Hugging Face cache.

```bash
routefm-predict --encoder qwen --input my_episode.npz \
  --output predictions.json --device cpu
```

To prefetch both released variants for offline jobs:

```bash
routefm-download --encoder all
```

The input schema, embedding instructions, observation masks, local checkpoint
overrides, and output format are documented in
[`docs/CUSTOM_DATA.md`](https://github.com/LAMDA-Model-Reuse/RouteFM/blob/v1.0.1/docs/CUSTOM_DATA.md).

## Training

Both variants were trained from random initialization in one continuous
10,000-update run. All router modules were trained jointly, while the external
query encoder remained frozen. The configured pretraining mixture is
LLMRouterBench 45%, RouterBench 25%, RouterEval 25%, and MixInstruct 5%.
MMR-Bench is not a pretraining source.

The complete data contract, source proportions, episode construction,
curriculum, and optimization commands are available in
[`docs/PRETRAINING.md`](https://github.com/LAMDA-Model-Reuse/RouteFM/blob/v1.0.1/docs/PRETRAINING.md).
Underlying third-party datasets and embedding models are not redistributed.

## Evaluation

The paper evaluates the frozen Qwen-based RouteFM on held-out RouterEval
queries and on cross-modal transfer to MMR-Bench.

| Evaluation | K=8 | K=16 | K=32 | K=64 | Large |
| --- | ---: | ---: | ---: | ---: | ---: |
| RouterEval (in-domain) | **0.6192** | **0.6433** | **0.6627** | **0.6611** | — |
| MMR-Bench (cross-modal) | **0.7323** | **0.7457** | **0.7487** | **0.7504** | **0.7614** |

These are the paper's matched-protocol results. MMR-Bench uses dataset-wise
five-fold context/target splits in the limited-observation setting and a
40%/60% split in the large-observation setting. RouteFM remains frozen and
uses no evaluation labels for parameter updates.

The GitHub release additionally publishes a fixed seed-31010 split for
artifact auditing. That split is retrospective and post-selected, differs
from the paper's five-fold protocol, and should not be compared directly with
the table above. Its exact IDs, hashes, aggregation rules, and reference
outputs are documented in
[`docs/MMRBENCH_V1.md`](https://github.com/LAMDA-Model-Reuse/RouteFM/blob/v1.0.1/docs/MMRBENCH_V1.md).

## Intended use and limitations

RouteFM is intended for research on routing among a user-supplied candidate
pool with observed context outcomes. It does not call candidate models,
generate query embeddings, or establish that a candidate is safe or suitable.

- Candidate order and embedding family must remain consistent within an episode.
- Each candidate needs at least one valid context observation.
- Predicted cost is relative, not a calibrated monetary or latency estimate.
- The default decision rule ignores predicted cost and selects maximum predicted quality.
- Performance may change across encoder revisions, languages, domains, candidate pools, and context sizes outside the training distribution.

## License

RouteFM source code and released routing weights are licensed under
Apache-2.0. Third-party assets retain their own licenses and terms; see
[`THIRD_PARTY_NOTICES.md`](https://github.com/LAMDA-Model-Reuse/RouteFM/blob/v1.0.1/THIRD_PARTY_NOTICES.md).

## Citation

If you use RouteFM in your research, please cite:

```bibtex
@misc{lai2026pretrainoncerouteanywhere,
  title         = {Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing},
  author        = {Guannan Lai and Han-Jia Ye},
  year          = {2026},
  eprint        = {2609.37362},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2609.37362}
}
```