File size: 5,609 Bytes
a0e8555 cbe9506 a0e8555 cbe9506 a0e8555 5ebfffe a0e8555 5ebfffe a0e8555 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | # OVERFLOWGUARD Checkpoints
Router checkpoints for **OVERFLOWGUARD**, a routing framework for detecting when compressed-context representations are likely to fail and selectively falling back to the full context.
This repository contains trained routing classifiers for multiple compressors, language models, and training datasets.
## Repository Structure
Each checkpoint contains two files:
```text
{compressor}/{model}/{dataset}/
├── router_config.json
└── routing_clf.pt
```
Available compressors:
- `xrag`
- `pisco`
- `oscar`
Available models:
- `mistral-7b`
- `mixtral-8x7b`
- `mistral-24b`
- `qwen-7b`
- `solar-7b`
- `llama-8b`
Available training datasets:
- `squad`
- `hotpotqa`
- `triviaqa`
- `combined`
Not every compressor/model/dataset combination is available.
## Loading a Checkpoint
The implementation of the `PiscoRouter` class, along with two other router classes, is available in our [Git repository](https://github.com/s-nlp/overflowguard/tree/main/examples).
The routers can be loaded directly from the Hugging Face repository using `from_pretrained`:
```python
from train_pisco import PiscoRouter
router = PiscoRouter.from_pretrained(
"s-nlp/OVERFLOWGUARD-checkpoints",
hf_local_path="pisco/mistral-7b/squad",
)
```
For other compressors, use the corresponding router class:
```python
router = XragRouter.from_pretrained(
"s-nlp/OVERFLOWGUARD-checkpoints",
hf_local_path="xrag/mistral-7b/squad",
)
```
Only the requested `router_config.json` and `routing_clf.pt` are downloaded from the repository.
The underlying base language model specified in `router_config.json` is loaded separately.
## Results
The table below reports the performance of the released routers.
### Metrics
- **AUC** — ROC-AUC of the routing classifier.
- **Acc_c → Acc_r** — accuracy before routing (`Acc_c`, compressed context) and after routing (`Acc_r`, routed between compressed and full context).
- **ΔAcc** — absolute accuracy improvement from routing.
- **R_c** — percentage of examples initially handled using compressed context.
- **S_tok** — percentage of input tokens saved relative to always using the full context.
### SQuADv2
| Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
|---|---:|---:|---:|---:|---:|
| xRAG · mistral-7b | 72.3 | 38.5 → **79.0** | +40.6 | 44.3 | 42.6 |
| xRAG · mixtral-8x7b | 72.0 | 45.5 → **80.4** | +35.0 | 48.1 | 48.0 |
| OSCAR · mistral-7b | 77.1 | 85.6 → **91.6** | +6.0 | 66.5 | 60.4 |
| OSCAR · mistral-24b | 80.7 | 87.5 → **94.9** | +7.4 | 66.0 | 59.2 |
| OSCAR · qwen-7b | 80.1 | 81.7 → **89.8** | +8.1 | 75.0 | 67.2 |
| PISCO · mistral-7b | 69.2 | 76.5 → **88.8** | +12.3 | 62.1 | 56.1 |
| PISCO · solar-7b | 64.4 | 80.0 → **88.4** | +8.4 | 68.7 | 60.8 |
| PISCO · llama-8b | 75.1 | 77.6 → **91.2** | +13.7 | 58.0 | 51.3 |
### HotpotQA
| Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
|---|---:|---:|---:|---:|---:|
| xRAG · mistral-7b | 73.6 | 50.1 → **79.5** | +29.4 | 49.6 | 48.3 |
| xRAG · mixtral-8x7b | 74.9 | 56.4 → **79.8** | +23.4 | 56.3 | 55.1 |
| OSCAR · mistral-7b | 81.2 | 77.0 → **86.7** | +9.7 | 61.4 | 53.5 |
| OSCAR · mistral-24b | 82.9 | 77.9 → **87.8** | +9.9 | 73.2 | 64.3 |
| OSCAR · qwen-7b | 80.4 | 72.1 → **84.7** | +12.6 | 60.8 | 52.5 |
| PISCO · mistral-7b | 74.0 | 76.7 → **87.0** | +10.3 | 59.5 | 52.2 |
| PISCO · solar-7b | 67.8 | 82.6 → **89.3** | +6.8 | 60.8 | 53.5 |
| PISCO · llama-8b | 74.0 | 78.1 → **88.0** | +10.0 | 56.9 | 50.8 |
### TriviaQA
| Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
|---|---:|---:|---:|---:|---:|
| xRAG · mistral-7b | 71.5 | 71.7 → **75.1** | +3.4 | 54.8 | 54.5 |
| xRAG · mixtral-8x7b | 69.0 | 79.7 → **79.5** | -0.2 | 55.2 | 56.7 |
| OSCAR · mistral-7b | 90.1 | 68.5 → **69.7** | +1.1 | 66.4 | 61.5 |
| OSCAR · mistral-24b | 92.9 | 72.3 → **73.3** | +1.0 | 66.6 | 61.6 |
| OSCAR · qwen-7b | 89.7 | 67.2 → **69.7** | +2.5 | 60.9 | 55.9 |
| PISCO · mistral-7b | 70.5 | 71.1 → **72.8** | +1.7 | 54.0 | 48.6 |
| PISCO · solar-7b | 71.8 | 71.4 → **73.4** | +2.0 | 41.5 | 37.5 |
| PISCO · llama-8b | 68.2 | 72.3 → **71.9** | -0.4 | 54.9 | 49.9 |
### Combined
The combined dataset pools examples from SQuADv2, HotpotQA, and TriviaQA.
| Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
|---|---:|---:|---:|---:|---:|
| xRAG · mistral-7b | 80.7 | 53.7 → **80.1** | +26.4 | 54.9 | 53.4 |
| xRAG · mixtral-8x7b | 82.9 | 60.7 → **85.4** | +24.6 | 51.5 | 50.5 |
| OSCAR · mistral-7b | 76.9 | 76.2 → **81.9** | +5.7 | 71.8 | 64.3 |
| OSCAR · mistral-24b | 86.1 | 78.6 → **84.6** | +6.0 | 68.3 | 61.7 |
| OSCAR · qwen-7b | 83.5 | 73.4 → **81.6** | +8.2 | 63.3 | 56.9 |
| PISCO · mistral-7b | 74.1 | 74.2 → **83.1** | +8.9 | 62.8 | 55.6 |
| PISCO · solar-7b | 72.0 | 78.0 → **85.3** | +7.4 | 56.3 | 49.3 |
| PISCO · llama-8b | 76.1 | 76.7 → **85.1** | +8.4 | 54.0 | 47.6 |
## Notes
* Except for PISCO-Llama-8B, routers **transfer** across models, generally with only a modest few-percentage-point drop in AUC. **However**, the decision threshold does not transfer, rendering routing degenerate.
* The routers are trained independently for each compressor/model/dataset configuration. The `combined` routers are trained on the combined SQuADv2, HotpotQA, and TriviaQA data rather than being transferred from a single-dataset router.
* The repository contains only the router configuration and classifier checkpoint; the underlying language models are not included.
|