wexumin commited on
Commit
a0e8555
·
verified ·
1 Parent(s): 0dff9cc

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +139 -0
README.md ADDED
@@ -0,0 +1,139 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # OVERFLOWGUARD Checkpoints
2
+
3
+ Router checkpoints for **OVERFLOWGUARD**, a routing framework for detecting when compressed-context representations are likely to fail and selectively falling back to the full context.
4
+
5
+ This repository contains trained routing classifiers for multiple compressors, language models, and training datasets.
6
+
7
+ ## Repository Structure
8
+
9
+ Each checkpoint contains two files:
10
+
11
+ ```text
12
+ {compressor}/{model}/{dataset}/
13
+ ├── router_config.json
14
+ └── routing_clf.pt
15
+ ```
16
+
17
+ Available compressors:
18
+
19
+ - `xrag`
20
+ - `pisco`
21
+ - `oscar`
22
+
23
+ Available models:
24
+
25
+ - `mistral-7b`
26
+ - `mixtral-8x7b`
27
+ - `mistral-24b`
28
+ - `qwen-7b`
29
+ - `solar-7b`
30
+ - `llama-8b`
31
+
32
+ Available training datasets:
33
+
34
+ - `squad`
35
+ - `hotpotqa`
36
+ - `triviaqa`
37
+ - `combined`
38
+
39
+ Not every compressor/model/dataset combination is available.
40
+
41
+ ## Loading a Checkpoint
42
+
43
+ The routers can be loaded directly from the Hugging Face repository using `from_pretrained`:
44
+
45
+ ```python
46
+ from your_router_module import PiscoRouter
47
+
48
+ router = PiscoRouter.from_pretrained(
49
+ "wexumin/OVERFLOWGUARD-checkpoints",
50
+ hf_local_path="pisco/mistral-7b/squad",
51
+ )
52
+ ```
53
+
54
+ For other compressors, use the corresponding router class:
55
+
56
+ ```python
57
+ router = XragRouter.from_pretrained(
58
+ "wexumin/OVERFLOWGUARD-checkpoints",
59
+ hf_local_path="xrag/mistral-7b/squad",
60
+ )
61
+ ```
62
+
63
+ Only the requested `router_config.json` and `routing_clf.pt` are downloaded from the repository.
64
+
65
+ The underlying base language model specified in `router_config.json` is loaded separately.
66
+
67
+ ## Results
68
+
69
+ The table below reports the performance of the released routers.
70
+
71
+ ### Metrics
72
+
73
+ - **AUC** — ROC-AUC of the routing classifier.
74
+ - **Acc_c → Acc_r** — accuracy before routing (`Acc_c`, compressed context) and after routing (`Acc_r`, routed between compressed and full context).
75
+ - **ΔAcc** — absolute accuracy improvement from routing.
76
+ - **R_c** — percentage of examples initially handled using compressed context.
77
+ - **S_tok** — percentage of input tokens saved relative to always using the full context.
78
+
79
+ ### SQuADv2
80
+
81
+ | Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
82
+ |---|---:|---:|---:|---:|---:|
83
+ | xRAG · mistral-7b | 72.3 | 38.5 → **79.0** | +40.6 | 44.3 | 42.6 |
84
+ | xRAG · mixtral-8x7b | 72.0 | 45.5 → **80.4** | +35.0 | 48.1 | 48.0 |
85
+ | OSCAR · mistral-7b | 77.1 | 85.6 → **91.6** | +6.0 | 66.5 | 60.4 |
86
+ | OSCAR · mistral-24b | 80.7 | 87.5 → **94.9** | +7.4 | 66.0 | 59.2 |
87
+ | OSCAR · qwen-7b | 80.1 | 81.7 → **89.8** | +8.1 | 75.0 | 67.2 |
88
+ | PISCO · mistral-7b | 69.2 | 76.5 → **88.8** | +12.3 | 62.1 | 56.1 |
89
+ | PISCO · solar-7b | 64.4 | 80.0 → **88.4** | +8.4 | 68.7 | 60.8 |
90
+ | PISCO · llama-8b | 75.1 | 77.6 → **91.2** | +13.7 | 58.0 | 51.3 |
91
+
92
+ ### HotpotQA
93
+
94
+ | Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
95
+ |---|---:|---:|---:|---:|---:|
96
+ | xRAG · mistral-7b | 73.6 | 50.1 → **79.5** | +29.4 | 49.6 | 48.3 |
97
+ | xRAG · mixtral-8x7b | 74.9 | 56.4 → **79.8** | +23.4 | 56.3 | 55.1 |
98
+ | OSCAR · mistral-7b | 81.2 | 77.0 → **86.7** | +9.7 | 61.4 | 53.5 |
99
+ | OSCAR · mistral-24b | 82.9 | 77.9 → **87.8** | +9.9 | 73.2 | 64.3 |
100
+ | OSCAR · qwen-7b | 80.4 | 72.1 → **84.7** | +12.6 | 60.8 | 52.5 |
101
+ | PISCO · mistral-7b | 74.0 | 76.7 → **87.0** | +10.3 | 59.5 | 52.2 |
102
+ | PISCO · solar-7b | 67.8 | 82.6 → **89.3** | +6.8 | 60.8 | 53.5 |
103
+ | PISCO · llama-8b | 74.0 | 78.1 → **88.0** | +10.0 | 56.9 | 50.8 |
104
+
105
+ ### TriviaQA
106
+
107
+ | Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
108
+ |---|---:|---:|---:|---:|---:|
109
+ | xRAG · mistral-7b | 71.5 | 71.7 → **75.1** | +3.4 | 54.8 | 54.5 |
110
+ | xRAG · mixtral-8x7b | 69.0 | 79.7 → **79.5** | -0.2 | 55.2 | 56.7 |
111
+ | OSCAR · mistral-7b | 90.1 | 68.5 → **69.7** | +1.1 | 66.4 | 61.5 |
112
+ | OSCAR · mistral-24b | 92.9 | 72.3 → **73.3** | +1.0 | 66.6 | 61.6 |
113
+ | OSCAR · qwen-7b | 89.7 | 67.2 → **69.7** | +2.5 | 60.9 | 55.9 |
114
+ | PISCO · mistral-7b | 70.5 | 71.1 → **72.8** | +1.7 | 54.0 | 48.6 |
115
+ | PISCO · solar-7b | 71.8 | 71.4 → **73.4** | +2.0 | 41.5 | 37.5 |
116
+ | PISCO · llama-8b | 68.2 | 72.3 → **71.9** | -0.4 | 54.9 | 49.9 |
117
+
118
+ ### Combined
119
+
120
+ The combined dataset pools examples from SQuADv2, HotpotQA, and TriviaQA.
121
+
122
+ | Compressor · Model | AUC | Acc_c → Acc_r | ΔAcc | R_c | S_tok |
123
+ |---|---:|---:|---:|---:|---:|
124
+ | xRAG · mistral-7b | 80.7 | 53.7 → **80.1** | +26.4 | 54.9 | 53.4 |
125
+ | xRAG · mixtral-8x7b | 82.9 | 60.7 → **85.4** | +24.6 | 51.5 | 50.5 |
126
+ | OSCAR · mistral-7b | 76.9 | 76.2 → **81.9** | +5.7 | 71.8 | 64.3 |
127
+ | OSCAR · mistral-24b | 86.1 | 78.6 → **84.6** | +6.0 | 68.3 | 61.7 |
128
+ | OSCAR · qwen-7b | 83.5 | 73.4 → **81.6** | +8.2 | 63.3 | 56.9 |
129
+ | PISCO · mistral-7b | 74.1 | 74.2 → **83.1** | +8.9 | 62.8 | 55.6 |
130
+ | PISCO · solar-7b | 72.0 | 78.0 → **85.3** | +7.4 | 56.3 | 49.3 |
131
+ | PISCO · llama-8b | 76.1 | 76.7 → **85.1** | +8.4 | 54.0 | 47.6 |
132
+
133
+ ## Notes
134
+ * Except for PISCO-Llama-8B, routers **transfer** across models, generally with only a modest few-percentage-point drop in AUC. **However**, the decision threshold does not transfer, rendering routing degenerate.
135
+
136
+ * The routers are trained independently for each compressor/model/dataset configuration. The `combined` routers are trained on the combined SQuADv2, HotpotQA, and TriviaQA data rather than being transferred from a single-dataset router.
137
+
138
+ * The repository contains only the router configuration and classifier checkpoint; the underlying language models are not included.
139
+