File size: 9,600 Bytes
d7ae9a7
968c251
 
 
 
d7ae9a7
 
968c251
 
 
d7ae9a7
 
 
 
 
 
 
 
 
 
 
 
 
 
968c251
 
d7ae9a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
aa3d42d
 
 
 
d7ae9a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
968c251
 
d7ae9a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16854e1
 
 
d7ae9a7
 
 
 
 
 
 
 
 
968c251
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
---
base_model: mixedbread-ai/mxbai-edge-colbert-v0-32m
datasets:
- reasonir/reasonir-data
- hanhainebula/bge-reasoner-data
language:
- en
library_name: PyLate
license: cc-by-nc-4.0
pipeline_tag: text-retrieval
tags:
- ColBERT
- PyLate
- sentence-transformers
- sentence-similarity
- feature-extraction
- late-interaction
- reasoning-retrieval
- edge
- loss:CachedContrastive
---

# SmallReason-ColBERT (32M)

**Paper**: [SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning Intensive Retrieval](https://huggingface.co/papers/2609.29652)

An ultra-small late-interaction retriever for **reasoning-intensive** retrieval.
32M parameters, plus a **129-parameter query-side importance head**.

**21.41 mean nDCG@10 on BRIGHT** β€” above every ≀33M ColBERT we evaluated, and within
1.21 of the 4.7Γ— larger 150M Reason-ModernColBERT.

---

## ⚠ Read this before loading

This model is a ColBERT base **plus a small importance head** stored in
`importance_head/`. The head is *not* part of `modules.json`, so a standard PyLate /
sentence-transformers load **silently ignores it** and gives you the un-headed base:

| How you load it | What you get | BRIGHT mean |
|---|---|---:|
| `pylate.models.ColBERT(...)` β€” plain load | base only, head ignored, **no error** | 19.61 |
| `WeightedColBERT.from_base(...)` β€” see below | full model | **21.41** |

There is no warning when the head is skipped, so if you are reproducing the paper
number, use the second path. `WeightedColBERT.from_base` resolves the head from this
repo automatically and **raises** if it cannot find one, so that path cannot fail
silently. Pass `require_head=False` if you deliberately want the base.

---

## Usage

The loader is a single file, [`weighted_colbert.py`](https://github.com/DataScience-UIBK/SmallReason-ColBERT/blob/main/src/weighted_colbert.py),
from the companion repository.

```python
from weighted_colbert import WeightedColBERT

model = WeightedColBERT.from_base(
    "DataScience-UIBK/SmallReason-ColBERT-32M",   # auto-detects importance_head/
    query_length=256,
    document_length=2048,
    device="cuda:0",
)

queries = ["What factors affect the number of Hadley cells a planet has, and how?"]
docs = [
    "Hadley cells are driven by differential solar heating; their number scales with "
    "planetary rotation rate and atmospheric depth.",
    "The best pasta recipe uses semolina flour and plenty of salted boiling water.",
]

q_embs, q_weights = model.encode(queries, is_query=True, return_weights=True)
d_embs = model.encode(docs, is_query=False)

for i, d in enumerate(d_embs):
    score = WeightedColBERT.weighted_maxsim(q_embs[0], q_weights[0], d)
    print(i, float(score))
```

`weighted_maxsim` implements the evaluation-time score

$$s(q,d) = \frac{\sum_t w_t \cdot \max_j \mathbf{Q}_t \cdot \mathbf{D}_j}{\sum_t w_t}$$

where $w_t = \sigma(\mathbf{W}\mathbf{Q}_t + b)$ is the learned per-query-token gate.
The `1/\sum_t w_t` factor is constant across documents for a fixed query, so it does not
change ranking β€” it only keeps scores comparable across queries of different length.

### Base only (no head)

If you want the reasoning-tuned base without the gate (19.61 on BRIGHT), load it as an
ordinary PyLate ColBERT β€” the head files are simply unused:

```python
from pylate import models
base = models.ColBERT("DataScience-UIBK/SmallReason-ColBERT-32M",
                      query_length=256, document_length=2048)
```

---

## Results

### BRIGHT (nDCG@10 Γ—100)

Evaluated with brute-force MaxSim, `query_length=256` (Pony: 32), `document_length=2048`.

| Split | upstream 32M | base (no head) | **SmallReason-ColBERT** |
|---|---:|---:|---:|
| biology | 28.70 | 33.16 | **34.17** |
| earth_science | 42.29 | 44.28 | **45.03** |
| economics | 17.65 | **20.25** | 19.99 |
| psychology | 21.93 | 24.91 | **24.94** |
| robotics | 18.09 | **18.65** | 18.14 |
| stackoverflow | 16.49 | 16.66 | **17.21** |
| sustainable_living | 18.64 | 20.11 | **21.07** |
| pony | 12.90 | **22.77** | 19.33 |
| leetcode | 16.15 | 17.40 | **29.98** |
| aops | 9.80 | 4.89 | **10.29** |
| theoremqa_questions | 12.51 | 9.04 | **13.00** |
| theoremqa_theorems | 2.76 | 3.19 | **3.74** |
| **Mean** | 18.16 | 19.61 | **21.41** |

The head is worth **+1.80** mean nDCG@10 over the same base, concentrated in the long,
symbol-dense splits: LeetCode +12.58, AoPS +5.40, TheoremQA-questions +3.96.

### Reference points

| Model | Params | BRIGHT mean |
|---|---:|---:|
| **SmallReason-ColBERT** | **32M** | **21.41** |
| answerai-colbert-small-v1 | 33M | 18.49 |
| mxbai-edge-colbert-v0-17m | 17M | 18.60 |
| GTE-ModernColBERT-v1 | 150M | 21.72 |
| Reason-ModernColBERT | 150M | 21.97 (our protocol) / 22.62 (published) |

### NanoBEIR sanity (classical IR)

The gate is trained on long reasoning queries, so it is expected to give a little back
on short keyword queries. It does, but not much:

| Model | All 13 | Excl. Touche-2020 |
|---|---:|---:|
| upstream 32M | 60.47 | 65.51 |
| base (no head) | 60.93 | 65.35 |
| **SmallReason-ColBERT** | 60.00 | 65.00 |

---

## How it works

Three stages, on top of `mixedbread-ai/mxbai-edge-colbert-v0-32m`:

1. **Widen the projection** 64 β†’ 128 dims. The first 64 rows are inherited; the new 64
   are initialised from `N(0, σ²)` with `Οƒ` at 10% of the original weight-matrix std β€”
   small enough to leave MaxSim β‰ˆ unchanged at step 0, non-zero so the new channels
   actually receive gradient.
2. **Two-stage base training** β€” a varied-length warmup on ReasonIR-VL, then a
   hard-negative polish on merged ReasonIR-HQ + BGE-Reasoner. Both stages use PyLate's
   `CachedContrastive` loss over in-batch negatives.
3. **Importance head** β€” freeze the base, train a single `Linear(128, 1)` + sigmoid
   (129 parameters) to weight each query token.

### The one non-obvious trick

The head is **trained against the un-normalised** weighted score
`Ξ£ w_t Β· max_j(Q_t Β· D_j)` but **evaluated against the length-normalised** one.

This asymmetry is the single most consequential choice in the recipe. Train against the
normalised score instead and the per-pair score difference is bounded by one token's
cosine range, the cross-entropy gradient collapses, the loss stalls near `ln 2`, the
gates never leave their initialisation β€” and BRIGHT drops by **3.59** nDCG@10.

The head is initialised `W = 0`, `b = 5`, so every gate starts at `Οƒ(5) β‰ˆ 0.993` and the
head is a no-op against the frozen base at step zero.

### What the head actually learns

Not soft-IDF. Across ~199K BRIGHT query tokens the gate–IDF Spearman correlation is
**ρ = βˆ’0.02** β€” statistically detectable, practically zero. Per-split mean gate sits in
0.43–0.47 with std β‰ˆ 0.10: the head is a soft re-weighting, not a selector. A fixed
IDF gate on the same base reaches only 20.06, against 21.41 for the learned head.

---

## Training

| | Warmup | Polish | Head |
|---|---|---|---|
| Data | ReasonIR-VL (~245K) | merged ReasonIR-HQ + BGE-Reasoner (~2.7M) | same merged set |
| Loss | CachedContrastive | CachedContrastive | CE over `[s_pos, s_neg]` |
| LR | 1e-5 | 5e-6 | 5e-4 (AdamW, wd=0) |
| Batch | 32/GPU Γ— accum 4 Γ— 8 GPU | 32/GPU Γ— accum 2 Γ— 8 GPU | 16 triples/step, 1 GPU |
| Steps | 1 epoch (~8 h) | 1 epoch (~16 h) | 3,000 steps (~12 min) |
| Lengths | q 256 / doc 2048 | q 256 / doc 2048 | q 256 / doc 2048 |
| Precision | bf16 + FA2 | bf16 + FA2 | fp32 head, frozen bf16 base |

Base training: 8Γ— H100 across two nodes, ~24 h total. Head training: one H100, ~12 min.

---

## Limitations

- **Scale.** The recipe was developed and validated at 32M. It does not transfer for
  free β€” the same head at 17M gives **no** gain.
- **Frozen base.** The head is trained on a frozen base; joint fine-tuning is unexplored.
- **Late-interaction cost.** The head is nearly free, but the model still carries
  multi-vector storage and scoring costs. The efficiency claim is about parameter count,
  not about matching single-vector retrieval.
- **Short queries.** Pony (32-token queries) regresses relative to the un-headed base β€”
  a per-token gate needs tokens to discriminate between.
- **Oblique queries.** On OBLIQ-Bench (stance / intent / tip-of-the-tongue) the model is
  near zero (mean 3.66) and is beaten by every baseline there. Reported as a deliberate
  negative result; embedding similarity is the wrong tool for that class of query.
- **Synthetic teacher data.** Training data is synthetic with cross-encoder-mined hard
  negatives; biases in that mining can propagate.

## License

**CC-BY-NC-4.0**, inherited from the ReasonIR and BGE-Reasoner training data.
The upstream base model (`mixedbread-ai/mxbai-edge-colbert-v0-32m`) is Apache-2.0, and
the companion training/inference **code** is released under Apache-2.0 β€” but these
**weights** are non-commercial.

## Citation

```bibtex
@inproceedings{smallreason-colbert,
  title     = {SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever
               for Reasoning Intensive Retrieval},
  author    = {Abdallah, Abdelrahman and Ali, Mohammed and Jatowt, Adam},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
               Natural Language Processing (EMNLP)},
  year      = {2026}
}
```

## Acknowledgements

Thanks to Antoine Chaffin (LightOn, Reason-ModernColBERT) for flagging the upstream
`2_Dense/use_residual` config bug in `mxbai-edge-colbert-v0-32m` β€” the base weights were
trained with a residual on that layer while the shipped config said otherwise. This
model uses the patched config (`use_residual: true`).