cazyundee commited on
Commit
184d40c
Β·
1 Parent(s): 17b0658

feat(search): relevance reranking with Laya, and restore files the last commit clobbered

Browse files

Adds the reranker, and puts back what 4dba37a overwrote.

4dba37a was a mistake: it synced the whole of backend/space/ over the Space,
and this repo's copy of that folder is a STALE FORK. It reverted app.py's
_KEY_HEADERS alias (the x-respite-key / x-ael-key fix), its _is_protected fix,
the raised request_timeout in searxng_settings.yml, and the explicit
mojeek/marginalia/rightdao/mwmbl/presearch engine list. Restored here from
17b0658; this commit touches only the three files it meant to.

The reranker itself:
- rerank.py scores each (query, passage) pair with convaiinnovations/laya, a
non-autoregressive System 1 decision model (421M English / 322M multilingual,
Apache-2.0), using a `noul` question for the calibrated P(yes) that the
passage contains what the query needs.
- The race only ever asked "is this set big enough", so the winner was whoever
answered first and a small index won on speed alone. "what is the capital of
Peru" came back as NPR transcripts about mummy lice and an article about
mosquitoes in Delhi: eight results, none answering it, degraded: false. A
count check cannot see that and neither can lexical overlap, since
"Viceroyalty of Peru" really does contain both words.
- Strictly fail-open. Model missing, disabled, scoring exception, over budget,
or a model that rejects everything: the original set comes back untouched.
A reranker is the only part of search that can make results worse, so every
failure mode has to be a no-op.
- POST /respite/v2/searx/rerank, under the prefix the key gate already covers,
so it needs no second authentication to get right.
- GET /respite/v2/searx/rerank/recent serves the recent score distribution,
because AEL_RERANK_DROP is a guess that should be chosen from real traffic.

πŸ€– Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>

Files changed (3) hide show
  1. requirements.txt +6 -0
  2. rerank.py +231 -0
  3. searx_bridge.py +43 -1
requirements.txt CHANGED
@@ -10,3 +10,9 @@ ddgs
10
  transformers>=4.51.0
11
  accelerate>=1.6.0
12
  httpx
 
 
 
 
 
 
 
10
  transformers>=4.51.0
11
  accelerate>=1.6.0
12
  httpx
13
+
14
+ # Relevance reranker for search results (backend/space/rerank.py). Apache-2.0,
15
+ # non-autoregressive, 421M (english) / 322M (multilingual), one forward pass per
16
+ # document. rerank.py degrades to "no opinion" if this import fails, so a
17
+ # failed install costs search quality, not search itself.
18
+ laya>=0.1.0
rerank.py ADDED
@@ -0,0 +1,231 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Relevance reranking for search results, using Laya (convaiinnovations).
2
+
3
+ Why this exists
4
+ ---------------
5
+ The SearXNG bridge races several engines and hands back the first set that
6
+ looks big enough. Nothing in that path asks whether the results are *about* the
7
+ query, so the fastest engine wins β€” and a small independent index is fast.
8
+ Measured on production, 2026-10-02, from the route's own `source` field:
9
+
10
+ "what is the capital of Peru" -> mwmbl -> NPR transcripts about
11
+ mummy lice, an article about mosquitoes in Delhi, "The Capital Cycle".
12
+ Eight results. None of them answers the question.
13
+ "Peru capital city" -> mwmbl -> a martial arts studio
14
+ in Wellington, an NScale railroad modellers' club.
15
+ "Donald Trump AI regulations policies" -> bing -> "Donald Trump -
16
+ Wikipedia", AP's Trump hub, Reuters' Trump hub.
17
+
18
+ None of that is an outage, so nothing raised an alarm: HTTP 200, `degraded:
19
+ false`, a full result page. It is the failure mode where the search path is
20
+ technically healthy and completely useless, and a count-based "is the set big
21
+ enough" check cannot see it β€” a lexical overlap check cannot either, because
22
+ "Viceroyalty of Peru" really does contain the words "Peru" and "capital" and
23
+ still does not answer the question.
24
+
25
+ A cross-encoder can, because it reads the query and the passage together and
26
+ answers "does this passage answer this question" rather than "do these strings
27
+ overlap". Laya is the right size for that: a non-autoregressive System 1
28
+ decision model, 421M params in English (ModernBERT-large) and 322M for 100+
29
+ languages (mmBERT-base), one forward pass per document, Apache-2.0, calibrated
30
+ probabilities via RLCD. It never generates text, so there is nothing to parse
31
+ and nothing to hallucinate.
32
+
33
+ Design constraints, all of them learned the hard way
34
+ ----------------------------------------------------
35
+ * **Fail open, always.** If the model is missing, slow, or the import breaks,
36
+ the caller gets its results back untouched. Search must not depend on a
37
+ component that is described as "kind of dumb" and can be swapped out.
38
+ * **Never return an empty page.** If the model rejects every result, that is
39
+ not a confident judgement, it is a model that did not understand the query β€”
40
+ so the original set is returned and `starved` is set. That flag is the signal
41
+ to retune DROP_BELOW, not a silent success.
42
+ An earlier version of this file had a floor of 3 surviving results, on the
43
+ theory that a short page is a bad page. That was wrong in the exact case
44
+ this module exists for: the mwmbl result set for "what is the capital of Peru"
45
+ is eight results of which perhaps two are worth showing, and a floor of 3
46
+ throws those two away and serves the eight instead. A short honest page beats
47
+ a full wrong one. The only floor that is genuinely wrong is zero.
48
+ * **Log every score.** The threshold is not a guess to be shipped and admired;
49
+ it is a number to be chosen from the distribution below. RECENT keeps the
50
+ last N (query, scores) so the threshold can be tuned from real traffic.
51
+ * **Deadline-bounded.** Reranking is an enhancement, never a reason to make a
52
+ user wait: the caller passes a budget and this module stops at it.
53
+ """
54
+
55
+ import os
56
+ import threading
57
+ import time
58
+ from collections import deque
59
+
60
+ # ── configuration ───────────────────────────────────────────────────────────
61
+ # Off by default so a bad deploy cannot take search down; the route treats a
62
+ # disabled reranker as "no opinion" and returns what SearXNG gave it.
63
+ ENABLED = os.environ.get("AEL_RERANK", "1") not in ("0", "false", "no")
64
+
65
+ # "english" is 421M ModernBERT-large, 512 tokens, the strongest.
66
+ # "multilingual" is 322M mmBERT-base, ~2x faster, 100+ languages. The route
67
+ # sends whichever it wants; the default here is the stronger English one
68
+ # because Ael's corpus is overwhelmingly English, and non-English queries are
69
+ # still scored correctly by it far more often than not.
70
+ MODEL = os.environ.get("AEL_RERANK_MODEL", "english")
71
+
72
+ # Drop a result when the model thinks it is more likely NOT to answer the query
73
+ # than to answer it. Deliberately conservative: 0.5 is the neutral point of a
74
+ # calibrated probability, so this only removes what the model actively rejects.
75
+ DROP_BELOW = float(os.environ.get("AEL_RERANK_DROP", "0.5"))
76
+
77
+ # Never return an empty page. See the starvation note above β€” this is a floor of
78
+ # one, not a target: a short page is fine, an empty one is not.
79
+ MIN_RESULTS = 1
80
+
81
+ # Recent (query, scores) for threshold tuning. Bounded so a long-lived Space
82
+ # cannot grow without limit.
83
+ RECENT = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))
84
+
85
+ _router = None
86
+ _router_lock = threading.Lock()
87
+ _load_error = None
88
+
89
+
90
+ def _get_router():
91
+ """Load the checkpoint once, lazily. Returns None if it cannot be loaded.
92
+
93
+ Deliberately never raises: a Space that cannot afford the model still
94
+ serves search, and the route needs to be able to ask "any opinion?"
95
+ without handling an exception.
96
+ """
97
+ global _router, _load_error
98
+ if _router is not None or _load_error is not None:
99
+ return _router
100
+ with _router_lock:
101
+ if _router is not None or _load_error is not None:
102
+ return _router
103
+ try:
104
+ from laya import Router
105
+
106
+ # max_loaded=1: one checkpoint resident, which is all this uses.
107
+ _router = Router(max_loaded=1)
108
+ _router.preload([MODEL])
109
+ except Exception as e: # noqa: BLE001 - reported, never raised
110
+ _load_error = f"{type(e).__name__}"
111
+ print(f"[rerank] disabled: cannot load laya ({_load_error}): {e}", flush=True)
112
+ return _router
113
+
114
+
115
+ def status():
116
+ """For the /status route: is this thing actually working right now."""
117
+ router = _get_router()
118
+ return {
119
+ "enabled": ENABLED,
120
+ "ready": router is not None,
121
+ "model": MODEL,
122
+ "drop_below": DROP_BELOW,
123
+ "min_results": MIN_RESULTS,
124
+ "load_error": _load_error,
125
+ "samples_logged": len(RECENT),
126
+ }
127
+
128
+
129
+ def recent_scores(limit=20):
130
+ """Recent (query, scores) pairs, newest first β€” the threshold-tuning data."""
131
+ items = list(RECENT)[-limit:][::-1]
132
+ return [{"q": q, "scores": s} for q, s in items]
133
+
134
+
135
+ def _question(query):
136
+ """The typed question. `noul` returns the calibrated probability of YES.
137
+
138
+ The query is repeated in the instructions rather than passed as context
139
+ because that is the documented shape: state is the passage, questions carry
140
+ the ask. Asking in the imperative ("does this passage contain the answer")
141
+ is what makes it a relevance judgement and not a topic match.
142
+ """
143
+ return {
144
+ "relevant": {
145
+ "type": "noul",
146
+ "instructions": (
147
+ "Does this passage contain the information needed to answer the "
148
+ "web search query below?\n"
149
+ f"Query: {query}\n"
150
+ "Answer no if the passage is about a related topic but does not "
151
+ "actually address the query, and no if it is a navigation page, "
152
+ "a category listing, or a site homepage."
153
+ ),
154
+ }
155
+ }
156
+
157
+
158
+ def _passage(result):
159
+ """The text the model judges. Title carries most of the topical signal;
160
+ snippet is what we have of the page without fetching it."""
161
+ title = str(result.get("title") or "").strip()
162
+ snippet = str(result.get("snippet") or "").strip()
163
+ return (title + "\n" + snippet).strip()
164
+
165
+
166
+ def rerank(query, results, budget_ms=800):
167
+ """Score, reorder and optionally drop `results` for `query`.
168
+
169
+ Returns a dict; never raises. On any problem `results` comes back in its
170
+ original order with `ok: false`, because a reranker that degrades search
171
+ is worse than no reranker.
172
+ """
173
+ out = {
174
+ "ok": False,
175
+ "results": results,
176
+ "model": MODEL,
177
+ "drop_below": DROP_BELOW,
178
+ "dropped": 0,
179
+ "starved": False,
180
+ "scores": [],
181
+ }
182
+ if not ENABLED or not results:
183
+ return out
184
+ router = _get_router()
185
+ if router is None:
186
+ return out
187
+
188
+ t0 = time.time()
189
+ scored = []
190
+ for i, r in enumerate(results):
191
+ passage = _passage(r)
192
+ if not passage:
193
+ # No text to judge. Keep it, unscored, and let the ordering stand.
194
+ scored.append((None, i, r))
195
+ continue
196
+ try:
197
+ res = router.predict(passage, _question(query))
198
+ p = float(res["answers"]["relevant"]["noul"])
199
+ except Exception as e: # noqa: BLE001
200
+ print(f"[rerank] score failed for result {i}: {type(e).__name__}: {e}", flush=True)
201
+ return dict(out, results=results) # fail open, whole set untouched
202
+ if (time.time() - t0) * 1000 > budget_ms:
203
+ # Over budget: return what we have not yet reordered, in order.
204
+ print(f"[rerank] over budget ({budget_ms}ms) after {i} results", flush=True)
205
+ return dict(out, results=results)
206
+ scored.append((p, i, r))
207
+
208
+ if not any(p is not None for p, _, _ in scored):
209
+ return out
210
+
211
+ # Highest score first. Unscored results keep their relative position at the
212
+ # end rather than being dropped: we have no opinion about them.
213
+ scored.sort(key=lambda t: (t[0] is None, -(t[0] if t[0] is not None else 0.0), t[1]))
214
+ kept = [r for p, _, r in scored if p is None or p >= DROP_BELOW]
215
+ dropped = len(scored) - len(kept)
216
+
217
+ out["scores"] = [
218
+ {"title": str(r.get("title") or "")[:80], "score": p}
219
+ for p, _, r in scored
220
+ if p is not None
221
+ ]
222
+ RECENT.append((query[:120], [p for p, _, _ in scored if p is not None]))
223
+
224
+ if dropped and len(kept) < MIN_RESULTS:
225
+ # The model rejected everything, which is a model that did not
226
+ # understand the query rather than a set with nothing in it. Say so
227
+ # instead of pretending.
228
+ return dict(out, results=results, starved=True)
229
+
230
+ out.update(ok=True, results=kept, dropped=dropped)
231
+ return out
searx_bridge.py CHANGED
@@ -2,17 +2,26 @@
2
  container (userspace venv, no root) and proxies queries to it.
3
 
4
  Mounted under /respite/v2/searx/*:
5
- /status β†’ bootstrap state of the runtime
6
  /search β†’ proxy ?q=... queries to SearXNG's JSON API
 
7
 
8
  Engine selection comes from searxng_settings.yml (DDG + Google enabled);
9
  the frontend picks engines per request via the standard `engines` param.
 
 
 
 
10
  """
11
 
 
 
 
12
  import httpx
13
  from fastapi import Request
14
  from fastapi.responses import JSONResponse
15
 
 
16
  import searx_runtime
17
 
18
 
@@ -24,6 +33,7 @@ def register_routes(fa_app):
24
  "installed": searx_runtime.searxng_ready(),
25
  "boot_error": searx_runtime.boot_error(),
26
  "port": searx_runtime.SEARXNG_PORT,
 
27
  })
28
 
29
  @fa_app.get("/respite/v2/searx/search")
@@ -42,3 +52,35 @@ def register_routes(fa_app):
42
  return JSONResponse(r.json(), status_code=r.status_code)
43
  except Exception as e:
44
  return JSONResponse({"error": str(e)}, status_code=502)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  container (userspace venv, no root) and proxies queries to it.
3
 
4
  Mounted under /respite/v2/searx/*:
5
+ /status β†’ bootstrap state of the runtime, and of the reranker
6
  /search β†’ proxy ?q=... queries to SearXNG's JSON API
7
+ /rerank β†’ score/reorder/drop a result set for a query (see rerank.py)
8
 
9
  Engine selection comes from searxng_settings.yml (DDG + Google enabled);
10
  the frontend picks engines per request via the standard `engines` param.
11
+
12
+ /rerank lives under the same prefix as /search on purpose: the Space's key gate
13
+ guards /respite/v2/searx/*, so the reranker inherits that authentication rather
14
+ than needing a second one to be got right.
15
  """
16
 
17
+ import asyncio
18
+ import json
19
+
20
  import httpx
21
  from fastapi import Request
22
  from fastapi.responses import JSONResponse
23
 
24
+ import rerank
25
  import searx_runtime
26
 
27
 
 
33
  "installed": searx_runtime.searxng_ready(),
34
  "boot_error": searx_runtime.boot_error(),
35
  "port": searx_runtime.SEARXNG_PORT,
36
+ "reranker": rerank.status(),
37
  })
38
 
39
  @fa_app.get("/respite/v2/searx/search")
 
52
  return JSONResponse(r.json(), status_code=r.status_code)
53
  except Exception as e:
54
  return JSONResponse({"error": str(e)}, status_code=502)
55
+
56
+ @fa_app.post("/respite/v2/searx/rerank")
57
+ async def _rerank(request: Request):
58
+ """Rerank a result set. Always answers 200 with the results in some
59
+ order β€” the client is a search path, and an exception here would look
60
+ like an outage to whoever is trying to read the page."""
61
+ try:
62
+ body = await request.body()
63
+ payload = json.loads(body or b"{}")
64
+ except Exception as e:
65
+ return JSONResponse({"ok": False, "error": f"bad request: {e}", "results": []})
66
+
67
+ q = str(payload.get("q") or "").strip()
68
+ results = payload.get("results")
69
+ if not q or not isinstance(results, list) or not results:
70
+ return JSONResponse({"ok": False, "error": "q and results are required", "results": results or []})
71
+
72
+ budget_ms = int(payload.get("budget_ms") or 800)
73
+ # rerank() is synchronous and CPU-bound (torch), so it must not run on
74
+ # the event loop β€” that would stall every other Space request, audio
75
+ # generation included, for as long as the model takes.
76
+ try:
77
+ out = await asyncio.to_thread(rerank.rerank, q, results, budget_ms)
78
+ except Exception as e: # noqa: BLE001
79
+ return JSONResponse({"ok": False, "error": f"rerank failed: {e}", "results": results})
80
+ return JSONResponse(out)
81
+
82
+ @fa_app.get("/respite/v2/searx/rerank/recent")
83
+ def _rerank_recent(limit: int = 20):
84
+ """Recent relevance scores, for choosing DROP_BELOW from real traffic
85
+ rather than from a guess."""
86
+ return JSONResponse({"scores": rerank.recent_scores(limit)})