Spaces:
Running on Zero
feat(search): relevance reranking with Laya, and restore files the last commit clobbered
Browse filesAdds the reranker, and puts back what 4dba37a overwrote.
4dba37a was a mistake: it synced the whole of backend/space/ over the Space,
and this repo's copy of that folder is a STALE FORK. It reverted app.py's
_KEY_HEADERS alias (the x-respite-key / x-ael-key fix), its _is_protected fix,
the raised request_timeout in searxng_settings.yml, and the explicit
mojeek/marginalia/rightdao/mwmbl/presearch engine list. Restored here from
17b0658; this commit touches only the three files it meant to.
The reranker itself:
- rerank.py scores each (query, passage) pair with convaiinnovations/laya, a
non-autoregressive System 1 decision model (421M English / 322M multilingual,
Apache-2.0), using a `noul` question for the calibrated P(yes) that the
passage contains what the query needs.
- The race only ever asked "is this set big enough", so the winner was whoever
answered first and a small index won on speed alone. "what is the capital of
Peru" came back as NPR transcripts about mummy lice and an article about
mosquitoes in Delhi: eight results, none answering it, degraded: false. A
count check cannot see that and neither can lexical overlap, since
"Viceroyalty of Peru" really does contain both words.
- Strictly fail-open. Model missing, disabled, scoring exception, over budget,
or a model that rejects everything: the original set comes back untouched.
A reranker is the only part of search that can make results worse, so every
failure mode has to be a no-op.
- POST /respite/v2/searx/rerank, under the prefix the key gate already covers,
so it needs no second authentication to get right.
- GET /respite/v2/searx/rerank/recent serves the recent score distribution,
because AEL_RERANK_DROP is a guess that should be chosen from real traffic.
π€ Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
- requirements.txt +6 -0
- rerank.py +231 -0
- searx_bridge.py +43 -1
|
@@ -10,3 +10,9 @@ ddgs
|
|
| 10 |
transformers>=4.51.0
|
| 11 |
accelerate>=1.6.0
|
| 12 |
httpx
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
transformers>=4.51.0
|
| 11 |
accelerate>=1.6.0
|
| 12 |
httpx
|
| 13 |
+
|
| 14 |
+
# Relevance reranker for search results (backend/space/rerank.py). Apache-2.0,
|
| 15 |
+
# non-autoregressive, 421M (english) / 322M (multilingual), one forward pass per
|
| 16 |
+
# document. rerank.py degrades to "no opinion" if this import fails, so a
|
| 17 |
+
# failed install costs search quality, not search itself.
|
| 18 |
+
laya>=0.1.0
|
|
@@ -0,0 +1,231 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Relevance reranking for search results, using Laya (convaiinnovations).
|
| 2 |
+
|
| 3 |
+
Why this exists
|
| 4 |
+
---------------
|
| 5 |
+
The SearXNG bridge races several engines and hands back the first set that
|
| 6 |
+
looks big enough. Nothing in that path asks whether the results are *about* the
|
| 7 |
+
query, so the fastest engine wins β and a small independent index is fast.
|
| 8 |
+
Measured on production, 2026-10-02, from the route's own `source` field:
|
| 9 |
+
|
| 10 |
+
"what is the capital of Peru" -> mwmbl -> NPR transcripts about
|
| 11 |
+
mummy lice, an article about mosquitoes in Delhi, "The Capital Cycle".
|
| 12 |
+
Eight results. None of them answers the question.
|
| 13 |
+
"Peru capital city" -> mwmbl -> a martial arts studio
|
| 14 |
+
in Wellington, an NScale railroad modellers' club.
|
| 15 |
+
"Donald Trump AI regulations policies" -> bing -> "Donald Trump -
|
| 16 |
+
Wikipedia", AP's Trump hub, Reuters' Trump hub.
|
| 17 |
+
|
| 18 |
+
None of that is an outage, so nothing raised an alarm: HTTP 200, `degraded:
|
| 19 |
+
false`, a full result page. It is the failure mode where the search path is
|
| 20 |
+
technically healthy and completely useless, and a count-based "is the set big
|
| 21 |
+
enough" check cannot see it β a lexical overlap check cannot either, because
|
| 22 |
+
"Viceroyalty of Peru" really does contain the words "Peru" and "capital" and
|
| 23 |
+
still does not answer the question.
|
| 24 |
+
|
| 25 |
+
A cross-encoder can, because it reads the query and the passage together and
|
| 26 |
+
answers "does this passage answer this question" rather than "do these strings
|
| 27 |
+
overlap". Laya is the right size for that: a non-autoregressive System 1
|
| 28 |
+
decision model, 421M params in English (ModernBERT-large) and 322M for 100+
|
| 29 |
+
languages (mmBERT-base), one forward pass per document, Apache-2.0, calibrated
|
| 30 |
+
probabilities via RLCD. It never generates text, so there is nothing to parse
|
| 31 |
+
and nothing to hallucinate.
|
| 32 |
+
|
| 33 |
+
Design constraints, all of them learned the hard way
|
| 34 |
+
----------------------------------------------------
|
| 35 |
+
* **Fail open, always.** If the model is missing, slow, or the import breaks,
|
| 36 |
+
the caller gets its results back untouched. Search must not depend on a
|
| 37 |
+
component that is described as "kind of dumb" and can be swapped out.
|
| 38 |
+
* **Never return an empty page.** If the model rejects every result, that is
|
| 39 |
+
not a confident judgement, it is a model that did not understand the query β
|
| 40 |
+
so the original set is returned and `starved` is set. That flag is the signal
|
| 41 |
+
to retune DROP_BELOW, not a silent success.
|
| 42 |
+
An earlier version of this file had a floor of 3 surviving results, on the
|
| 43 |
+
theory that a short page is a bad page. That was wrong in the exact case
|
| 44 |
+
this module exists for: the mwmbl result set for "what is the capital of Peru"
|
| 45 |
+
is eight results of which perhaps two are worth showing, and a floor of 3
|
| 46 |
+
throws those two away and serves the eight instead. A short honest page beats
|
| 47 |
+
a full wrong one. The only floor that is genuinely wrong is zero.
|
| 48 |
+
* **Log every score.** The threshold is not a guess to be shipped and admired;
|
| 49 |
+
it is a number to be chosen from the distribution below. RECENT keeps the
|
| 50 |
+
last N (query, scores) so the threshold can be tuned from real traffic.
|
| 51 |
+
* **Deadline-bounded.** Reranking is an enhancement, never a reason to make a
|
| 52 |
+
user wait: the caller passes a budget and this module stops at it.
|
| 53 |
+
"""
|
| 54 |
+
|
| 55 |
+
import os
|
| 56 |
+
import threading
|
| 57 |
+
import time
|
| 58 |
+
from collections import deque
|
| 59 |
+
|
| 60 |
+
# ββ configuration βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 61 |
+
# Off by default so a bad deploy cannot take search down; the route treats a
|
| 62 |
+
# disabled reranker as "no opinion" and returns what SearXNG gave it.
|
| 63 |
+
ENABLED = os.environ.get("AEL_RERANK", "1") not in ("0", "false", "no")
|
| 64 |
+
|
| 65 |
+
# "english" is 421M ModernBERT-large, 512 tokens, the strongest.
|
| 66 |
+
# "multilingual" is 322M mmBERT-base, ~2x faster, 100+ languages. The route
|
| 67 |
+
# sends whichever it wants; the default here is the stronger English one
|
| 68 |
+
# because Ael's corpus is overwhelmingly English, and non-English queries are
|
| 69 |
+
# still scored correctly by it far more often than not.
|
| 70 |
+
MODEL = os.environ.get("AEL_RERANK_MODEL", "english")
|
| 71 |
+
|
| 72 |
+
# Drop a result when the model thinks it is more likely NOT to answer the query
|
| 73 |
+
# than to answer it. Deliberately conservative: 0.5 is the neutral point of a
|
| 74 |
+
# calibrated probability, so this only removes what the model actively rejects.
|
| 75 |
+
DROP_BELOW = float(os.environ.get("AEL_RERANK_DROP", "0.5"))
|
| 76 |
+
|
| 77 |
+
# Never return an empty page. See the starvation note above β this is a floor of
|
| 78 |
+
# one, not a target: a short page is fine, an empty one is not.
|
| 79 |
+
MIN_RESULTS = 1
|
| 80 |
+
|
| 81 |
+
# Recent (query, scores) for threshold tuning. Bounded so a long-lived Space
|
| 82 |
+
# cannot grow without limit.
|
| 83 |
+
RECENT = deque(maxlen=int(os.environ.get("AEL_RERANK_LOG", "200")))
|
| 84 |
+
|
| 85 |
+
_router = None
|
| 86 |
+
_router_lock = threading.Lock()
|
| 87 |
+
_load_error = None
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def _get_router():
|
| 91 |
+
"""Load the checkpoint once, lazily. Returns None if it cannot be loaded.
|
| 92 |
+
|
| 93 |
+
Deliberately never raises: a Space that cannot afford the model still
|
| 94 |
+
serves search, and the route needs to be able to ask "any opinion?"
|
| 95 |
+
without handling an exception.
|
| 96 |
+
"""
|
| 97 |
+
global _router, _load_error
|
| 98 |
+
if _router is not None or _load_error is not None:
|
| 99 |
+
return _router
|
| 100 |
+
with _router_lock:
|
| 101 |
+
if _router is not None or _load_error is not None:
|
| 102 |
+
return _router
|
| 103 |
+
try:
|
| 104 |
+
from laya import Router
|
| 105 |
+
|
| 106 |
+
# max_loaded=1: one checkpoint resident, which is all this uses.
|
| 107 |
+
_router = Router(max_loaded=1)
|
| 108 |
+
_router.preload([MODEL])
|
| 109 |
+
except Exception as e: # noqa: BLE001 - reported, never raised
|
| 110 |
+
_load_error = f"{type(e).__name__}"
|
| 111 |
+
print(f"[rerank] disabled: cannot load laya ({_load_error}): {e}", flush=True)
|
| 112 |
+
return _router
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def status():
|
| 116 |
+
"""For the /status route: is this thing actually working right now."""
|
| 117 |
+
router = _get_router()
|
| 118 |
+
return {
|
| 119 |
+
"enabled": ENABLED,
|
| 120 |
+
"ready": router is not None,
|
| 121 |
+
"model": MODEL,
|
| 122 |
+
"drop_below": DROP_BELOW,
|
| 123 |
+
"min_results": MIN_RESULTS,
|
| 124 |
+
"load_error": _load_error,
|
| 125 |
+
"samples_logged": len(RECENT),
|
| 126 |
+
}
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
def recent_scores(limit=20):
|
| 130 |
+
"""Recent (query, scores) pairs, newest first β the threshold-tuning data."""
|
| 131 |
+
items = list(RECENT)[-limit:][::-1]
|
| 132 |
+
return [{"q": q, "scores": s} for q, s in items]
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
def _question(query):
|
| 136 |
+
"""The typed question. `noul` returns the calibrated probability of YES.
|
| 137 |
+
|
| 138 |
+
The query is repeated in the instructions rather than passed as context
|
| 139 |
+
because that is the documented shape: state is the passage, questions carry
|
| 140 |
+
the ask. Asking in the imperative ("does this passage contain the answer")
|
| 141 |
+
is what makes it a relevance judgement and not a topic match.
|
| 142 |
+
"""
|
| 143 |
+
return {
|
| 144 |
+
"relevant": {
|
| 145 |
+
"type": "noul",
|
| 146 |
+
"instructions": (
|
| 147 |
+
"Does this passage contain the information needed to answer the "
|
| 148 |
+
"web search query below?\n"
|
| 149 |
+
f"Query: {query}\n"
|
| 150 |
+
"Answer no if the passage is about a related topic but does not "
|
| 151 |
+
"actually address the query, and no if it is a navigation page, "
|
| 152 |
+
"a category listing, or a site homepage."
|
| 153 |
+
),
|
| 154 |
+
}
|
| 155 |
+
}
|
| 156 |
+
|
| 157 |
+
|
| 158 |
+
def _passage(result):
|
| 159 |
+
"""The text the model judges. Title carries most of the topical signal;
|
| 160 |
+
snippet is what we have of the page without fetching it."""
|
| 161 |
+
title = str(result.get("title") or "").strip()
|
| 162 |
+
snippet = str(result.get("snippet") or "").strip()
|
| 163 |
+
return (title + "\n" + snippet).strip()
|
| 164 |
+
|
| 165 |
+
|
| 166 |
+
def rerank(query, results, budget_ms=800):
|
| 167 |
+
"""Score, reorder and optionally drop `results` for `query`.
|
| 168 |
+
|
| 169 |
+
Returns a dict; never raises. On any problem `results` comes back in its
|
| 170 |
+
original order with `ok: false`, because a reranker that degrades search
|
| 171 |
+
is worse than no reranker.
|
| 172 |
+
"""
|
| 173 |
+
out = {
|
| 174 |
+
"ok": False,
|
| 175 |
+
"results": results,
|
| 176 |
+
"model": MODEL,
|
| 177 |
+
"drop_below": DROP_BELOW,
|
| 178 |
+
"dropped": 0,
|
| 179 |
+
"starved": False,
|
| 180 |
+
"scores": [],
|
| 181 |
+
}
|
| 182 |
+
if not ENABLED or not results:
|
| 183 |
+
return out
|
| 184 |
+
router = _get_router()
|
| 185 |
+
if router is None:
|
| 186 |
+
return out
|
| 187 |
+
|
| 188 |
+
t0 = time.time()
|
| 189 |
+
scored = []
|
| 190 |
+
for i, r in enumerate(results):
|
| 191 |
+
passage = _passage(r)
|
| 192 |
+
if not passage:
|
| 193 |
+
# No text to judge. Keep it, unscored, and let the ordering stand.
|
| 194 |
+
scored.append((None, i, r))
|
| 195 |
+
continue
|
| 196 |
+
try:
|
| 197 |
+
res = router.predict(passage, _question(query))
|
| 198 |
+
p = float(res["answers"]["relevant"]["noul"])
|
| 199 |
+
except Exception as e: # noqa: BLE001
|
| 200 |
+
print(f"[rerank] score failed for result {i}: {type(e).__name__}: {e}", flush=True)
|
| 201 |
+
return dict(out, results=results) # fail open, whole set untouched
|
| 202 |
+
if (time.time() - t0) * 1000 > budget_ms:
|
| 203 |
+
# Over budget: return what we have not yet reordered, in order.
|
| 204 |
+
print(f"[rerank] over budget ({budget_ms}ms) after {i} results", flush=True)
|
| 205 |
+
return dict(out, results=results)
|
| 206 |
+
scored.append((p, i, r))
|
| 207 |
+
|
| 208 |
+
if not any(p is not None for p, _, _ in scored):
|
| 209 |
+
return out
|
| 210 |
+
|
| 211 |
+
# Highest score first. Unscored results keep their relative position at the
|
| 212 |
+
# end rather than being dropped: we have no opinion about them.
|
| 213 |
+
scored.sort(key=lambda t: (t[0] is None, -(t[0] if t[0] is not None else 0.0), t[1]))
|
| 214 |
+
kept = [r for p, _, r in scored if p is None or p >= DROP_BELOW]
|
| 215 |
+
dropped = len(scored) - len(kept)
|
| 216 |
+
|
| 217 |
+
out["scores"] = [
|
| 218 |
+
{"title": str(r.get("title") or "")[:80], "score": p}
|
| 219 |
+
for p, _, r in scored
|
| 220 |
+
if p is not None
|
| 221 |
+
]
|
| 222 |
+
RECENT.append((query[:120], [p for p, _, _ in scored if p is not None]))
|
| 223 |
+
|
| 224 |
+
if dropped and len(kept) < MIN_RESULTS:
|
| 225 |
+
# The model rejected everything, which is a model that did not
|
| 226 |
+
# understand the query rather than a set with nothing in it. Say so
|
| 227 |
+
# instead of pretending.
|
| 228 |
+
return dict(out, results=results, starved=True)
|
| 229 |
+
|
| 230 |
+
out.update(ok=True, results=kept, dropped=dropped)
|
| 231 |
+
return out
|
|
@@ -2,17 +2,26 @@
|
|
| 2 |
container (userspace venv, no root) and proxies queries to it.
|
| 3 |
|
| 4 |
Mounted under /respite/v2/searx/*:
|
| 5 |
-
/status β bootstrap state of the runtime
|
| 6 |
/search β proxy ?q=... queries to SearXNG's JSON API
|
|
|
|
| 7 |
|
| 8 |
Engine selection comes from searxng_settings.yml (DDG + Google enabled);
|
| 9 |
the frontend picks engines per request via the standard `engines` param.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
"""
|
| 11 |
|
|
|
|
|
|
|
|
|
|
| 12 |
import httpx
|
| 13 |
from fastapi import Request
|
| 14 |
from fastapi.responses import JSONResponse
|
| 15 |
|
|
|
|
| 16 |
import searx_runtime
|
| 17 |
|
| 18 |
|
|
@@ -24,6 +33,7 @@ def register_routes(fa_app):
|
|
| 24 |
"installed": searx_runtime.searxng_ready(),
|
| 25 |
"boot_error": searx_runtime.boot_error(),
|
| 26 |
"port": searx_runtime.SEARXNG_PORT,
|
|
|
|
| 27 |
})
|
| 28 |
|
| 29 |
@fa_app.get("/respite/v2/searx/search")
|
|
@@ -42,3 +52,35 @@ def register_routes(fa_app):
|
|
| 42 |
return JSONResponse(r.json(), status_code=r.status_code)
|
| 43 |
except Exception as e:
|
| 44 |
return JSONResponse({"error": str(e)}, status_code=502)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
container (userspace venv, no root) and proxies queries to it.
|
| 3 |
|
| 4 |
Mounted under /respite/v2/searx/*:
|
| 5 |
+
/status β bootstrap state of the runtime, and of the reranker
|
| 6 |
/search β proxy ?q=... queries to SearXNG's JSON API
|
| 7 |
+
/rerank β score/reorder/drop a result set for a query (see rerank.py)
|
| 8 |
|
| 9 |
Engine selection comes from searxng_settings.yml (DDG + Google enabled);
|
| 10 |
the frontend picks engines per request via the standard `engines` param.
|
| 11 |
+
|
| 12 |
+
/rerank lives under the same prefix as /search on purpose: the Space's key gate
|
| 13 |
+
guards /respite/v2/searx/*, so the reranker inherits that authentication rather
|
| 14 |
+
than needing a second one to be got right.
|
| 15 |
"""
|
| 16 |
|
| 17 |
+
import asyncio
|
| 18 |
+
import json
|
| 19 |
+
|
| 20 |
import httpx
|
| 21 |
from fastapi import Request
|
| 22 |
from fastapi.responses import JSONResponse
|
| 23 |
|
| 24 |
+
import rerank
|
| 25 |
import searx_runtime
|
| 26 |
|
| 27 |
|
|
|
|
| 33 |
"installed": searx_runtime.searxng_ready(),
|
| 34 |
"boot_error": searx_runtime.boot_error(),
|
| 35 |
"port": searx_runtime.SEARXNG_PORT,
|
| 36 |
+
"reranker": rerank.status(),
|
| 37 |
})
|
| 38 |
|
| 39 |
@fa_app.get("/respite/v2/searx/search")
|
|
|
|
| 52 |
return JSONResponse(r.json(), status_code=r.status_code)
|
| 53 |
except Exception as e:
|
| 54 |
return JSONResponse({"error": str(e)}, status_code=502)
|
| 55 |
+
|
| 56 |
+
@fa_app.post("/respite/v2/searx/rerank")
|
| 57 |
+
async def _rerank(request: Request):
|
| 58 |
+
"""Rerank a result set. Always answers 200 with the results in some
|
| 59 |
+
order β the client is a search path, and an exception here would look
|
| 60 |
+
like an outage to whoever is trying to read the page."""
|
| 61 |
+
try:
|
| 62 |
+
body = await request.body()
|
| 63 |
+
payload = json.loads(body or b"{}")
|
| 64 |
+
except Exception as e:
|
| 65 |
+
return JSONResponse({"ok": False, "error": f"bad request: {e}", "results": []})
|
| 66 |
+
|
| 67 |
+
q = str(payload.get("q") or "").strip()
|
| 68 |
+
results = payload.get("results")
|
| 69 |
+
if not q or not isinstance(results, list) or not results:
|
| 70 |
+
return JSONResponse({"ok": False, "error": "q and results are required", "results": results or []})
|
| 71 |
+
|
| 72 |
+
budget_ms = int(payload.get("budget_ms") or 800)
|
| 73 |
+
# rerank() is synchronous and CPU-bound (torch), so it must not run on
|
| 74 |
+
# the event loop β that would stall every other Space request, audio
|
| 75 |
+
# generation included, for as long as the model takes.
|
| 76 |
+
try:
|
| 77 |
+
out = await asyncio.to_thread(rerank.rerank, q, results, budget_ms)
|
| 78 |
+
except Exception as e: # noqa: BLE001
|
| 79 |
+
return JSONResponse({"ok": False, "error": f"rerank failed: {e}", "results": results})
|
| 80 |
+
return JSONResponse(out)
|
| 81 |
+
|
| 82 |
+
@fa_app.get("/respite/v2/searx/rerank/recent")
|
| 83 |
+
def _rerank_recent(limit: int = 20):
|
| 84 |
+
"""Recent relevance scores, for choosing DROP_BELOW from real traffic
|
| 85 |
+
rather than from a guess."""
|
| 86 |
+
return JSONResponse({"scores": rerank.recent_scores(limit)})
|