Claude Claude Opus 5 commited on
Commit
e44fdef
·
unverified ·
1 Parent(s): 9841346

Move orchestration out of the route handlers into a service layer

Browse files

The backend policy - which backend to try, in what order, when to fall
back, what counts as evidence a backend is unhealthy, when a patent is
absent rather than merely unreachable - was written inline in FastAPI route
handlers. It is not HTTP-specific: it would be equally true behind a CLI or
a queue consumer, and putting it there meant exercising any of it cost an
ASGI round-trip plus monkeypatching module globals.

It now lives in services.py as SearchService and PatentService, which take
their stateful collaborators (HTTP client, browser provider, circuit
breaker, OPS credentials) as constructor arguments. app.py keeps the
module-level instances and wires them in through `Depends`, so
`app.dependency_overrides` replaces them in a test.

The services raise domain errors - PatentNotFound, UpstreamUnavailable,
OPSUnconfigured - rather than HTTPException. Mapping those onto status
codes is the HTTP layer's job and is now three exception handlers instead
of the same branching repeated across two endpoints. app.py no longer
imports HTTPException at all.

What this buys, concretely:

- The autouse fixture that reset a process-wide circuit breaker between
tests is gone; each test builds a service with its own breaker.
- The pokes at ops token_manager._key / ._secret from two different test
files are gone; tests pass a small stand-in with a `configured` flag.
- 40 tests that needed an ASGI client are now plain unit tests. The
SearchService suite runs 22 tests in 0.21s.
- test_app.py is now about what app.py does: endpoint-to-service dispatch,
validation, dependency wiring, and the lifespan.

Also closes the coverage the refactor would otherwise have lost.
Splitting the scrapers left the four-line composition in each `query_*`
untested, so tests/test_serp_navigation.py covers it with a recording
stand-in browser - asserting each scraper navigates to the URL its own
builder produced. That is precisely the link the deleted goto-patching
harness could never test, since it accepted the URL and discarded it.

The full mutation battery from the review now stands at 18 of 18 killed,
against 6 of 13 when the review was written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MSNSYnceqvdVz7Csis4K9e

.coverage ADDED
Binary file (53.2 kB). View file
 
CONTRIBUTING.md CHANGED
@@ -30,7 +30,7 @@ ruff check . # same lint CI runs
30
  pytest
31
  ```
32
 
33
- 199 tests, a few seconds, fully offline. All outbound HTTP is mocked with
34
  `respx`; the Playwright-driven scrapers (Bing, Brave, Google Scholar,
35
  Google Patents search) run against a real headless Chromium but navigate to
36
  local fixture HTML instead of the live sites — see `tests/helpers.py` for
@@ -55,6 +55,34 @@ bug here — fix it the same way as anything else (see below): update the
55
  fixture HTML to match the new real markup, confirm the test fails against
56
  the current selectors, fix the selectors, confirm green.
57
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
  ## The development loop (TDD, not test-after)
59
 
60
  Every change in this codebase's history followed the same loop, and new
 
30
  pytest
31
  ```
32
 
33
+ 238 tests, a few seconds, fully offline. All outbound HTTP is mocked with
34
  `respx`; the Playwright-driven scrapers (Bing, Brave, Google Scholar,
35
  Google Patents search) run against a real headless Chromium but navigate to
36
  local fixture HTML instead of the live sites — see `tests/helpers.py` for
 
55
  fixture HTML to match the new real markup, confirm the test fails against
56
  the current selectors, fix the selectors, confirm green.
57
 
58
+ ## How the code is laid out
59
+
60
+ - **`app.py`** is HTTP wiring only: request validation, dependency
61
+ providers, the lifespan, and mapping domain errors onto status codes.
62
+ - **`services.py`** holds the orchestration policy - which backend to try,
63
+ in what order, when to fall back, what counts as evidence a backend is
64
+ unhealthy, and when a patent is genuinely absent rather than merely
65
+ unreachable. None of that is HTTP-specific, and keeping it out of route
66
+ handlers is why most of its tests need no ASGI client.
67
+ - **`serp.py` / `scrap.py` / `ops.py`** are the backends. Each scraper is
68
+ split into a pure URL builder, a navigation step, and an `_extract_*`
69
+ function taking an already-loaded page, so all three can be tested
70
+ separately.
71
+ - **`utils.py` / `circuit_breaker.py`** are dependency-free helpers.
72
+
73
+ The services take their stateful collaborators (HTTP client, browser,
74
+ circuit breaker, OPS credentials) as constructor arguments. That is
75
+ deliberate and worth preserving: it is what lets a test build a service
76
+ with its own circuit breaker instead of resetting a process-wide singleton
77
+ between tests, and hand in a stand-in for the OPS credentials instead of
78
+ assigning to private attributes on the real token manager. `app.py`'s
79
+ `get_search_service` / `get_patent_service` providers exist so
80
+ `app.dependency_overrides` can replace them wholesale in a test.
81
+
82
+ The stateless backend functions stay module-level imports in
83
+ `services.py`; monkeypatching one of those in a test is fine, because they
84
+ hold no state to leak between tests.
85
+
86
  ## The development loop (TDD, not test-after)
87
 
88
  Every change in this codebase's history followed the same loop, and new
app.py CHANGED
@@ -1,24 +1,24 @@
1
- import asyncio
2
  import logging
3
  import os
4
  import secrets
5
  from contextlib import asynccontextmanager
6
  from typing import Annotated, Optional
7
- from fastapi import FastAPI, HTTPException, Request
 
 
 
8
  from fastapi.responses import JSONResponse
9
  from fastapi.routing import APIRouter
10
- import httpx
11
- from httpx import HTTPStatusError
12
  from pydantic import BaseModel, Field, StringConstraints
13
- from playwright.async_api import async_playwright, Browser
14
- import uvicorn
15
 
16
- from circuit_breaker import CircuitBreaker, CircuitOpenError
17
- from scrap import PatentScrapBulkResponse, PatentScrapResult, scrap_patent_async, scrap_patent_bulk_async
18
- from serp import EmptyResultsError, PATENT_ID_CORE, SerpQuery, SerpResults, query_arxiv, query_bing_search, query_brave_search, query_ddg_search, query_google_patents, query_google_scholar
19
- from ops import OPSBulkResponse, OPSNotConfigured, ops_scrap_patent, ops_scrap_patent_bulk, ops_search, token_manager as ops_token_manager
20
- from utils import log_gathered_exceptions
21
  from mcp_server import mount_mcp_server
 
 
 
 
 
22
 
23
  # Anchored version of serp.py's PATENT_ID_CORE: validates a whole
24
  # user-supplied patent id rather than finding one inside free text. Rejects
@@ -26,6 +26,11 @@ from mcp_server import mount_mcp_server
26
  # instead of it reaching Google Patents/OPS as a confusing request.
27
  PatentId = Annotated[str, StringConstraints(pattern=rf"^{PATENT_ID_CORE}$")]
28
 
 
 
 
 
 
29
  logging.basicConfig(
30
  level=logging.INFO,
31
  format='[%(asctime)s][%(levelname)s][%(filename)s:%(lineno)d]: %(message)s',
@@ -36,12 +41,19 @@ logging.basicConfig(
36
  playwright = None
37
  pw_browser: Optional[Browser] = None
38
 
39
- # httpx client. Constructed at import so module-level handlers can close
40
- # over it, but its lifetime is owned by `api_lifespan`, which closes it on
41
- # shutdown alongside the browser.
42
  httpx_client = httpx.AsyncClient(timeout=30, limits=httpx.Limits(
43
  max_connections=30, max_keepalive_connections=20))
44
 
 
 
 
 
 
 
 
45
  # ===================== Optional API key protection =====================
46
  # Unset by default, which keeps the API exactly as open as before. Set
47
  # SERPENT_API_KEY on the deployment to require every request (REST and MCP
@@ -97,6 +109,32 @@ app = FastAPI(lifespan=api_lifespan, docs_url="/",
97
  title="SERPent", description=_load_docs())
98
 
99
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
100
  @app.middleware("http")
101
  async def api_key_guard(request: Request, call_next):
102
  """No-op unless SERPENT_API_KEY is set; then gates every path but the docs."""
@@ -114,261 +152,101 @@ async def api_key_guard(request: Request, call_next):
114
  )
115
  return await call_next(request)
116
 
117
- # Router for scrapping related endpoints
118
- scrap_router = APIRouter(prefix="/scrap", tags=["scrapping"])
119
- # Router for SERP-scrapping related endpoints
120
- serp_router = APIRouter(prefix="/serp", tags=["serp scrapping"])
121
- # Router for EPO OPS (official patent API) endpoints
122
- ops_router = APIRouter(prefix="/ops", tags=["EPO OPS"])
123
 
124
- # ===================== Search endpoints =====================
 
 
 
 
125
 
126
 
127
- def _shape_serp_results(results: list) -> SerpResults:
128
- """Flatten a list of per-query result-lists (or Exceptions, from a
129
- `return_exceptions=True` gather) into one SerpResults, surfacing the
130
- last error only when every query failed.
131
- """
132
- if not results:
133
- # SerpQuery bounds the query list, so this is unreachable from the
134
- # endpoints; keep the helper total anyway rather than indexing into
135
- # an empty list, since every search endpoint shares it.
136
- return SerpResults(results=[], error="No queries were provided.")
137
 
138
- filtered_results = [r for r in results if not isinstance(r, Exception)]
139
- flattened_results = [
140
- item for sublist in filtered_results for item in sublist]
141
 
142
- if len(filtered_results) == 0:
143
- return SerpResults(results=[], error=str(results[-1]))
 
144
 
145
- return SerpResults(results=flattened_results, error=None)
146
 
 
 
 
147
 
148
- async def _run_serp_queries(fn, params: SerpQuery, context: str) -> SerpResults:
149
- """Run `fn(query, n_results)` concurrently for every query in `params`,
150
- log any failures, and shape the result the way every simple /serp and
151
- /ops search endpoint below expects.
152
- """
153
- results = await asyncio.gather(
154
- *[fn(q, params.n_results) for q in params.queries], return_exceptions=True)
155
- log_gathered_exceptions(results, context, params.queries)
156
- return _shape_serp_results(results)
157
 
158
 
159
  @serp_router.post("/search_scholar")
160
- async def search_google_scholar(params: SerpQuery) -> SerpResults:
161
  """Queries google scholar for the specified query"""
162
  logging.info(f"Searching Google Scholar for queries: {params.queries}")
163
- return await _run_serp_queries(
164
- lambda q, n: query_google_scholar(pw_browser, q, n), params, "google scholar search")
165
 
166
 
167
  @serp_router.post("/search_arxiv")
168
- async def search_arxiv(params: SerpQuery) -> SerpResults:
169
  """Searches arxiv for the specified queries and returns the found documents."""
170
  logging.info(f"Searching Arxiv for queries: {params.queries}")
171
- return await _run_serp_queries(
172
- lambda q, n: query_arxiv(httpx_client, q, n), params, "arxiv search")
173
 
174
 
175
  @serp_router.post("/search_patents")
176
- async def search_patents(params: SerpQuery) -> SerpResults:
177
  """Searches google patents for the specified queries and returns the found documents.
178
 
179
  Falls back to the EPO OPS API for any query Google Patents returns nothing
180
  for, when OPS credentials are configured.
181
  """
182
  logging.info(f"Searching Google Patents for queries: {params.queries}")
183
- results = await asyncio.gather(*[query_google_patents(pw_browser, q, params.n_results) for q in params.queries], return_exceptions=True)
184
- log_gathered_exceptions(results, "google patent search", params.queries)
185
-
186
- # Fall back to OPS for queries that errored or returned no results.
187
- # Gathered rather than awaited one at a time: this path exists for when
188
- # Google Patents is unavailable, so it is exactly when a request is
189
- # least able to afford N sequential round-trips to the slower backend.
190
- if ops_token_manager.configured:
191
- needs_fallback = [
192
- i for i, res in enumerate(results)
193
- if isinstance(res, Exception) or not res]
194
- if needs_fallback:
195
- logging.info(
196
- f"Google Patents empty for {len(needs_fallback)} quer(y/ies), trying OPS.")
197
- fallback_results = await asyncio.gather(
198
- *[ops_search(httpx_client, params.queries[i], params.n_results)
199
- for i in needs_fallback],
200
- return_exceptions=True)
201
- # strict: gather returns exactly one result per index, so a
202
- # length mismatch here would be a bug worth surfacing.
203
- for i, res in zip(needs_fallback, fallback_results, strict=True):
204
- if isinstance(res, Exception):
205
- logging.warning(f"OPS fallback failed for `{params.queries[i]}`: {res}")
206
- else:
207
- results[i] = res
208
-
209
- return _shape_serp_results(results)
210
 
211
 
212
  @serp_router.post("/search_brave")
213
- async def search_brave(params: SerpQuery) -> SerpResults:
214
  """Searches brave search for the specified queries and returns the found documents."""
215
  logging.info(f"Searching Brave Search for queries: {params.queries}")
216
- return await _run_serp_queries(
217
- lambda q, n: query_brave_search(pw_browser, q, n), params, "brave search")
218
 
219
 
220
  @serp_router.post("/search_bing")
221
- async def search_bing(params: SerpQuery) -> SerpResults:
222
  """Searches Bing search for the specified queries and returns the found documents."""
223
  logging.info(f"Searching Bing Search for queries: {params.queries}")
224
- return await _run_serp_queries(
225
- lambda q, n: query_bing_search(pw_browser, q, n), params, "bing search")
226
 
227
 
228
  @serp_router.post("/search_duck")
229
- async def search_duck(params: SerpQuery) -> SerpResults:
230
  """Searches duckduckgo for the specified queries and returns the found documents"""
231
  logging.info(f"Searching DuckDuckGo for queries: {params.queries}")
232
- return await _run_serp_queries(
233
- lambda q, n: query_ddg_search(q, n), params, "duckduckgo search")
234
-
235
-
236
- # Shared across all queries and requests: once a backend has failed
237
- # `failure_threshold` times in a row, skip it for `cooldown_seconds` instead
238
- # of attempting (and paying the timeout cost of) another call that's very
239
- # likely to fail - and, more importantly, stop hammering a backend that may
240
- # already be rate-limiting or blocking this deployment's IP.
241
- _backend_circuit_breaker = CircuitBreaker(failure_threshold=3, cooldown_seconds=60.0)
242
-
243
-
244
- async def _search_one(q: str, n_results: int) -> tuple[str, list[dict], Optional[str]]:
245
- """Try DDG, then Brave, then Bing for a single query; stop at the first success."""
246
- backends = [
247
- ("DuckDuckGo", lambda: query_ddg_search(q, n_results)),
248
- ("Brave Search", lambda: query_brave_search(pw_browser, q, n_results)),
249
- ("Bing", lambda: query_bing_search(pw_browser, q, n_results)),
250
- ]
251
- last_error: Optional[Exception] = None
252
- for name, call in backends:
253
- try:
254
- _backend_circuit_breaker.before_call(name)
255
- except CircuitOpenError as e:
256
- logging.info(f"Skipping {name} for query `{q}`: {e}")
257
- last_error = e
258
- continue
259
-
260
- try:
261
- logging.info(f"Querying {name} with query: `{q}`")
262
- result = await call()
263
- _backend_circuit_breaker.record_success(name)
264
- return q, result, None
265
- except EmptyResultsError as e:
266
- # Zero results is ambiguous - no hits, or a soft block that
267
- # parsed to nothing. Move on to the next backend, but don't
268
- # hold it against this one: the breaker is process-wide, so
269
- # counting it would let a few unusual-but-legitimate queries
270
- # disable the primary backend for every user. A hard block
271
- # normally raises a real error, which is handled below.
272
- logging.info(f"{name} returned no results for query `{q}`")
273
- last_error = e
274
- except Exception as e:
275
- logging.error(f"Failed to query {name} with query `{q}`: {e}")
276
- _backend_circuit_breaker.record_failure(name)
277
- last_error = e
278
-
279
- return q, [], f"All backends failed for query '{q}': {last_error}"
280
 
281
 
282
  @serp_router.post("/search")
283
- async def search(params: SerpQuery) -> SerpResults:
284
  """Attempts to search the specified queries using ALL backends"""
285
- outcomes = await asyncio.gather(*[_search_one(q, params.n_results) for q in params.queries])
286
-
287
- results: list[dict] = []
288
- errors: list[str] = []
289
- for _q, res, err in outcomes:
290
- results.extend(res)
291
- if err:
292
- errors.append(err)
293
-
294
- if len(results) == 0:
295
- return SerpResults(results=[], error="; ".join(errors) if errors else "All backends are rate-limited.")
296
-
297
- return SerpResults(results=results, error="; ".join(errors) if errors else None)
298
 
299
  # =========================== Scrapping endpoints ===========================
300
 
301
 
302
- def _is_not_found(exc: Exception) -> bool:
303
- """True only when a backend positively said the document is absent.
304
-
305
- Anything else - a 5xx, a rate-limit, a timeout, an unparseable page -
306
- is a failure to answer, not an answer.
307
- """
308
- return isinstance(exc, HTTPStatusError) and exc.response.status_code == 404
309
-
310
-
311
- def _upstream_status(exc: Exception) -> int:
312
- """The status to report for a failed upstream call. Never 404.
313
-
314
- 404 is a claim that the patent does not exist, and MCP_INSTRUCTIONS
315
- tells agents to believe it and move on without retrying - so it has to
316
- be reserved for a backend actually saying so. A transient outage
317
- reported as 404 teaches an agent that a real patent is absent.
318
- """
319
- return 504 if isinstance(exc, httpx.TimeoutException) else 502
320
-
321
-
322
  @scrap_router.get("/scrap_patent/{patent_id}")
323
- async def scrap_patent(patent_id: PatentId) -> PatentScrapResult:
324
  """Scraps the specified patent from Google Patents.
325
 
326
  Falls back to the EPO OPS API (which covers patents missing from Google
327
  Patents) when the scrape fails and OPS credentials are configured.
328
  """
329
- google_error: Exception
330
- try:
331
- return await scrap_patent_async(httpx_client, f"https://patents.google.com/patent/{patent_id}/en")
332
- except HTTPStatusError as e:
333
- google_error = e
334
- logging.warning(
335
- f"Google Patents returned {e.response.status_code} for {patent_id}.")
336
- except Exception as e:
337
- google_error = e
338
- logging.warning(f"Failed to scrap patent {patent_id}: {e}")
339
-
340
- if not ops_token_manager.configured:
341
- if _is_not_found(google_error):
342
- raise HTTPException(
343
- status_code=404,
344
- detail=f"Patent '{patent_id}' not found on Google Patents (EPO OPS fallback not configured).")
345
- raise HTTPException(
346
- status_code=_upstream_status(google_error),
347
- detail=f"Google Patents request failed for '{patent_id}' and the EPO OPS fallback is not configured: {google_error}")
348
-
349
- try:
350
- logging.info(f"Trying OPS for patent {patent_id}.")
351
- return await ops_scrap_patent(httpx_client, patent_id)
352
- except Exception as e:
353
- logging.warning(f"OPS fallback failed for {patent_id}: {e}")
354
- # Only claim the patent doesn't exist when *both* backends said so.
355
- if _is_not_found(google_error) and _is_not_found(e):
356
- raise HTTPException(
357
- status_code=404,
358
- detail=f"Patent '{patent_id}' not found on Google Patents or EPO OPS.") from e
359
- if isinstance(e, HTTPStatusError):
360
- raise HTTPException(
361
- status_code=_upstream_status(e),
362
- detail=f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
363
- raise HTTPException(
364
- status_code=_upstream_status(e),
365
- detail=f"EPO OPS request failed for '{patent_id}': {e}") from e
366
-
367
-
368
- # Same reasoning as MAX_QUERIES_PER_REQUEST in serp.py: one accepted
369
- # request must not be able to queue unbounded outbound work. The OPS
370
- # variant is the expensive one - three requests per id.
371
- MAX_BULK_PATENT_IDS = 50
372
 
373
 
374
  class ScrapPatentsRequest(BaseModel):
@@ -380,51 +258,32 @@ class ScrapPatentsRequest(BaseModel):
380
 
381
 
382
  @scrap_router.post("/scrap_patents_bulk", response_model=PatentScrapBulkResponse)
383
- async def scrap_patents(params: ScrapPatentsRequest) -> PatentScrapBulkResponse:
 
384
  """Scraps multiple patents from Google Patents."""
385
- patents = await scrap_patent_bulk_async(httpx_client, params.patent_ids)
386
- return patents
387
 
388
  # =========================== EPO OPS endpoints ===========================
389
 
390
 
391
  @ops_router.post("/search")
392
- async def ops_keyword_search(params: SerpQuery) -> SerpResults:
393
  """Keyword-searches patents via the official EPO OPS API."""
394
  logging.info(f"Searching EPO OPS for queries: {params.queries}")
395
- return await _run_serp_queries(
396
- lambda q, n: ops_search(httpx_client, q, n), params, "OPS search")
397
 
398
 
399
  @ops_router.get("/scrap_patent/{patent_id}")
400
- async def ops_get_patent(patent_id: PatentId) -> PatentScrapResult:
401
  """Retrieves a patent (biblio + abstract + claims + description) via EPO OPS."""
402
- if not ops_token_manager.configured:
403
- raise HTTPException(
404
- status_code=503,
405
- detail="EPO OPS is not configured (OPS_CONSUMER_KEY / OPS_CONSUMER_SECRET missing).")
406
- try:
407
- return await ops_scrap_patent(httpx_client, patent_id)
408
- except OPSNotConfigured as e:
409
- raise HTTPException(status_code=503, detail="EPO OPS is not configured.") from e
410
- except HTTPStatusError as e:
411
- if e.response.status_code == 404:
412
- raise HTTPException(
413
- status_code=404, detail=f"Patent '{patent_id}' not found in EPO OPS.") from e
414
- raise HTTPException(
415
- status_code=_upstream_status(e),
416
- detail=f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
417
- except Exception as e:
418
- logging.warning(f"Failed to retrieve patent {patent_id} from OPS: {e}")
419
- raise HTTPException(
420
- status_code=_upstream_status(e),
421
- detail=f"EPO OPS request failed for '{patent_id}': {e}") from e
422
 
423
 
424
  @ops_router.post("/scrap_patents_bulk", response_model=OPSBulkResponse)
425
- async def ops_get_patents_bulk(params: ScrapPatentsRequest) -> OPSBulkResponse:
 
426
  """Retrieves multiple patents via EPO OPS."""
427
- return await ops_scrap_patent_bulk(httpx_client, params.patent_ids)
428
 
429
  # ===============================================================================
430
 
 
 
1
  import logging
2
  import os
3
  import secrets
4
  from contextlib import asynccontextmanager
5
  from typing import Annotated, Optional
6
+
7
+ import httpx
8
+ import uvicorn
9
+ from fastapi import Depends, FastAPI, Request
10
  from fastapi.responses import JSONResponse
11
  from fastapi.routing import APIRouter
12
+ from playwright.async_api import Browser, async_playwright
 
13
  from pydantic import BaseModel, Field, StringConstraints
 
 
14
 
15
+ from circuit_breaker import CircuitBreaker
 
 
 
 
16
  from mcp_server import mount_mcp_server
17
+ from ops import OPSBulkResponse, token_manager as ops_token_manager
18
+ from scrap import PatentScrapBulkResponse, PatentScrapResult
19
+ from serp import PATENT_ID_CORE, SerpQuery, SerpResults
20
+ from services import (OPSUnconfigured, PatentNotFound, PatentService,
21
+ SearchService, UpstreamUnavailable)
22
 
23
  # Anchored version of serp.py's PATENT_ID_CORE: validates a whole
24
  # user-supplied patent id rather than finding one inside free text. Rejects
 
26
  # instead of it reaching Google Patents/OPS as a confusing request.
27
  PatentId = Annotated[str, StringConstraints(pattern=rf"^{PATENT_ID_CORE}$")]
28
 
29
+ # Same reasoning as MAX_QUERIES_PER_REQUEST in serp.py: one accepted
30
+ # request must not be able to queue unbounded outbound work. The OPS
31
+ # variant is the expensive one - three requests per id.
32
+ MAX_BULK_PATENT_IDS = 50
33
+
34
  logging.basicConfig(
35
  level=logging.INFO,
36
  format='[%(asctime)s][%(levelname)s][%(filename)s:%(lineno)d]: %(message)s',
 
41
  playwright = None
42
  pw_browser: Optional[Browser] = None
43
 
44
+ # httpx client. Constructed at import so the dependency providers below can
45
+ # close over it, but its lifetime is owned by `api_lifespan`, which closes
46
+ # it on shutdown alongside the browser.
47
  httpx_client = httpx.AsyncClient(timeout=30, limits=httpx.Limits(
48
  max_connections=30, max_keepalive_connections=20))
49
 
50
+ # Shared across all queries and requests: once a backend has failed
51
+ # `failure_threshold` times in a row, skip it for `cooldown_seconds` instead
52
+ # of attempting (and paying the timeout cost of) another call that's very
53
+ # likely to fail - and, more importantly, stop hammering a backend that may
54
+ # already be rate-limiting or blocking this deployment's IP.
55
+ _backend_circuit_breaker = CircuitBreaker(failure_threshold=3, cooldown_seconds=60.0)
56
+
57
  # ===================== Optional API key protection =====================
58
  # Unset by default, which keeps the API exactly as open as before. Set
59
  # SERPENT_API_KEY on the deployment to require every request (REST and MCP
 
109
  title="SERPent", description=_load_docs())
110
 
111
 
112
+ # ============================ service dependencies ============================
113
+ # The orchestration itself lives in services.py; these just wire the app's
114
+ # collaborators into it. Overriding one of these via
115
+ # `app.dependency_overrides` is how a test supplies its own circuit breaker
116
+ # or OPS credentials, rather than reassigning module globals.
117
+
118
+
119
+ def get_search_service() -> SearchService:
120
+ return SearchService(
121
+ http_client=httpx_client,
122
+ # A provider, not the browser itself: the lifespan starts it after
123
+ # import, and leaves it None when Playwright fails to start.
124
+ browser_provider=lambda: pw_browser,
125
+ circuit_breaker=_backend_circuit_breaker,
126
+ ops_tokens=ops_token_manager,
127
+ )
128
+
129
+
130
+ def get_patent_service() -> PatentService:
131
+ return PatentService(http_client=httpx_client, ops_tokens=ops_token_manager)
132
+
133
+
134
+ SearchServiceDep = Annotated[SearchService, Depends(get_search_service)]
135
+ PatentServiceDep = Annotated[PatentService, Depends(get_patent_service)]
136
+
137
+
138
  @app.middleware("http")
139
  async def api_key_guard(request: Request, call_next):
140
  """No-op unless SERPENT_API_KEY is set; then gates every path but the docs."""
 
152
  )
153
  return await call_next(request)
154
 
 
 
 
 
 
 
155
 
156
+ # ========================= domain errors -> status codes =========================
157
+ # 404 is a claim that the patent does not exist, and MCP_INSTRUCTIONS tells
158
+ # agents to believe it and move on without retrying - so the services raise
159
+ # PatentNotFound only when a backend positively said so, and everything else
160
+ # arrives here as UpstreamUnavailable.
161
 
162
 
163
+ @app.exception_handler(PatentNotFound)
164
+ async def _patent_not_found_handler(request: Request, exc: PatentNotFound):
165
+ return JSONResponse({"detail": str(exc)}, status_code=404)
 
 
 
 
 
 
 
166
 
 
 
 
167
 
168
+ @app.exception_handler(UpstreamUnavailable)
169
+ async def _upstream_unavailable_handler(request: Request, exc: UpstreamUnavailable):
170
+ return JSONResponse({"detail": str(exc)}, status_code=504 if exc.timeout else 502)
171
 
 
172
 
173
+ @app.exception_handler(OPSUnconfigured)
174
+ async def _ops_unconfigured_handler(request: Request, exc: OPSUnconfigured):
175
+ return JSONResponse({"detail": str(exc)}, status_code=503)
176
 
177
+
178
+ # Router for scrapping related endpoints
179
+ scrap_router = APIRouter(prefix="/scrap", tags=["scrapping"])
180
+ # Router for SERP-scrapping related endpoints
181
+ serp_router = APIRouter(prefix="/serp", tags=["serp scrapping"])
182
+ # Router for EPO OPS (official patent API) endpoints
183
+ ops_router = APIRouter(prefix="/ops", tags=["EPO OPS"])
184
+
185
+ # ===================== Search endpoints =====================
186
 
187
 
188
  @serp_router.post("/search_scholar")
189
+ async def search_google_scholar(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
190
  """Queries google scholar for the specified query"""
191
  logging.info(f"Searching Google Scholar for queries: {params.queries}")
192
+ return await service.google_scholar(params)
 
193
 
194
 
195
  @serp_router.post("/search_arxiv")
196
+ async def search_arxiv(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
197
  """Searches arxiv for the specified queries and returns the found documents."""
198
  logging.info(f"Searching Arxiv for queries: {params.queries}")
199
+ return await service.arxiv(params)
 
200
 
201
 
202
  @serp_router.post("/search_patents")
203
+ async def search_patents(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
204
  """Searches google patents for the specified queries and returns the found documents.
205
 
206
  Falls back to the EPO OPS API for any query Google Patents returns nothing
207
  for, when OPS credentials are configured.
208
  """
209
  logging.info(f"Searching Google Patents for queries: {params.queries}")
210
+ return await service.patents(params)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
211
 
212
 
213
  @serp_router.post("/search_brave")
214
+ async def search_brave(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
215
  """Searches brave search for the specified queries and returns the found documents."""
216
  logging.info(f"Searching Brave Search for queries: {params.queries}")
217
+ return await service.brave(params)
 
218
 
219
 
220
  @serp_router.post("/search_bing")
221
+ async def search_bing(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
222
  """Searches Bing search for the specified queries and returns the found documents."""
223
  logging.info(f"Searching Bing Search for queries: {params.queries}")
224
+ return await service.bing(params)
 
225
 
226
 
227
  @serp_router.post("/search_duck")
228
+ async def search_duck(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
229
  """Searches duckduckgo for the specified queries and returns the found documents"""
230
  logging.info(f"Searching DuckDuckGo for queries: {params.queries}")
231
+ return await service.duckduckgo(params)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
232
 
233
 
234
  @serp_router.post("/search")
235
+ async def search(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
236
  """Attempts to search the specified queries using ALL backends"""
237
+ return await service.search(params)
 
 
 
 
 
 
 
 
 
 
 
 
238
 
239
  # =========================== Scrapping endpoints ===========================
240
 
241
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
242
  @scrap_router.get("/scrap_patent/{patent_id}")
243
+ async def scrap_patent(patent_id: PatentId, service: PatentServiceDep) -> PatentScrapResult:
244
  """Scraps the specified patent from Google Patents.
245
 
246
  Falls back to the EPO OPS API (which covers patents missing from Google
247
  Patents) when the scrape fails and OPS credentials are configured.
248
  """
249
+ return await service.scrap(patent_id)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
250
 
251
 
252
  class ScrapPatentsRequest(BaseModel):
 
258
 
259
 
260
  @scrap_router.post("/scrap_patents_bulk", response_model=PatentScrapBulkResponse)
261
+ async def scrap_patents(params: ScrapPatentsRequest,
262
+ service: PatentServiceDep) -> PatentScrapBulkResponse:
263
  """Scraps multiple patents from Google Patents."""
264
+ return await service.scrap_bulk(params.patent_ids)
 
265
 
266
  # =========================== EPO OPS endpoints ===========================
267
 
268
 
269
  @ops_router.post("/search")
270
+ async def ops_keyword_search(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
271
  """Keyword-searches patents via the official EPO OPS API."""
272
  logging.info(f"Searching EPO OPS for queries: {params.queries}")
273
+ return await service.ops_keyword_search(params)
 
274
 
275
 
276
  @ops_router.get("/scrap_patent/{patent_id}")
277
+ async def ops_get_patent(patent_id: PatentId, service: PatentServiceDep) -> PatentScrapResult:
278
  """Retrieves a patent (biblio + abstract + claims + description) via EPO OPS."""
279
+ return await service.ops_scrap(patent_id)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
280
 
281
 
282
  @ops_router.post("/scrap_patents_bulk", response_model=OPSBulkResponse)
283
+ async def ops_get_patents_bulk(params: ScrapPatentsRequest,
284
+ service: PatentServiceDep) -> OPSBulkResponse:
285
  """Retrieves multiple patents via EPO OPS."""
286
+ return await service.ops_scrap_bulk(params.patent_ids)
287
 
288
  # ===============================================================================
289
 
services.py ADDED
@@ -0,0 +1,337 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Backend orchestration, separated from HTTP wiring.
2
+
3
+ The policy in here - which backend to try, in what order, when to fall back
4
+ to another, what counts as evidence a backend is unhealthy, and when a
5
+ patent is genuinely absent rather than merely unreachable - is domain
6
+ policy. It would be equally true behind a CLI or a queue consumer, and it
7
+ was previously written inline in FastAPI route handlers, which meant
8
+ exercising any of it cost an ASGI round-trip.
9
+
10
+ The stateful collaborators (HTTP client, browser, circuit breaker, OPS
11
+ credentials) are constructor arguments rather than module globals reached
12
+ into at call time. That is what lets a test build a service with its own
13
+ circuit breaker instead of resetting a process-wide singleton between
14
+ tests, and hand in a stand-in for the OPS credentials instead of assigning
15
+ to private attributes on the real token manager.
16
+
17
+ The stateless backend functions (`query_arxiv` and friends) stay
18
+ module-level imports: they hold no state, so a test that wants to replace
19
+ one can monkeypatch it here without any of the problems shared mutable
20
+ state causes.
21
+
22
+ Services raise domain errors - `PatentNotFound`, `UpstreamUnavailable` -
23
+ rather than `HTTPException`. Mapping those onto status codes is the HTTP
24
+ layer's job, and lives in app.py.
25
+ """
26
+
27
+ import asyncio
28
+ import logging
29
+ from typing import Callable, Optional
30
+
31
+ import httpx
32
+ from httpx import AsyncClient, HTTPStatusError
33
+ from playwright.async_api import Browser
34
+
35
+ from circuit_breaker import CircuitBreaker, CircuitOpenError
36
+ from ops import (OPSBulkResponse, ops_scrap_patent, ops_scrap_patent_bulk,
37
+ ops_search)
38
+ from scrap import (PatentScrapBulkResponse, PatentScrapResult,
39
+ scrap_patent_async, scrap_patent_bulk_async)
40
+ from serp import (EmptyResultsError, SerpQuery, SerpResults, query_arxiv,
41
+ query_bing_search, query_brave_search, query_ddg_search,
42
+ query_google_patents, query_google_scholar)
43
+ from utils import log_gathered_exceptions
44
+
45
+ logger = logging.getLogger(__name__)
46
+
47
+
48
+ # ================================ domain errors ================================
49
+
50
+
51
+ class PatentNotFound(Exception):
52
+ """Every backend consulted positively said the document is absent.
53
+
54
+ Distinct from UpstreamUnavailable on purpose: MCP_INSTRUCTIONS tells an
55
+ agent that a not-found patent is genuinely absent everywhere and to
56
+ move on without retrying, so this must never stand in for "we couldn't
57
+ reach anyone".
58
+ """
59
+
60
+
61
+ class UpstreamUnavailable(Exception):
62
+ """A backend failed to answer. Says nothing about whether the document
63
+ exists."""
64
+
65
+ def __init__(self, message: str, *, timeout: bool = False):
66
+ super().__init__(message)
67
+ self.timeout = timeout
68
+
69
+
70
+ class OPSUnconfigured(Exception):
71
+ """An OPS-only operation was requested without credentials configured."""
72
+
73
+
74
+ def _is_not_found(exc: BaseException) -> bool:
75
+ """True only when a backend positively said the document is absent."""
76
+ return isinstance(exc, HTTPStatusError) and exc.response.status_code == 404
77
+
78
+
79
+ def _as_upstream(exc: BaseException, message: str) -> UpstreamUnavailable:
80
+ return UpstreamUnavailable(message, timeout=isinstance(exc, httpx.TimeoutException))
81
+
82
+
83
+ # ================================ result shaping ================================
84
+
85
+
86
+ def shape_serp_results(results: list) -> SerpResults:
87
+ """Flatten a list of per-query result-lists (or Exceptions, from a
88
+ `return_exceptions=True` gather) into one SerpResults, surfacing the
89
+ last error only when every query failed.
90
+ """
91
+ if not results:
92
+ # The request models bound the query list, so this is unreachable
93
+ # from the endpoints; keep the helper total anyway rather than
94
+ # indexing into an empty list, since every search path shares it.
95
+ return SerpResults(results=[], error="No queries were provided.")
96
+
97
+ filtered_results = [r for r in results if not isinstance(r, Exception)]
98
+ flattened_results = [
99
+ item for sublist in filtered_results for item in sublist]
100
+
101
+ if len(filtered_results) == 0:
102
+ return SerpResults(results=[], error=str(results[-1]))
103
+
104
+ return SerpResults(results=flattened_results, error=None)
105
+
106
+
107
+ # ================================= search service =================================
108
+
109
+
110
+ class SearchService:
111
+ """Runs SERP queries across the available backends."""
112
+
113
+ def __init__(self, *, http_client: AsyncClient,
114
+ browser_provider: Callable[[], Optional[Browser]],
115
+ circuit_breaker: CircuitBreaker, ops_tokens):
116
+ self._http = http_client
117
+ # A provider rather than the browser itself: the browser is started
118
+ # by the app's lifespan, after this service may already exist, and
119
+ # is None when Playwright failed to start.
120
+ self._browser_provider = browser_provider
121
+ self._breaker = circuit_breaker
122
+ self._ops_tokens = ops_tokens
123
+
124
+ @property
125
+ def browser(self) -> Optional[Browser]:
126
+ return self._browser_provider()
127
+
128
+ async def _run_queries(self, fn, params: SerpQuery, context: str) -> SerpResults:
129
+ """Run `fn(query, n_results)` concurrently for every query, log any
130
+ failures, and shape the result the way every simple search path
131
+ expects.
132
+ """
133
+ results = await asyncio.gather(
134
+ *[fn(q, params.n_results) for q in params.queries], return_exceptions=True)
135
+ log_gathered_exceptions(results, context, params.queries)
136
+ return shape_serp_results(results)
137
+
138
+ # ------------------------------ single backend ------------------------------
139
+
140
+ async def google_scholar(self, params: SerpQuery) -> SerpResults:
141
+ return await self._run_queries(
142
+ lambda q, n: query_google_scholar(self.browser, q, n), params,
143
+ "google scholar search")
144
+
145
+ async def arxiv(self, params: SerpQuery) -> SerpResults:
146
+ return await self._run_queries(
147
+ lambda q, n: query_arxiv(self._http, q, n), params, "arxiv search")
148
+
149
+ async def brave(self, params: SerpQuery) -> SerpResults:
150
+ return await self._run_queries(
151
+ lambda q, n: query_brave_search(self.browser, q, n), params, "brave search")
152
+
153
+ async def bing(self, params: SerpQuery) -> SerpResults:
154
+ return await self._run_queries(
155
+ lambda q, n: query_bing_search(self.browser, q, n), params, "bing search")
156
+
157
+ async def duckduckgo(self, params: SerpQuery) -> SerpResults:
158
+ return await self._run_queries(
159
+ lambda q, n: query_ddg_search(q, n), params, "duckduckgo search")
160
+
161
+ async def ops_keyword_search(self, params: SerpQuery) -> SerpResults:
162
+ return await self._run_queries(
163
+ lambda q, n: ops_search(self._http, q, n), params, "OPS search")
164
+
165
+ # -------------------------------- patents --------------------------------
166
+
167
+ async def patents(self, params: SerpQuery) -> SerpResults:
168
+ """Google Patents, falling back to EPO OPS for any query it returns
169
+ nothing for (when OPS credentials are configured)."""
170
+ results = await asyncio.gather(
171
+ *[query_google_patents(self.browser, q, params.n_results)
172
+ for q in params.queries],
173
+ return_exceptions=True)
174
+ log_gathered_exceptions(results, "google patent search", params.queries)
175
+
176
+ # Gathered rather than awaited one at a time: this path exists for
177
+ # when Google Patents is unavailable, so it is exactly when a
178
+ # request can least afford N sequential round-trips to the slower
179
+ # backend.
180
+ if self._ops_tokens.configured:
181
+ needs_fallback = [
182
+ i for i, res in enumerate(results)
183
+ if isinstance(res, Exception) or not res]
184
+ if needs_fallback:
185
+ logger.info(
186
+ f"Google Patents empty for {len(needs_fallback)} quer(y/ies), trying OPS.")
187
+ fallback = await asyncio.gather(
188
+ *[ops_search(self._http, params.queries[i], params.n_results)
189
+ for i in needs_fallback],
190
+ return_exceptions=True)
191
+ # strict: gather returns exactly one result per index, so a
192
+ # length mismatch here would be a bug worth surfacing.
193
+ for i, res in zip(needs_fallback, fallback, strict=True):
194
+ if isinstance(res, Exception):
195
+ logger.warning(f"OPS fallback failed for `{params.queries[i]}`: {res}")
196
+ else:
197
+ results[i] = res
198
+
199
+ return shape_serp_results(results)
200
+
201
+ # ----------------------------- all backends -----------------------------
202
+
203
+ async def _search_one(self, q: str, n_results: int) -> tuple[str, list[dict], Optional[str]]:
204
+ """Try DDG, then Brave, then Bing for a single query; stop at the
205
+ first success."""
206
+ backends = [
207
+ ("DuckDuckGo", lambda: query_ddg_search(q, n_results)),
208
+ ("Brave Search", lambda: query_brave_search(self.browser, q, n_results)),
209
+ ("Bing", lambda: query_bing_search(self.browser, q, n_results)),
210
+ ]
211
+ last_error: Optional[Exception] = None
212
+ for name, call in backends:
213
+ try:
214
+ self._breaker.before_call(name)
215
+ except CircuitOpenError as e:
216
+ logger.info(f"Skipping {name} for query `{q}`: {e}")
217
+ last_error = e
218
+ continue
219
+
220
+ try:
221
+ logger.info(f"Querying {name} with query: `{q}`")
222
+ result = await call()
223
+ self._breaker.record_success(name)
224
+ return q, result, None
225
+ except EmptyResultsError as e:
226
+ # Zero results is ambiguous - no hits, or a soft block that
227
+ # parsed to nothing. Move on to the next backend, but don't
228
+ # hold it against this one: the breaker is shared across
229
+ # every request, so counting it would let a few
230
+ # unusual-but-legitimate queries disable the primary backend
231
+ # for everyone. A hard block normally raises a real error,
232
+ # which is handled below.
233
+ logger.info(f"{name} returned no results for query `{q}`")
234
+ last_error = e
235
+ except Exception as e:
236
+ logger.error(f"Failed to query {name} with query `{q}`: {e}")
237
+ self._breaker.record_failure(name)
238
+ last_error = e
239
+
240
+ return q, [], f"All backends failed for query '{q}': {last_error}"
241
+
242
+ async def search(self, params: SerpQuery) -> SerpResults:
243
+ """Search every query across the full backend fallback chain."""
244
+ outcomes = await asyncio.gather(
245
+ *[self._search_one(q, params.n_results) for q in params.queries])
246
+
247
+ results: list[dict] = []
248
+ errors: list[str] = []
249
+ for _q, res, err in outcomes:
250
+ results.extend(res)
251
+ if err:
252
+ errors.append(err)
253
+
254
+ if len(results) == 0:
255
+ return SerpResults(
256
+ results=[],
257
+ error="; ".join(errors) if errors else "All backends are rate-limited.")
258
+
259
+ return SerpResults(results=results, error="; ".join(errors) if errors else None)
260
+
261
+
262
+ # ================================= patent service =================================
263
+
264
+
265
+ class PatentService:
266
+ """Retrieves patent full text, from Google Patents or EPO OPS."""
267
+
268
+ def __init__(self, *, http_client: AsyncClient, ops_tokens):
269
+ self._http = http_client
270
+ self._ops_tokens = ops_tokens
271
+
272
+ async def scrap(self, patent_id: str) -> PatentScrapResult:
273
+ """Scrape from Google Patents, falling back to EPO OPS.
274
+
275
+ Raises PatentNotFound only when every backend consulted said the
276
+ document is absent; anything else raises UpstreamUnavailable.
277
+ """
278
+ google_error: BaseException
279
+ try:
280
+ return await scrap_patent_async(
281
+ self._http, f"https://patents.google.com/patent/{patent_id}/en")
282
+ except HTTPStatusError as e:
283
+ google_error = e
284
+ logger.warning(
285
+ f"Google Patents returned {e.response.status_code} for {patent_id}.")
286
+ except Exception as e:
287
+ google_error = e
288
+ logger.warning(f"Failed to scrap patent {patent_id}: {e}")
289
+
290
+ if not self._ops_tokens.configured:
291
+ if _is_not_found(google_error):
292
+ raise PatentNotFound(
293
+ f"Patent '{patent_id}' not found on Google Patents "
294
+ "(EPO OPS fallback not configured).")
295
+ raise _as_upstream(
296
+ google_error,
297
+ f"Google Patents request failed for '{patent_id}' and the EPO OPS "
298
+ f"fallback is not configured: {google_error}")
299
+
300
+ try:
301
+ logger.info(f"Trying OPS for patent {patent_id}.")
302
+ return await ops_scrap_patent(self._http, patent_id)
303
+ except Exception as e:
304
+ logger.warning(f"OPS fallback failed for {patent_id}: {e}")
305
+ # Only claim the patent doesn't exist when *both* backends said so.
306
+ if _is_not_found(google_error) and _is_not_found(e):
307
+ raise PatentNotFound(
308
+ f"Patent '{patent_id}' not found on Google Patents or EPO OPS.") from e
309
+ if isinstance(e, HTTPStatusError):
310
+ raise _as_upstream(
311
+ e, f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
312
+ raise _as_upstream(e, f"EPO OPS request failed for '{patent_id}': {e}") from e
313
+
314
+ async def scrap_bulk(self, patent_ids: list[str]) -> PatentScrapBulkResponse:
315
+ return await scrap_patent_bulk_async(self._http, patent_ids)
316
+
317
+ async def ops_scrap(self, patent_id: str) -> PatentScrapResult:
318
+ """Retrieve via EPO OPS only, with no Google Patents fallback."""
319
+ if not self._ops_tokens.configured:
320
+ raise OPSUnconfigured(
321
+ "EPO OPS is not configured (OPS_CONSUMER_KEY / OPS_CONSUMER_SECRET missing).")
322
+ try:
323
+ return await ops_scrap_patent(self._http, patent_id)
324
+ except HTTPStatusError as e:
325
+ if e.response.status_code == 404:
326
+ # OPS is the only backend asked here, so its 404 is the
327
+ # whole answer.
328
+ raise PatentNotFound(
329
+ f"Patent '{patent_id}' not found in EPO OPS.") from e
330
+ raise _as_upstream(
331
+ e, f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
332
+ except Exception as e:
333
+ logger.warning(f"Failed to retrieve patent {patent_id} from OPS: {e}")
334
+ raise _as_upstream(e, f"EPO OPS request failed for '{patent_id}': {e}") from e
335
+
336
+ async def ops_scrap_bulk(self, patent_ids: list[str]) -> OPSBulkResponse:
337
+ return await ops_scrap_patent_bulk(self._http, patent_ids)
tests/test_app.py CHANGED
@@ -1,12 +1,10 @@
1
- """Characterization tests for app.py's search endpoints.
2
-
3
- These monkeypatch the underlying query_*/ops_search functions rather than
4
- hitting real search engines - app.py's own job here is just wiring
5
- (gather queries concurrently, log failures, flatten results, decide when to
6
- report an error), which is exactly what's being tested. The app's lifespan
7
- (which starts a real Playwright browser) is never triggered: ASGITransport
8
- doesn't run it, and these tests don't need it since pw_browser is never
9
- actually used once the query functions are patched out.
10
  """
11
 
12
  import httpx
@@ -14,9 +12,8 @@ import pytest
14
  from httpx import ASGITransport
15
 
16
  import app as app_module
17
- from circuit_breaker import CircuitBreaker
18
- from scrap import PatentScrapResult
19
- from serp import SerpQuery
20
 
21
 
22
  @pytest.fixture
@@ -26,151 +23,134 @@ async def client():
26
  yield c
27
 
28
 
29
- @pytest.fixture(autouse=True)
30
- def _fresh_backend_circuit_breaker(monkeypatch):
31
- """`_backend_circuit_breaker` is a module-level singleton shared across
32
- every request in production, which is the point - but it means a
33
- backend failure recorded by one test would otherwise leak into the
34
- next. Give every test its own instance.
35
- """
36
- monkeypatch.setattr(app_module, "_backend_circuit_breaker", CircuitBreaker())
37
-
38
-
39
- # ------------------------- _shape_serp_results / _run_serp_queries -------------------------
40
 
 
 
41
 
42
- def test_shape_serp_results_flattens_successful_lists():
43
- result = app_module._shape_serp_results([[{"a": 1}], [{"b": 2}, {"c": 3}]])
44
- assert result.results == [{"a": 1}, {"b": 2}, {"c": 3}]
45
- assert result.error is None
46
 
 
 
47
 
48
- def test_shape_serp_results_surfaces_last_error_when_all_failed():
49
- result = app_module._shape_serp_results([ValueError("first"), RuntimeError("second")])
50
- assert result.results == []
51
- assert "second" in result.error
52
 
 
 
 
 
 
53
 
54
- def test_shape_serp_results_ignores_failures_when_some_succeed():
55
- result = app_module._shape_serp_results([[{"a": 1}], RuntimeError("boom")])
56
- assert result.results == [{"a": 1}]
57
- assert result.error is None
58
 
59
 
60
- async def test_run_serp_queries_calls_fn_per_query_and_shapes_result():
61
- params = SerpQuery(queries=["x", "y"], n_results=5)
62
- calls = []
63
 
64
- async def fn(q, n):
65
- calls.append((q, n))
66
- return [{"q": q}]
67
 
68
- result = await app_module._run_serp_queries(fn, params, "test context")
 
 
69
 
70
- assert sorted(calls) == [("x", 5), ("y", 5)]
71
- assert sorted(result.results, key=str) == sorted([{"q": "x"}, {"q": "y"}], key=str)
 
72
 
 
 
 
73
 
74
- # ------------------------------------- endpoints -------------------------------------
75
 
 
76
 
77
- async def test_search_arxiv_flattens_results_from_all_queries(client, monkeypatch):
78
- async def fake_query_arxiv(client_arg, q, n):
79
- return [{"title": f"paper about {q}", "href": "x", "body": "y", "id": "1"}]
80
 
81
- monkeypatch.setattr(app_module, "query_arxiv", fake_query_arxiv)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
 
83
- resp = await client.post("/serp/search_arxiv", json={"queries": ["a", "b"], "n_results": 5})
84
 
85
  assert resp.status_code == 200
86
- data = resp.json()
87
- assert data["error"] is None
88
- assert {r["title"] for r in data["results"]} == {"paper about a", "paper about b"}
89
-
90
 
91
- async def test_search_arxiv_all_queries_failing_returns_error(client, monkeypatch):
92
- async def failing_query_arxiv(client_arg, q, n):
93
- raise RuntimeError(f"boom for {q}")
94
 
95
- monkeypatch.setattr(app_module, "query_arxiv", failing_query_arxiv)
 
 
 
 
 
 
 
96
 
97
- resp = await client.post("/serp/search_arxiv", json={"queries": ["a"], "n_results": 5})
98
 
99
  assert resp.status_code == 200
100
- data = resp.json()
101
- assert data["results"] == []
102
- assert "boom for a" in data["error"]
103
-
104
-
105
- async def test_search_patents_does_not_fall_back_when_google_patents_succeeds(client, monkeypatch):
106
- async def fake_google_patents(browser, q, n):
107
- return [{"id": "US1", "href": "h", "title": "t", "body": "b"}]
108
-
109
- ops_search_called = False
110
-
111
- async def fake_ops_search(client_arg, q, n):
112
- nonlocal ops_search_called
113
- ops_search_called = True
114
- return []
115
-
116
- monkeypatch.setattr(app_module, "query_google_patents", fake_google_patents)
117
- monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
118
- monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
119
- monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
120
-
121
- resp = await client.post("/serp/search_patents", json={"queries": ["widget"], "n_results": 5})
122
-
123
- data = resp.json()
124
- assert data["results"] == [{"id": "US1", "href": "h", "title": "t", "body": "b"}]
125
- assert ops_search_called is False
126
-
127
 
128
- async def test_search_patents_falls_back_to_ops_when_google_patents_empty(client, monkeypatch):
129
- async def fake_google_patents(browser, q, n):
130
- return []
131
 
132
- async def fake_ops_search(client_arg, q, n):
133
- return [{"title": "OPS result", "body": "", "href": "", "id": "US123"}]
 
 
 
 
 
 
134
 
135
- monkeypatch.setattr(app_module, "query_google_patents", fake_google_patents)
136
- monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
137
- monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
138
- monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
139
 
140
- resp = await client.post("/serp/search_patents", json={"queries": ["widget"], "n_results": 5})
141
-
142
- data = resp.json()
143
- assert data["results"] == [{"title": "OPS result", "body": "", "href": "", "id": "US123"}]
144
-
145
-
146
- async def test_search_patents_no_fallback_when_ops_not_configured(client, monkeypatch):
147
- async def fake_google_patents(browser, q, n):
148
- return []
149
-
150
- monkeypatch.setattr(app_module, "query_google_patents", fake_google_patents)
151
- monkeypatch.setattr(app_module.ops_token_manager, "_key", None)
152
- monkeypatch.setattr(app_module.ops_token_manager, "_secret", None)
153
-
154
- resp = await client.post("/serp/search_patents", json={"queries": ["widget"], "n_results": 5})
155
-
156
- data = resp.json()
157
- assert data["results"] == []
158
- assert data["error"] is None
159
 
160
 
161
  # ------------------------------- patent_id validation -------------------------------
162
 
163
 
164
- async def test_scrap_patent_rejects_malformed_patent_id(client):
165
- resp = await client.get("/scrap/scrap_patent/not-a-patent-id")
 
 
 
 
166
  assert resp.status_code == 422
167
 
168
 
169
- async def test_scrap_patent_accepts_well_formed_patent_id(client, monkeypatch):
170
- async def fake_scrap(client_arg, url):
171
- return PatentScrapResult(title="Widget apparatus")
 
 
172
 
173
- monkeypatch.setattr(app_module, "scrap_patent_async", fake_scrap)
 
 
174
 
175
  resp = await client.get("/scrap/scrap_patent/US11930446B2")
176
 
@@ -178,38 +158,47 @@ async def test_scrap_patent_accepts_well_formed_patent_id(client, monkeypatch):
178
  assert resp.json()["title"] == "Widget apparatus"
179
 
180
 
181
- async def test_ops_get_patent_rejects_malformed_patent_id(client):
182
- resp = await client.get("/ops/scrap_patent/not-a-patent-id")
183
- assert resp.status_code == 422
184
 
185
 
186
- async def test_scrap_patents_bulk_rejects_malformed_id_in_list(client):
187
- resp = await client.post(
188
- "/scrap/scrap_patents_bulk", json={"patent_ids": ["US11930446B2", "garbage!!"]})
189
- assert resp.status_code == 422
190
 
 
 
 
191
 
192
- async def test_ops_scrap_patents_bulk_rejects_malformed_id_in_list(client):
193
- resp = await client.post("/ops/scrap_patents_bulk", json={"patent_ids": ["garbage!!"]})
194
- assert resp.status_code == 422
195
 
 
 
 
 
 
 
196
 
197
- async def test_ops_keyword_search_flattens_results(client, monkeypatch):
198
- async def fake_ops_search(client_arg, q, n):
199
- return [{"title": f"OPS {q}", "body": "", "href": "", "id": "1"}]
200
 
201
- monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
202
 
203
- resp = await client.post("/ops/search", json={"queries": ["a", "b"], "n_results": 5})
 
204
 
205
- data = resp.json()
206
- assert data["error"] is None
207
- assert {r["title"] for r in data["results"]} == {"OPS a", "OPS b"}
208
 
209
 
210
  # ------------------------------------- api_lifespan -------------------------------------
211
 
212
 
 
 
 
 
 
 
 
 
213
  async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
214
  """The container runs as a non-root user (see the Dockerfile), and
215
  Chromium's own internal sandbox needs privileges a non-root container
@@ -255,24 +244,10 @@ async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
255
  assert "--no-sandbox" in launch_calls[0].get("args", [])
256
 
257
 
258
- class _RecordingClient:
259
- def __init__(self):
260
- self.closed = False
261
-
262
- async def aclose(self):
263
- self.closed = True
264
-
265
-
266
  async def test_api_lifespan_closes_the_http_client_on_shutdown(monkeypatch):
267
  """The client was created at import and never closed - the browser was
268
  torn down on shutdown but its connection pool was not.
269
  """
270
- class FakePlaywright:
271
- chromium = None
272
-
273
- async def stop(self):
274
- pass
275
-
276
  recording_client = _RecordingClient()
277
  monkeypatch.setattr(app_module, "pw_browser", None)
278
  monkeypatch.setattr(app_module, "playwright", None)
@@ -301,257 +276,3 @@ async def test_api_lifespan_survives_playwright_failing_to_start(monkeypatch):
301
 
302
  async with app_module.api_lifespan(app_module.app):
303
  assert app_module.pw_browser is None
304
-
305
-
306
- # --------------------------------- _search_one fallback ---------------------------------
307
-
308
-
309
- async def test_search_one_returns_first_backend_that_succeeds(monkeypatch):
310
- calls = []
311
-
312
- async def fake_ddg(q, n):
313
- calls.append("ddg")
314
- return [{"title": "from ddg"}]
315
-
316
- async def fake_brave(browser, q, n):
317
- calls.append("brave")
318
- return [{"title": "from brave"}]
319
-
320
- monkeypatch.setattr(app_module, "query_ddg_search", fake_ddg)
321
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
322
-
323
- q, results, err = await app_module._search_one("widget", 5)
324
-
325
- assert calls == ["ddg"]
326
- assert results == [{"title": "from ddg"}]
327
- assert err is None
328
-
329
-
330
- async def test_search_one_falls_through_to_the_next_backend_on_failure(monkeypatch):
331
- calls = []
332
-
333
- async def failing_ddg(q, n):
334
- calls.append("ddg")
335
- raise RuntimeError("ddg blocked")
336
-
337
- async def fake_brave(browser, q, n):
338
- calls.append("brave")
339
- return [{"title": "from brave"}]
340
-
341
- monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
342
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
343
-
344
- q, results, err = await app_module._search_one("widget", 5)
345
-
346
- assert calls == ["ddg", "brave"]
347
- assert results == [{"title": "from brave"}]
348
- assert err is None
349
-
350
-
351
- async def test_search_one_reports_an_error_when_every_backend_fails(monkeypatch):
352
- async def failing(*args):
353
- raise RuntimeError("blocked")
354
-
355
- monkeypatch.setattr(app_module, "query_ddg_search", failing)
356
- monkeypatch.setattr(app_module, "query_brave_search", failing)
357
- monkeypatch.setattr(app_module, "query_bing_search", failing)
358
-
359
- q, results, err = await app_module._search_one("widget", 5)
360
-
361
- assert results == []
362
- assert "All backends failed for query 'widget'" in err
363
- assert "blocked" in err
364
-
365
-
366
- async def test_search_one_skips_a_backend_whose_circuit_is_open(monkeypatch):
367
- """A backend that's failed `failure_threshold` times in a row shouldn't
368
- be attempted again until its cooldown elapses - repeatedly hammering a
369
- backend that's already rate-limiting us only makes that worse.
370
- """
371
- from circuit_breaker import CircuitBreaker
372
-
373
- fresh_breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
374
- monkeypatch.setattr(app_module, "_backend_circuit_breaker", fresh_breaker)
375
-
376
- ddg_calls = []
377
-
378
- async def failing_ddg(q, n):
379
- ddg_calls.append(q)
380
- raise RuntimeError("ddg blocked")
381
-
382
- async def fake_brave(browser, q, n):
383
- return [{"title": "from brave"}]
384
-
385
- monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
386
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
387
-
388
- # First call: DDG fails, opening its circuit (threshold=1); Brave picks it up.
389
- await app_module._search_one("first query", 5)
390
- assert ddg_calls == ["first query"]
391
-
392
- # Second call: DDG's circuit is now open, so it should be skipped
393
- # entirely (not called again) and go straight to Brave.
394
- q, results, err = await app_module._search_one("second query", 5)
395
- assert ddg_calls == ["first query"] # unchanged - DDG was skipped
396
- assert results == [{"title": "from brave"}]
397
-
398
-
399
- # ------------------------------- /serp/search endpoint -------------------------------
400
- # `_search_one` is covered above; this is the aggregation wrapped around it,
401
- # which had no coverage - dropping every per-query error left the suite green.
402
-
403
-
404
- async def test_search_endpoint_merges_results_from_every_query(client, monkeypatch):
405
- async def fake_ddg(q, n):
406
- return [{"title": f"result for {q}"}]
407
-
408
- monkeypatch.setattr(app_module, "query_ddg_search", fake_ddg)
409
-
410
- resp = await client.post("/serp/search", json={"queries": ["a", "b"], "n_results": 5})
411
-
412
- data = resp.json()
413
- assert {r["title"] for r in data["results"]} == {"result for a", "result for b"}
414
- assert data["error"] is None
415
-
416
-
417
- async def test_search_endpoint_reports_partial_failure_alongside_results(client, monkeypatch):
418
- """A query that no backend could answer must be reported even when
419
- another query succeeded - otherwise the caller silently receives fewer
420
- results than they asked for with no indication why.
421
- """
422
- async def selective_ddg(q, n):
423
- if q == "bad":
424
- raise RuntimeError("blocked")
425
- return [{"title": f"result for {q}"}]
426
-
427
- async def failing(*args):
428
- raise RuntimeError("blocked")
429
-
430
- monkeypatch.setattr(app_module, "query_ddg_search", selective_ddg)
431
- monkeypatch.setattr(app_module, "query_brave_search", failing)
432
- monkeypatch.setattr(app_module, "query_bing_search", failing)
433
-
434
- resp = await client.post("/serp/search", json={"queries": ["good", "bad"], "n_results": 5})
435
-
436
- data = resp.json()
437
- assert [r["title"] for r in data["results"]] == ["result for good"]
438
- assert "bad" in data["error"]
439
-
440
-
441
- async def test_search_endpoint_reports_an_error_when_everything_fails(client, monkeypatch):
442
- async def failing(*args):
443
- raise RuntimeError("blocked")
444
-
445
- monkeypatch.setattr(app_module, "query_ddg_search", failing)
446
- monkeypatch.setattr(app_module, "query_brave_search", failing)
447
- monkeypatch.setattr(app_module, "query_bing_search", failing)
448
-
449
- resp = await client.post("/serp/search", json={"queries": ["a", "b"], "n_results": 5})
450
-
451
- data = resp.json()
452
- assert data["results"] == []
453
- # Both failed queries are named, not just the last one.
454
- assert "'a'" in data["error"] and "'b'" in data["error"]
455
-
456
-
457
- async def test_search_endpoint_falls_through_backends_per_query(client, monkeypatch):
458
- """The fallback chain is per query, so one blocked backend shouldn't
459
- cost results for queries the next backend can answer."""
460
- async def failing_ddg(q, n):
461
- raise RuntimeError("ddg blocked")
462
-
463
- async def fake_brave(browser, q, n):
464
- return [{"title": f"brave has {q}"}]
465
-
466
- monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
467
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
468
-
469
- resp = await client.post("/serp/search", json={"queries": ["a"], "n_results": 5})
470
-
471
- data = resp.json()
472
- assert data["results"] == [{"title": "brave has a"}]
473
- assert data["error"] is None
474
-
475
-
476
- # --------------------- what counts as a backend failure ---------------------
477
-
478
-
479
- async def test_a_zero_result_query_does_not_open_the_backends_circuit(monkeypatch):
480
- """An empty result is ambiguous - it can mean the query genuinely has
481
- no hits, or that the backend served an anti-bot page that parsed to
482
- nothing. The circuit breaker is a process-wide singleton, so counting
483
- the ambiguous case meant three unusual-but-legitimate queries in a row
484
- could disable the primary backend for every user of the deployment for
485
- a full cooldown.
486
- """
487
- from circuit_breaker import CircuitBreaker
488
- from serp import DuckDuckGoBlockedException
489
-
490
- breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
491
- monkeypatch.setattr(app_module, "_backend_circuit_breaker", breaker)
492
-
493
- ddg_calls = []
494
-
495
- async def empty_ddg(q, n):
496
- ddg_calls.append(q)
497
- raise DuckDuckGoBlockedException()
498
-
499
- async def fake_brave(browser, q, n):
500
- return [{"title": "from brave"}]
501
-
502
- monkeypatch.setattr(app_module, "query_ddg_search", empty_ddg)
503
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
504
-
505
- await app_module._search_one("obscure query", 5)
506
- await app_module._search_one("another obscure query", 5)
507
-
508
- # Still being tried: an empty result is not evidence the backend is down.
509
- assert ddg_calls == ["obscure query", "another obscure query"]
510
-
511
-
512
- async def test_an_empty_result_still_falls_through_to_the_next_backend(monkeypatch):
513
- """Not counting it against the backend must not stop the fallback: the
514
- caller still needs results from somewhere."""
515
- from serp import DuckDuckGoBlockedException
516
-
517
- async def empty_ddg(q, n):
518
- raise DuckDuckGoBlockedException()
519
-
520
- async def fake_brave(browser, q, n):
521
- return [{"title": "from brave"}]
522
-
523
- monkeypatch.setattr(app_module, "query_ddg_search", empty_ddg)
524
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
525
-
526
- _q, results, err = await app_module._search_one("widget", 5)
527
-
528
- assert results == [{"title": "from brave"}]
529
- assert err is None
530
-
531
-
532
- async def test_a_real_backend_error_still_opens_the_circuit(monkeypatch):
533
- """The distinction has to cut both ways: a backend that raises a
534
- genuine error (a rate limit, a transport failure) is still evidence,
535
- and must still trip the breaker.
536
- """
537
- from circuit_breaker import CircuitBreaker
538
-
539
- breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
540
- monkeypatch.setattr(app_module, "_backend_circuit_breaker", breaker)
541
-
542
- ddg_calls = []
543
-
544
- async def failing_ddg(q, n):
545
- ddg_calls.append(q)
546
- raise RuntimeError("rate limited")
547
-
548
- async def fake_brave(browser, q, n):
549
- return [{"title": "from brave"}]
550
-
551
- monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
552
- monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
553
-
554
- await app_module._search_one("first", 5)
555
- await app_module._search_one("second", 5)
556
-
557
- assert ddg_calls == ["first"] # skipped the second time
 
1
+ """app.py: HTTP wiring only.
2
+
3
+ The orchestration these endpoints used to contain now lives in
4
+ services.py and is tested in test_search_service.py and
5
+ test_patent_fallback.py without an HTTP layer. What's left here is what
6
+ app.py is actually responsible for: request validation, dependency wiring,
7
+ the lifespan, and turning domain results into responses.
 
 
8
  """
9
 
10
  import httpx
 
12
  from httpx import ASGITransport
13
 
14
  import app as app_module
15
+ from scrap import PatentScrapBulkResponse, PatentScrapResult
16
+ from serp import SerpResults
 
17
 
18
 
19
  @pytest.fixture
 
23
  yield c
24
 
25
 
26
+ @pytest.fixture
27
+ def override_services():
28
+ """Install stand-in services for the duration of one test."""
29
+ def _install(*, search=None, patent=None):
30
+ if search is not None:
31
+ app_module.app.dependency_overrides[app_module.get_search_service] = lambda: search
32
+ if patent is not None:
33
+ app_module.app.dependency_overrides[app_module.get_patent_service] = lambda: patent
 
 
 
34
 
35
+ yield _install
36
+ app_module.app.dependency_overrides.clear()
37
 
 
 
 
 
38
 
39
+ class _RecordingSearchService:
40
+ """Records which service method each endpoint dispatched to."""
41
 
42
+ def __init__(self):
43
+ self.calls = []
 
 
44
 
45
+ def _record(self, name):
46
+ async def _fn(params):
47
+ self.calls.append((name, list(params.queries), params.n_results))
48
+ return SerpResults(results=[{"title": name}], error=None)
49
+ return _fn
50
 
51
+ def __getattr__(self, name):
52
+ return self._record(name)
 
 
53
 
54
 
55
+ class _StubPatentService:
56
+ def __init__(self):
57
+ self.calls = []
58
 
59
+ async def scrap(self, patent_id):
60
+ self.calls.append(("scrap", patent_id))
61
+ return PatentScrapResult(title="Widget apparatus")
62
 
63
+ async def scrap_bulk(self, patent_ids):
64
+ self.calls.append(("scrap_bulk", patent_ids))
65
+ return PatentScrapBulkResponse(patents=[], failed_ids=list(patent_ids))
66
 
67
+ async def ops_scrap(self, patent_id):
68
+ self.calls.append(("ops_scrap", patent_id))
69
+ return PatentScrapResult(title="From OPS")
70
 
71
+ async def ops_scrap_bulk(self, patent_ids):
72
+ self.calls.append(("ops_scrap_bulk", patent_ids))
73
+ return PatentScrapBulkResponse(patents=[], failed_ids=list(patent_ids))
74
 
 
75
 
76
+ # ------------------------------- endpoint dispatch -------------------------------
77
 
 
 
 
78
 
79
+ @pytest.mark.parametrize("path, expected_method", [
80
+ ("/serp/search_scholar", "google_scholar"),
81
+ ("/serp/search_arxiv", "arxiv"),
82
+ ("/serp/search_patents", "patents"),
83
+ ("/serp/search_brave", "brave"),
84
+ ("/serp/search_bing", "bing"),
85
+ ("/serp/search_duck", "duckduckgo"),
86
+ ("/serp/search", "search"),
87
+ ("/ops/search", "ops_keyword_search"),
88
+ ])
89
+ async def test_each_search_endpoint_dispatches_to_its_service_method(
90
+ client, override_services, path, expected_method):
91
+ """Guards the wiring itself: every endpoint reaching the right method
92
+ with the request's queries and result count intact."""
93
+ service = _RecordingSearchService()
94
+ override_services(search=service)
95
 
96
+ resp = await client.post(path, json={"queries": ["a", "b"], "n_results": 25})
97
 
98
  assert resp.status_code == 200
99
+ assert service.calls == [(expected_method, ["a", "b"], 25)]
100
+ assert resp.json()["results"] == [{"title": expected_method}]
 
 
101
 
 
 
 
102
 
103
+ @pytest.mark.parametrize("method, path, expected_call", [
104
+ ("get", "/scrap/scrap_patent/US11930446B2", ("scrap", "US11930446B2")),
105
+ ("get", "/ops/scrap_patent/US11930446B2", ("ops_scrap", "US11930446B2")),
106
+ ])
107
+ async def test_patent_endpoints_dispatch_to_their_service_method(
108
+ client, override_services, method, path, expected_call):
109
+ service = _StubPatentService()
110
+ override_services(patent=service)
111
 
112
+ resp = await getattr(client, method)(path)
113
 
114
  assert resp.status_code == 200
115
+ assert service.calls == [expected_call]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
116
 
 
 
 
117
 
118
+ @pytest.mark.parametrize("path, expected_method", [
119
+ ("/scrap/scrap_patents_bulk", "scrap_bulk"),
120
+ ("/ops/scrap_patents_bulk", "ops_scrap_bulk"),
121
+ ])
122
+ async def test_bulk_endpoints_dispatch_to_their_service_method(
123
+ client, override_services, path, expected_method):
124
+ service = _StubPatentService()
125
+ override_services(patent=service)
126
 
127
+ resp = await client.post(path, json={"patent_ids": ["US11930446B2"]})
 
 
 
128
 
129
+ assert resp.status_code == 200
130
+ assert service.calls == [(expected_method, ["US11930446B2"])]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
131
 
132
 
133
  # ------------------------------- patent_id validation -------------------------------
134
 
135
 
136
+ @pytest.mark.parametrize("path", [
137
+ "/scrap/scrap_patent/not-a-patent-id",
138
+ "/ops/scrap_patent/not-a-patent-id",
139
+ ])
140
+ async def test_malformed_patent_id_is_rejected(client, path):
141
+ resp = await client.get(path)
142
  assert resp.status_code == 422
143
 
144
 
145
+ @pytest.mark.parametrize("path", [
146
+ "/scrap/scrap_patents_bulk", "/ops/scrap_patents_bulk"])
147
+ async def test_malformed_patent_id_in_a_bulk_list_is_rejected(client, path):
148
+ resp = await client.post(path, json={"patent_ids": ["US11930446B2", "garbage!!"]})
149
+ assert resp.status_code == 422
150
 
151
+
152
+ async def test_well_formed_patent_id_is_accepted(client, override_services):
153
+ override_services(patent=_StubPatentService())
154
 
155
  resp = await client.get("/scrap/scrap_patent/US11930446B2")
156
 
 
158
  assert resp.json()["title"] == "Widget apparatus"
159
 
160
 
161
+ # ---------------------------- default dependency wiring ----------------------------
 
 
162
 
163
 
164
+ def test_the_default_search_service_gets_the_apps_collaborators():
165
+ """The providers exist so tests can override them; this pins that the
166
+ un-overridden ones still hand the service the real collaborators."""
167
+ service = app_module.get_search_service()
168
 
169
+ assert service._http is app_module.httpx_client
170
+ assert service._breaker is app_module._backend_circuit_breaker
171
+ assert service._ops_tokens is app_module.ops_token_manager
172
 
 
 
 
173
 
174
+ def test_the_search_service_reads_the_browser_lazily(monkeypatch):
175
+ """The browser is started by the lifespan, after the service may already
176
+ exist, and stays None when Playwright fails to start - so the service
177
+ has to read it per call rather than capture it at construction."""
178
+ service = app_module.get_search_service()
179
+ monkeypatch.setattr(app_module, "pw_browser", "a-browser")
180
 
181
+ assert service.browser == "a-browser"
 
 
182
 
 
183
 
184
+ def test_the_default_patent_service_gets_the_apps_collaborators():
185
+ service = app_module.get_patent_service()
186
 
187
+ assert service._http is app_module.httpx_client
188
+ assert service._ops_tokens is app_module.ops_token_manager
 
189
 
190
 
191
  # ------------------------------------- api_lifespan -------------------------------------
192
 
193
 
194
+ class _RecordingClient:
195
+ def __init__(self):
196
+ self.closed = False
197
+
198
+ async def aclose(self):
199
+ self.closed = True
200
+
201
+
202
  async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
203
  """The container runs as a non-root user (see the Dockerfile), and
204
  Chromium's own internal sandbox needs privileges a non-root container
 
244
  assert "--no-sandbox" in launch_calls[0].get("args", [])
245
 
246
 
 
 
 
 
 
 
 
 
247
  async def test_api_lifespan_closes_the_http_client_on_shutdown(monkeypatch):
248
  """The client was created at import and never closed - the browser was
249
  torn down on shutdown but its connection pool was not.
250
  """
 
 
 
 
 
 
251
  recording_client = _RecordingClient()
252
  monkeypatch.setattr(app_module, "pw_browser", None)
253
  monkeypatch.setattr(app_module, "playwright", None)
 
276
 
277
  async with app_module.api_lifespan(app_module.app):
278
  assert app_module.pw_browser is None
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
tests/test_auth.py CHANGED
@@ -15,6 +15,7 @@ import pytest
15
  from httpx import ASGITransport
16
 
17
  import app as app_module
 
18
 
19
 
20
  @pytest.fixture
@@ -47,7 +48,7 @@ async def test_no_key_configured_leaves_the_api_open(client, monkeypatch):
47
  """The middleware is a no-op unless SERPENT_API_KEY is set - this keeps
48
  the default deployment exactly as open as it was before it existed.
49
  """
50
- monkeypatch.setattr(app_module, "query_arxiv", _stub_arxiv)
51
 
52
  resp = await client.post("/serp/search_arxiv", json={"queries": ["a"]})
53
 
@@ -76,7 +77,7 @@ async def test_wrong_key_is_rejected(client, api_key):
76
 
77
 
78
  async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, monkeypatch):
79
- monkeypatch.setattr(app_module, "query_arxiv", _stub_arxiv)
80
 
81
  resp = await client.post(
82
  "/serp/search_arxiv", json={"queries": ["a"]}, headers={"X-API-Key": api_key})
@@ -85,7 +86,7 @@ async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, mon
85
 
86
 
87
  async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkeypatch):
88
- monkeypatch.setattr(app_module, "query_arxiv", _stub_arxiv)
89
 
90
  resp = await client.post(
91
  "/serp/search_arxiv", json={"queries": ["a"]},
@@ -95,7 +96,7 @@ async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkey
95
 
96
 
97
  async def test_bearer_scheme_match_is_case_insensitive(client, api_key, monkeypatch):
98
- monkeypatch.setattr(app_module, "query_arxiv", _stub_arxiv)
99
 
100
  resp = await client.post(
101
  "/serp/search_arxiv", json={"queries": ["a"]},
 
15
  from httpx import ASGITransport
16
 
17
  import app as app_module
18
+ import services as services_module
19
 
20
 
21
  @pytest.fixture
 
48
  """The middleware is a no-op unless SERPENT_API_KEY is set - this keeps
49
  the default deployment exactly as open as it was before it existed.
50
  """
51
+ monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
52
 
53
  resp = await client.post("/serp/search_arxiv", json={"queries": ["a"]})
54
 
 
77
 
78
 
79
  async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, monkeypatch):
80
+ monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
81
 
82
  resp = await client.post(
83
  "/serp/search_arxiv", json={"queries": ["a"]}, headers={"X-API-Key": api_key})
 
86
 
87
 
88
  async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkeypatch):
89
+ monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
90
 
91
  resp = await client.post(
92
  "/serp/search_arxiv", json={"queries": ["a"]},
 
96
 
97
 
98
  async def test_bearer_scheme_match_is_case_insensitive(client, api_key, monkeypatch):
99
+ monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
100
 
101
  resp = await client.post(
102
  "/serp/search_arxiv", json={"queries": ["a"]},
tests/test_input_bounds.py CHANGED
@@ -16,6 +16,7 @@ import app as app_module
16
  import scrap as scrap_module
17
  from app import MAX_BULK_PATENT_IDS, ScrapPatentsRequest
18
  from serp import MAX_QUERIES_PER_REQUEST, SerpQuery
 
19
 
20
 
21
  @pytest.fixture
@@ -57,9 +58,9 @@ async def test_oversized_query_list_returns_422(client):
57
 
58
  def test_shape_serp_results_is_total_for_an_empty_list():
59
  """Defence in depth: the helper must not depend on its caller having
60
- validated the query list, since it is shared by every search endpoint.
61
  """
62
- result = app_module._shape_serp_results([])
63
 
64
  assert result.results == []
65
  assert result.error is not None
 
16
  import scrap as scrap_module
17
  from app import MAX_BULK_PATENT_IDS, ScrapPatentsRequest
18
  from serp import MAX_QUERIES_PER_REQUEST, SerpQuery
19
+ from services import shape_serp_results
20
 
21
 
22
  @pytest.fixture
 
58
 
59
  def test_shape_serp_results_is_total_for_an_empty_list():
60
  """Defence in depth: the helper must not depend on its caller having
61
+ validated the query list, since it is shared by every search path.
62
  """
63
+ result = shape_serp_results([])
64
 
65
  assert result.results == []
66
  assert result.error is not None
tests/test_patent_fallback.py CHANGED
@@ -1,13 +1,14 @@
1
- """The Google Patents -> EPO OPS fallback chain in app.py.
2
 
3
- This is the most heavily-branched code in the module - scrape, catch,
4
- check whether OPS is configured, fall back, then map OPS's own failures
5
- onto status codes - and it had no test coverage at all.
 
6
 
7
- The distinction these tests exist to protect: a 404 means "this patent
8
- does not exist", and an agent is explicitly told by MCP_INSTRUCTIONS to
9
- believe it and move on without retrying. Anything else - an upstream
10
- outage, a timeout, a blocked scrape - must not be reported that way.
11
  """
12
 
13
  import httpx
@@ -15,28 +16,26 @@ import pytest
15
  from httpx import ASGITransport, HTTPStatusError, Request, Response
16
 
17
  import app as app_module
 
18
  from scrap import PatentScrapResult
 
 
19
 
20
  PATENT_ID = "US11930446B2"
21
 
22
 
23
- @pytest.fixture
24
- async def client():
25
- transport = ASGITransport(app=app_module.app)
26
- async with httpx.AsyncClient(transport=transport, base_url="http://test") as c:
27
- yield c
28
-
29
 
30
- @pytest.fixture
31
- def ops_configured(monkeypatch):
32
- monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
33
- monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
34
 
35
 
36
- @pytest.fixture
37
- def ops_unconfigured(monkeypatch):
38
- monkeypatch.setattr(app_module.ops_token_manager, "_key", None)
39
- monkeypatch.setattr(app_module.ops_token_manager, "_secret", None)
40
 
41
 
42
  def _http_error(status: int) -> HTTPStatusError:
@@ -51,10 +50,16 @@ def _raises(exc):
51
  return _fn
52
 
53
 
 
 
 
 
 
 
54
  # ------------------------------- the happy path -------------------------------
55
 
56
 
57
- async def test_successful_scrape_never_touches_ops(client, ops_configured, monkeypatch):
58
  ops_called = False
59
 
60
  async def fake_ops(*args, **kwargs):
@@ -63,133 +68,165 @@ async def test_successful_scrape_never_touches_ops(client, ops_configured, monke
63
  return PatentScrapResult(title="from OPS")
64
 
65
  monkeypatch.setattr(
66
- app_module, "scrap_patent_async",
67
- lambda c, url: _ok(PatentScrapResult(title="from Google Patents")))
68
- monkeypatch.setattr(app_module, "ops_scrap_patent", fake_ops)
69
 
70
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
71
 
72
- assert resp.status_code == 200
73
- assert resp.json()["title"] == "from Google Patents"
74
  assert ops_called is False
75
 
76
 
77
- async def _ok(value):
78
- return value
79
-
80
-
81
  # ------------------------- genuine miss vs upstream failure -------------------------
82
 
83
 
84
- async def test_a_real_404_from_google_patents_is_a_404(client, ops_unconfigured, monkeypatch):
85
  """The one case where "not found" is the truth."""
86
- monkeypatch.setattr(app_module, "scrap_patent_async", _raises(_http_error(404)))
87
-
88
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
89
 
90
- assert resp.status_code == 404
 
91
 
92
 
93
  @pytest.mark.parametrize("status", [429, 500, 502, 503])
94
- async def test_an_upstream_failure_is_not_reported_as_not_found(
95
- client, ops_unconfigured, monkeypatch, status):
96
  """MCP_INSTRUCTIONS tells the model a "not found" patent is genuinely
97
  absent everywhere and to move on rather than retrying. Reporting a
98
  transient upstream failure that way teaches an agent - with the
99
  server's explicit encouragement - that a real patent does not exist.
100
  """
101
- monkeypatch.setattr(app_module, "scrap_patent_async", _raises(_http_error(status)))
102
 
103
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
 
104
 
105
- assert resp.status_code == 502, f"upstream {status} was reported as {resp.status_code}"
106
 
107
 
108
- async def test_a_timeout_is_reported_as_a_gateway_timeout(client, ops_unconfigured, monkeypatch):
109
  monkeypatch.setattr(
110
- app_module, "scrap_patent_async", _raises(httpx.ConnectTimeout("timed out")))
111
 
112
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
 
113
 
114
- assert resp.status_code == 504
115
 
116
 
117
- async def test_an_unparseable_page_is_reported_as_a_bad_gateway(
118
- client, ops_unconfigured, monkeypatch):
119
  """parse_patent_html raises ValueError when the page isn't a patent page
120
  (an interstitial, or a markup change). That is our problem or theirs,
121
  but it is not evidence the patent doesn't exist.
122
  """
123
  monkeypatch.setattr(
124
- app_module, "scrap_patent_async", _raises(ValueError("no DC.title")))
125
-
126
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
127
 
128
- assert resp.status_code == 502
 
129
 
130
 
131
  # --------------------------------- the OPS fallback ---------------------------------
132
 
133
 
134
- async def test_falls_back_to_ops_when_google_patents_fails(client, ops_configured, monkeypatch):
135
- monkeypatch.setattr(app_module, "scrap_patent_async", _raises(_http_error(503)))
136
  monkeypatch.setattr(
137
- app_module, "ops_scrap_patent",
138
- lambda c, pid: _ok(PatentScrapResult(title="from OPS")))
139
 
140
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
141
 
142
- assert resp.status_code == 200
143
- assert resp.json()["title"] == "from OPS"
144
 
145
 
146
- async def test_missing_from_both_backends_is_a_404(client, ops_configured, monkeypatch):
147
- monkeypatch.setattr(app_module, "scrap_patent_async", _raises(_http_error(404)))
148
- monkeypatch.setattr(app_module, "ops_scrap_patent", _raises(_http_error(404)))
149
 
150
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
 
151
 
152
- assert resp.status_code == 404
153
- assert "not found" in resp.json()["detail"].lower()
154
 
155
-
156
- async def test_ops_failing_after_a_google_patents_404_is_not_a_404(
157
- client, ops_configured, monkeypatch):
158
  """Google Patents says the patent is missing, but OPS - the backend that
159
  covers what Google Patents doesn't - never answered. That is unknown,
160
  not absent.
161
  """
162
- monkeypatch.setattr(app_module, "scrap_patent_async", _raises(_http_error(404)))
163
- monkeypatch.setattr(app_module, "ops_scrap_patent", _raises(_http_error(500)))
164
 
165
- resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
 
 
 
 
 
 
 
 
 
166
 
167
- assert resp.status_code == 502
168
 
 
 
 
169
 
170
- # ------------------------------ the direct /ops endpoint ------------------------------
 
171
 
172
 
173
- async def test_ops_endpoint_reports_a_timeout_as_a_gateway_timeout(
174
- client, ops_configured, monkeypatch):
175
  monkeypatch.setattr(
176
- app_module, "ops_scrap_patent", _raises(httpx.ConnectTimeout("timed out")))
177
 
178
- resp = await client.get(f"/ops/scrap_patent/{PATENT_ID}")
 
179
 
180
- assert resp.status_code == 504
181
 
182
 
183
- async def test_ops_endpoint_reports_a_real_miss_as_404(client, ops_configured, monkeypatch):
184
- """Here OPS is the only backend asked, so its 404 is the whole answer."""
185
- monkeypatch.setattr(app_module, "ops_scrap_patent", _raises(_http_error(404)))
186
 
187
- resp = await client.get(f"/ops/scrap_patent/{PATENT_ID}")
188
 
189
- assert resp.status_code == 404
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
190
 
 
 
 
191
 
192
- async def test_ops_endpoint_is_503_when_not_configured(client, ops_unconfigured):
193
- resp = await client.get(f"/ops/scrap_patent/{PATENT_ID}")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
194
 
195
- assert resp.status_code == 503
 
 
1
+ """The Google Patents -> EPO OPS fallback chain.
2
 
3
+ The policy lives in services.PatentService, so most of this is plain unit
4
+ testing with no ASGI client involved. Only the last section goes through
5
+ the app, and only to check that the domain errors the service raises map
6
+ onto the right status codes.
7
 
8
+ The distinction these tests exist to protect: a 404 means "this patent does
9
+ not exist", and an agent is explicitly told by MCP_INSTRUCTIONS to believe
10
+ it and move on without retrying. Anything else - an upstream outage, a
11
+ timeout, a blocked scrape - must not be reported that way.
12
  """
13
 
14
  import httpx
 
16
  from httpx import ASGITransport, HTTPStatusError, Request, Response
17
 
18
  import app as app_module
19
+ import services as services_module
20
  from scrap import PatentScrapResult
21
+ from services import (OPSUnconfigured, PatentNotFound, PatentService,
22
+ UpstreamUnavailable)
23
 
24
  PATENT_ID = "US11930446B2"
25
 
26
 
27
+ class FakeOPSTokens:
28
+ """Stand-in for the OPS token manager. Tests only care whether
29
+ credentials are configured, and this avoids assigning to private
30
+ attributes on the real module-level singleton.
31
+ """
 
32
 
33
+ def __init__(self, configured: bool):
34
+ self.configured = configured
 
 
35
 
36
 
37
+ def make_service(*, ops_configured: bool) -> PatentService:
38
+ return PatentService(http_client=None, ops_tokens=FakeOPSTokens(ops_configured))
 
 
39
 
40
 
41
  def _http_error(status: int) -> HTTPStatusError:
 
50
  return _fn
51
 
52
 
53
+ def _returns(value):
54
+ async def _fn(*args, **kwargs):
55
+ return value
56
+ return _fn
57
+
58
+
59
  # ------------------------------- the happy path -------------------------------
60
 
61
 
62
+ async def test_successful_scrape_never_touches_ops(monkeypatch):
63
  ops_called = False
64
 
65
  async def fake_ops(*args, **kwargs):
 
68
  return PatentScrapResult(title="from OPS")
69
 
70
  monkeypatch.setattr(
71
+ services_module, "scrap_patent_async",
72
+ _returns(PatentScrapResult(title="from Google Patents")))
73
+ monkeypatch.setattr(services_module, "ops_scrap_patent", fake_ops)
74
 
75
+ result = await make_service(ops_configured=True).scrap(PATENT_ID)
76
 
77
+ assert result.title == "from Google Patents"
 
78
  assert ops_called is False
79
 
80
 
 
 
 
 
81
  # ------------------------- genuine miss vs upstream failure -------------------------
82
 
83
 
84
+ async def test_a_real_404_from_google_patents_is_not_found(monkeypatch):
85
  """The one case where "not found" is the truth."""
86
+ monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
 
 
87
 
88
+ with pytest.raises(PatentNotFound):
89
+ await make_service(ops_configured=False).scrap(PATENT_ID)
90
 
91
 
92
  @pytest.mark.parametrize("status", [429, 500, 502, 503])
93
+ async def test_an_upstream_failure_is_not_reported_as_not_found(monkeypatch, status):
 
94
  """MCP_INSTRUCTIONS tells the model a "not found" patent is genuinely
95
  absent everywhere and to move on rather than retrying. Reporting a
96
  transient upstream failure that way teaches an agent - with the
97
  server's explicit encouragement - that a real patent does not exist.
98
  """
99
+ monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(status)))
100
 
101
+ with pytest.raises(UpstreamUnavailable) as exc_info:
102
+ await make_service(ops_configured=False).scrap(PATENT_ID)
103
 
104
+ assert exc_info.value.timeout is False
105
 
106
 
107
+ async def test_a_timeout_is_flagged_as_a_timeout(monkeypatch):
108
  monkeypatch.setattr(
109
+ services_module, "scrap_patent_async", _raises(httpx.ConnectTimeout("timed out")))
110
 
111
+ with pytest.raises(UpstreamUnavailable) as exc_info:
112
+ await make_service(ops_configured=False).scrap(PATENT_ID)
113
 
114
+ assert exc_info.value.timeout is True
115
 
116
 
117
+ async def test_an_unparseable_page_is_an_upstream_failure(monkeypatch):
 
118
  """parse_patent_html raises ValueError when the page isn't a patent page
119
  (an interstitial, or a markup change). That is our problem or theirs,
120
  but it is not evidence the patent doesn't exist.
121
  """
122
  monkeypatch.setattr(
123
+ services_module, "scrap_patent_async", _raises(ValueError("no DC.title")))
 
 
124
 
125
+ with pytest.raises(UpstreamUnavailable):
126
+ await make_service(ops_configured=False).scrap(PATENT_ID)
127
 
128
 
129
  # --------------------------------- the OPS fallback ---------------------------------
130
 
131
 
132
+ async def test_falls_back_to_ops_when_google_patents_fails(monkeypatch):
133
+ monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(503)))
134
  monkeypatch.setattr(
135
+ services_module, "ops_scrap_patent", _returns(PatentScrapResult(title="from OPS")))
 
136
 
137
+ result = await make_service(ops_configured=True).scrap(PATENT_ID)
138
 
139
+ assert result.title == "from OPS"
 
140
 
141
 
142
+ async def test_missing_from_both_backends_is_not_found(monkeypatch):
143
+ monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
144
+ monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(404)))
145
 
146
+ with pytest.raises(PatentNotFound):
147
+ await make_service(ops_configured=True).scrap(PATENT_ID)
148
 
 
 
149
 
150
+ async def test_ops_failing_after_a_google_patents_404_is_not_a_miss(monkeypatch):
 
 
151
  """Google Patents says the patent is missing, but OPS - the backend that
152
  covers what Google Patents doesn't - never answered. That is unknown,
153
  not absent.
154
  """
155
+ monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
156
+ monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(500)))
157
 
158
+ with pytest.raises(UpstreamUnavailable):
159
+ await make_service(ops_configured=True).scrap(PATENT_ID)
160
+
161
+
162
+ # ------------------------------ the OPS-only path ------------------------------
163
+
164
+
165
+ async def test_ops_only_retrieval_requires_credentials():
166
+ with pytest.raises(OPSUnconfigured):
167
+ await make_service(ops_configured=False).ops_scrap(PATENT_ID)
168
 
 
169
 
170
+ async def test_ops_only_retrieval_reports_a_real_miss(monkeypatch):
171
+ """Here OPS is the only backend asked, so its 404 is the whole answer."""
172
+ monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(404)))
173
 
174
+ with pytest.raises(PatentNotFound):
175
+ await make_service(ops_configured=True).ops_scrap(PATENT_ID)
176
 
177
 
178
+ async def test_ops_only_retrieval_flags_a_timeout(monkeypatch):
 
179
  monkeypatch.setattr(
180
+ services_module, "ops_scrap_patent", _raises(httpx.ConnectTimeout("timed out")))
181
 
182
+ with pytest.raises(UpstreamUnavailable) as exc_info:
183
+ await make_service(ops_configured=True).ops_scrap(PATENT_ID)
184
 
185
+ assert exc_info.value.timeout is True
186
 
187
 
188
+ # ---------------------------- domain error -> status code ----------------------------
 
 
189
 
 
190
 
191
+ @pytest.fixture
192
+ async def client():
193
+ transport = ASGITransport(app=app_module.app)
194
+ async with httpx.AsyncClient(transport=transport, base_url="http://test") as c:
195
+ yield c
196
+
197
+
198
+ @pytest.fixture
199
+ def override_patent_service():
200
+ """Swap the app's patent service for one the test controls, then put the
201
+ real provider back."""
202
+ def _install(service):
203
+ app_module.app.dependency_overrides[app_module.get_patent_service] = lambda: service
204
+
205
+ yield _install
206
+ app_module.app.dependency_overrides.clear()
207
+
208
 
209
+ class _StubService:
210
+ def __init__(self, error):
211
+ self._error = error
212
 
213
+ async def scrap(self, patent_id):
214
+ raise self._error
215
+
216
+ ops_scrap = scrap
217
+
218
+
219
+ @pytest.mark.parametrize("error, expected_status", [
220
+ (PatentNotFound("gone"), 404),
221
+ (UpstreamUnavailable("upstream broke"), 502),
222
+ (UpstreamUnavailable("slow", timeout=True), 504),
223
+ (OPSUnconfigured("no credentials"), 503),
224
+ ])
225
+ async def test_domain_errors_map_onto_status_codes(
226
+ client, override_patent_service, error, expected_status):
227
+ override_patent_service(_StubService(error))
228
+
229
+ resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
230
 
231
+ assert resp.status_code == expected_status
232
+ assert resp.json()["detail"] == str(error)
tests/test_scrap.py CHANGED
@@ -89,3 +89,15 @@ async def test_scrap_patent_bulk_async_separates_successes_from_failures():
89
  assert len(result.patents) == 1
90
  assert result.patents[0].title == "Widget with improved gadget mechanism"
91
  assert result.failed_ids == ["US_FAIL"]
 
 
 
 
 
 
 
 
 
 
 
 
 
89
  assert len(result.patents) == 1
90
  assert result.patents[0].title == "Widget with improved gadget mechanism"
91
  assert result.failed_ids == ["US_FAIL"]
92
+
93
+
94
+ def test_parse_patent_html_strips_whitespace_around_the_title():
95
+ """Google Patents' DC.title meta content is frequently padded with
96
+ newlines and indentation from the surrounding markup, which would
97
+ otherwise end up in the API response and in every downstream citation.
98
+ """
99
+ html = '<html><head><meta name="DC.title" content=" Widget apparatus\n "></head><body></body></html>'
100
+
101
+ result = parse_patent_html(html, "https://patents.google.com/patent/US1/en")
102
+
103
+ assert result.title == "Widget apparatus"
tests/test_search_service.py ADDED
@@ -0,0 +1,390 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """SearchService: the backend orchestration policy.
2
+
3
+ Which backend to try, in what order, when to fall back, and what counts as
4
+ evidence a backend is unhealthy. None of this needs an HTTP layer, so none
5
+ of these tests use one.
6
+
7
+ Each test builds a service with its own circuit breaker. That replaces the
8
+ autouse fixture that used to reset a process-wide singleton between tests -
9
+ the breaker being shared across every request is correct in production and
10
+ was only a problem because tests had no way to opt out of it.
11
+ """
12
+
13
+
14
+ import services as services_module
15
+ from circuit_breaker import CircuitBreaker
16
+ from serp import DuckDuckGoBlockedException, SerpQuery
17
+ from services import SearchService, shape_serp_results
18
+
19
+
20
+ class FakeOPSTokens:
21
+ def __init__(self, configured: bool):
22
+ self.configured = configured
23
+
24
+
25
+ def make_service(*, ops_configured: bool = False, breaker: CircuitBreaker = None,
26
+ browser=None) -> SearchService:
27
+ return SearchService(
28
+ http_client=None,
29
+ browser_provider=lambda: browser,
30
+ circuit_breaker=breaker or CircuitBreaker(),
31
+ ops_tokens=FakeOPSTokens(ops_configured),
32
+ )
33
+
34
+
35
+ # --------------------------------- result shaping ---------------------------------
36
+
37
+
38
+ class TestShapeSerpResults:
39
+ def test_flattens_successful_lists(self):
40
+ result = shape_serp_results([[{"a": 1}], [{"b": 2}, {"c": 3}]])
41
+ assert result.results == [{"a": 1}, {"b": 2}, {"c": 3}]
42
+ assert result.error is None
43
+
44
+ def test_surfaces_last_error_when_all_failed(self):
45
+ result = shape_serp_results([ValueError("first"), RuntimeError("second")])
46
+ assert result.results == []
47
+ assert "second" in result.error
48
+
49
+ def test_ignores_failures_when_some_succeed(self):
50
+ result = shape_serp_results([[{"a": 1}], RuntimeError("boom")])
51
+ assert result.results == [{"a": 1}]
52
+ assert result.error is None
53
+
54
+ def test_is_total_for_an_empty_list(self):
55
+ """Defence in depth: the request models bound the query list, but
56
+ this helper is shared by every search path and must not depend on
57
+ its caller having validated anything."""
58
+ result = shape_serp_results([])
59
+ assert result.results == []
60
+ assert result.error is not None
61
+
62
+
63
+ # ------------------------------- per-query dispatch -------------------------------
64
+
65
+
66
+ async def test_every_query_is_dispatched_with_the_result_count():
67
+ calls = []
68
+
69
+ async def fn(q, n):
70
+ calls.append((q, n))
71
+ return [{"q": q}]
72
+
73
+ result = await make_service()._run_queries(
74
+ fn, SerpQuery(queries=["x", "y"], n_results=5), "test context")
75
+
76
+ assert sorted(calls) == [("x", 5), ("y", 5)]
77
+ assert sorted(result.results, key=str) == sorted([{"q": "x"}, {"q": "y"}], key=str)
78
+
79
+
80
+ async def test_arxiv_flattens_results_from_all_queries(monkeypatch):
81
+ async def fake_query_arxiv(client, q, n):
82
+ return [{"title": f"paper about {q}"}]
83
+
84
+ monkeypatch.setattr(services_module, "query_arxiv", fake_query_arxiv)
85
+
86
+ result = await make_service().arxiv(SerpQuery(queries=["a", "b"], n_results=5))
87
+
88
+ assert {r["title"] for r in result.results} == {"paper about a", "paper about b"}
89
+ assert result.error is None
90
+
91
+
92
+ async def test_arxiv_reports_an_error_when_every_query_fails(monkeypatch):
93
+ async def failing(client, q, n):
94
+ raise RuntimeError(f"boom for {q}")
95
+
96
+ monkeypatch.setattr(services_module, "query_arxiv", failing)
97
+
98
+ result = await make_service().arxiv(SerpQuery(queries=["a"], n_results=5))
99
+
100
+ assert result.results == []
101
+ assert "boom for a" in result.error
102
+
103
+
104
+ # ------------------------------ patents + OPS fallback ------------------------------
105
+
106
+
107
+ async def test_patents_does_not_fall_back_when_google_patents_succeeds(monkeypatch):
108
+ ops_called = False
109
+
110
+ async def fake_ops(client, q, n):
111
+ nonlocal ops_called
112
+ ops_called = True
113
+ return []
114
+
115
+ monkeypatch.setattr(
116
+ services_module, "query_google_patents",
117
+ lambda b, q, n: _ok([{"id": "US1", "title": "t"}]))
118
+ monkeypatch.setattr(services_module, "ops_search", fake_ops)
119
+
120
+ result = await make_service(ops_configured=True).patents(
121
+ SerpQuery(queries=["widget"], n_results=5))
122
+
123
+ assert result.results == [{"id": "US1", "title": "t"}]
124
+ assert ops_called is False
125
+
126
+
127
+ async def _ok(value):
128
+ return value
129
+
130
+
131
+ async def test_patents_falls_back_to_ops_when_google_patents_is_empty(monkeypatch):
132
+ monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
133
+ monkeypatch.setattr(
134
+ services_module, "ops_search", lambda c, q, n: _ok([{"id": "US123", "title": "OPS"}]))
135
+
136
+ result = await make_service(ops_configured=True).patents(
137
+ SerpQuery(queries=["widget"], n_results=5))
138
+
139
+ assert result.results == [{"id": "US123", "title": "OPS"}]
140
+
141
+
142
+ async def test_patents_does_not_fall_back_when_ops_is_unconfigured(monkeypatch):
143
+ monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
144
+
145
+ result = await make_service(ops_configured=False).patents(
146
+ SerpQuery(queries=["widget"], n_results=5))
147
+
148
+ assert result.results == []
149
+ assert result.error is None
150
+
151
+
152
+ async def test_ops_fallback_runs_concurrently_not_one_query_at_a_time(monkeypatch):
153
+ """This path exists for when Google Patents is unavailable, which is
154
+ exactly when a request can least afford N sequential round-trips to the
155
+ slower backend."""
156
+ import asyncio
157
+
158
+ in_flight = 0
159
+ peak = 0
160
+
161
+ async def slow_ops(client, q, n):
162
+ nonlocal in_flight, peak
163
+ in_flight += 1
164
+ peak = max(peak, in_flight)
165
+ await asyncio.sleep(0.02)
166
+ in_flight -= 1
167
+ return [{"id": q}]
168
+
169
+ monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
170
+ monkeypatch.setattr(services_module, "ops_search", slow_ops)
171
+
172
+ await make_service(ops_configured=True).patents(
173
+ SerpQuery(queries=["a", "b", "c", "d"], n_results=5))
174
+
175
+ assert peak > 1, "OPS fallback queries were issued one at a time"
176
+
177
+
178
+ async def test_a_failing_ops_fallback_does_not_lose_other_queries_results(monkeypatch):
179
+ async def selective_ops(client, q, n):
180
+ if q == "bad":
181
+ raise RuntimeError("OPS down")
182
+ return [{"id": q}]
183
+
184
+ monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
185
+ monkeypatch.setattr(services_module, "ops_search", selective_ops)
186
+
187
+ result = await make_service(ops_configured=True).patents(
188
+ SerpQuery(queries=["good", "bad"], n_results=5))
189
+
190
+ assert result.results == [{"id": "good"}]
191
+
192
+
193
+ # -------------------------------- the fallback chain --------------------------------
194
+
195
+
196
+ async def test_returns_the_first_backend_that_succeeds(monkeypatch):
197
+ calls = []
198
+
199
+ async def fake_ddg(q, n):
200
+ calls.append("ddg")
201
+ return [{"title": "from ddg"}]
202
+
203
+ async def fake_brave(browser, q, n):
204
+ calls.append("brave")
205
+ return [{"title": "from brave"}]
206
+
207
+ monkeypatch.setattr(services_module, "query_ddg_search", fake_ddg)
208
+ monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
209
+
210
+ q, results, err = await make_service()._search_one("widget", 5)
211
+
212
+ assert calls == ["ddg"]
213
+ assert results == [{"title": "from ddg"}]
214
+ assert err is None
215
+
216
+
217
+ async def test_falls_through_to_the_next_backend_on_failure(monkeypatch):
218
+ calls = []
219
+
220
+ async def failing_ddg(q, n):
221
+ calls.append("ddg")
222
+ raise RuntimeError("ddg blocked")
223
+
224
+ async def fake_brave(browser, q, n):
225
+ calls.append("brave")
226
+ return [{"title": "from brave"}]
227
+
228
+ monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
229
+ monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
230
+
231
+ q, results, err = await make_service()._search_one("widget", 5)
232
+
233
+ assert calls == ["ddg", "brave"]
234
+ assert results == [{"title": "from brave"}]
235
+ assert err is None
236
+
237
+
238
+ async def test_reports_an_error_when_every_backend_fails(monkeypatch):
239
+ async def failing(*args):
240
+ raise RuntimeError("blocked")
241
+
242
+ for name in ("query_ddg_search", "query_brave_search", "query_bing_search"):
243
+ monkeypatch.setattr(services_module, name, failing)
244
+
245
+ q, results, err = await make_service()._search_one("widget", 5)
246
+
247
+ assert results == []
248
+ assert "All backends failed for query 'widget'" in err
249
+ assert "blocked" in err
250
+
251
+
252
+ # ------------------------------ the circuit breaker ------------------------------
253
+
254
+
255
+ async def test_skips_a_backend_whose_circuit_is_open(monkeypatch):
256
+ """A backend that's failed `failure_threshold` times in a row shouldn't
257
+ be attempted again until its cooldown elapses - repeatedly hammering a
258
+ backend that's already rate-limiting us only makes that worse.
259
+ """
260
+ ddg_calls = []
261
+
262
+ async def failing_ddg(q, n):
263
+ ddg_calls.append(q)
264
+ raise RuntimeError("ddg blocked")
265
+
266
+ async def fake_brave(browser, q, n):
267
+ return [{"title": "from brave"}]
268
+
269
+ monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
270
+ monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
271
+
272
+ service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
273
+
274
+ await service._search_one("first query", 5)
275
+ assert ddg_calls == ["first query"]
276
+
277
+ q, results, err = await service._search_one("second query", 5)
278
+
279
+ assert ddg_calls == ["first query"] # unchanged - DDG was skipped
280
+ assert results == [{"title": "from brave"}]
281
+
282
+
283
+ async def test_a_zero_result_query_does_not_open_the_circuit(monkeypatch):
284
+ """An empty result is ambiguous - it can mean the query genuinely has no
285
+ hits, or that the backend served an anti-bot page that parsed to
286
+ nothing. The breaker is shared across every request, so counting the
287
+ ambiguous case meant a few unusual-but-legitimate queries could disable
288
+ the primary backend for everyone for a full cooldown.
289
+ """
290
+ ddg_calls = []
291
+
292
+ async def empty_ddg(q, n):
293
+ ddg_calls.append(q)
294
+ raise DuckDuckGoBlockedException()
295
+
296
+ monkeypatch.setattr(services_module, "query_ddg_search", empty_ddg)
297
+ monkeypatch.setattr(
298
+ services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "brave"}]))
299
+
300
+ service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
301
+
302
+ await service._search_one("obscure", 5)
303
+ await service._search_one("another obscure", 5)
304
+
305
+ assert ddg_calls == ["obscure", "another obscure"]
306
+
307
+
308
+ async def test_an_empty_result_still_falls_through_to_the_next_backend(monkeypatch):
309
+ """Not counting it against the backend must not stop the fallback: the
310
+ caller still needs results from somewhere."""
311
+ async def empty_ddg(q, n):
312
+ raise DuckDuckGoBlockedException()
313
+
314
+ monkeypatch.setattr(services_module, "query_ddg_search", empty_ddg)
315
+ monkeypatch.setattr(
316
+ services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "from brave"}]))
317
+
318
+ _q, results, err = await make_service()._search_one("widget", 5)
319
+
320
+ assert results == [{"title": "from brave"}]
321
+ assert err is None
322
+
323
+
324
+ async def test_a_real_backend_error_still_opens_the_circuit(monkeypatch):
325
+ """The distinction has to cut both ways: a backend raising a genuine
326
+ error (a rate limit, a transport failure) is still evidence."""
327
+ ddg_calls = []
328
+
329
+ async def failing_ddg(q, n):
330
+ ddg_calls.append(q)
331
+ raise RuntimeError("rate limited")
332
+
333
+ monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
334
+ monkeypatch.setattr(
335
+ services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "brave"}]))
336
+
337
+ service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
338
+
339
+ await service._search_one("first", 5)
340
+ await service._search_one("second", 5)
341
+
342
+ assert ddg_calls == ["first"]
343
+
344
+
345
+ # ------------------------------- search aggregation -------------------------------
346
+
347
+
348
+ async def test_search_merges_results_from_every_query(monkeypatch):
349
+ monkeypatch.setattr(
350
+ services_module, "query_ddg_search", lambda q, n: _ok([{"title": f"result for {q}"}]))
351
+
352
+ result = await make_service().search(SerpQuery(queries=["a", "b"], n_results=5))
353
+
354
+ assert {r["title"] for r in result.results} == {"result for a", "result for b"}
355
+ assert result.error is None
356
+
357
+
358
+ async def test_search_reports_partial_failure_alongside_results(monkeypatch):
359
+ """A query no backend could answer must be reported even when another
360
+ query succeeded - otherwise the caller silently receives fewer results
361
+ than they asked for with no indication why."""
362
+ async def selective_ddg(q, n):
363
+ if q == "bad":
364
+ raise RuntimeError("blocked")
365
+ return [{"title": f"result for {q}"}]
366
+
367
+ async def failing(*args):
368
+ raise RuntimeError("blocked")
369
+
370
+ monkeypatch.setattr(services_module, "query_ddg_search", selective_ddg)
371
+ monkeypatch.setattr(services_module, "query_brave_search", failing)
372
+ monkeypatch.setattr(services_module, "query_bing_search", failing)
373
+
374
+ result = await make_service().search(SerpQuery(queries=["good", "bad"], n_results=5))
375
+
376
+ assert [r["title"] for r in result.results] == ["result for good"]
377
+ assert "bad" in result.error
378
+
379
+
380
+ async def test_search_names_every_failed_query_not_just_the_last(monkeypatch):
381
+ async def failing(*args):
382
+ raise RuntimeError("blocked")
383
+
384
+ for name in ("query_ddg_search", "query_brave_search", "query_bing_search"):
385
+ monkeypatch.setattr(services_module, name, failing)
386
+
387
+ result = await make_service().search(SerpQuery(queries=["a", "b"], n_results=5))
388
+
389
+ assert result.results == []
390
+ assert "'a'" in result.error and "'b'" in result.error
tests/test_serp_navigation.py ADDED
@@ -0,0 +1,156 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The glue holding each scraper together.
2
+
3
+ Every Playwright scraper is now: build a URL, open a page, block
4
+ decorative resources, navigate, extract. The builders are tested in
5
+ test_serp_urls.py and the extractors in test_serp_playwright.py; this
6
+ covers the composition - specifically that each scraper navigates to the
7
+ URL its own builder produced.
8
+
9
+ That link is the one the old goto-patching harness could not test, because
10
+ the stand-in it installed accepted the navigation URL and discarded it.
11
+ These use a recording stand-in browser instead: no real Chromium, and the
12
+ URL is the thing being asserted on.
13
+ """
14
+
15
+ import pytest
16
+
17
+ import serp
18
+ from serp import (bing_search_url, brave_search_url, google_patents_search_url,
19
+ google_scholar_url, playwright_open_page,
20
+ query_bing_search, query_brave_search, query_google_patents,
21
+ query_google_scholar)
22
+
23
+ SENTINEL = [{"title": "extracted"}]
24
+
25
+
26
+ class _RecordingPage:
27
+ def __init__(self):
28
+ self.goto_urls = []
29
+ self.route_patterns = []
30
+ self.closed = False
31
+
32
+ async def route(self, pattern, handler):
33
+ self.route_patterns.append(pattern)
34
+
35
+ async def goto(self, url, **kwargs):
36
+ self.goto_urls.append(url)
37
+ return None
38
+
39
+ async def close(self):
40
+ self.closed = True
41
+
42
+
43
+ class _RecordingContext:
44
+ def __init__(self, page):
45
+ self._page = page
46
+ self.closed = False
47
+
48
+ async def new_page(self):
49
+ return self._page
50
+
51
+ async def close(self):
52
+ self.closed = True
53
+
54
+
55
+ class _RecordingBrowser:
56
+ def __init__(self):
57
+ self.page = _RecordingPage()
58
+ self.context = _RecordingContext(self.page)
59
+
60
+ async def new_context(self, **kwargs):
61
+ return self.context
62
+
63
+
64
+ async def _sentinel_extractor(page, n_results):
65
+ return SENTINEL
66
+
67
+
68
+ SCRAPERS = [
69
+ ("query_google_scholar", query_google_scholar,
70
+ "_extract_google_scholar_results", google_scholar_url),
71
+ ("query_google_patents", query_google_patents,
72
+ "_extract_google_patents_results", google_patents_search_url),
73
+ ("query_brave_search", query_brave_search,
74
+ "_extract_brave_results", brave_search_url),
75
+ ("query_bing_search", query_bing_search,
76
+ "_extract_bing_results", bing_search_url),
77
+ ]
78
+
79
+
80
+ @pytest.mark.parametrize("name, scraper, extractor_attr, builder",
81
+ SCRAPERS, ids=[s[0] for s in SCRAPERS])
82
+ async def test_scraper_navigates_to_the_url_its_builder_produced(
83
+ monkeypatch, name, scraper, extractor_attr, builder):
84
+ monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
85
+ browser = _RecordingBrowser()
86
+
87
+ await scraper(browser, "agentic ai", 25)
88
+
89
+ assert browser.page.goto_urls == [builder("agentic ai", 25)]
90
+
91
+
92
+ @pytest.mark.parametrize("name, scraper, extractor_attr, builder",
93
+ SCRAPERS, ids=[s[0] for s in SCRAPERS])
94
+ async def test_scraper_returns_what_its_extractor_produced(
95
+ monkeypatch, name, scraper, extractor_attr, builder):
96
+ monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
97
+
98
+ results = await scraper(_RecordingBrowser(), "widgets", 10)
99
+
100
+ assert results == SENTINEL
101
+
102
+
103
+ @pytest.mark.parametrize("name, scraper, extractor_attr, builder",
104
+ SCRAPERS, ids=[s[0] for s in SCRAPERS])
105
+ async def test_scraper_blocks_decorative_resources(
106
+ monkeypatch, name, scraper, extractor_attr, builder):
107
+ """Skipping stylesheets and images meaningfully speeds up navigation,
108
+ and every scraper is supposed to do it."""
109
+ monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
110
+ browser = _RecordingBrowser()
111
+
112
+ await scraper(browser, "widgets", 10)
113
+
114
+ assert browser.page.route_patterns == ["**/*"]
115
+
116
+
117
+ @pytest.mark.parametrize("name, scraper, extractor_attr, builder",
118
+ SCRAPERS, ids=[s[0] for s in SCRAPERS])
119
+ async def test_scraper_closes_its_page_and_context(
120
+ monkeypatch, name, scraper, extractor_attr, builder):
121
+ monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
122
+ browser = _RecordingBrowser()
123
+
124
+ await scraper(browser, "widgets", 10)
125
+
126
+ assert browser.page.closed and browser.context.closed
127
+
128
+
129
+ async def test_page_and_context_are_closed_even_when_extraction_raises(monkeypatch):
130
+ async def exploding_extractor(page, n_results):
131
+ raise RuntimeError("selectors changed")
132
+
133
+ monkeypatch.setattr(serp, "_extract_bing_results", exploding_extractor)
134
+ browser = _RecordingBrowser()
135
+
136
+ with pytest.raises(RuntimeError):
137
+ await query_bing_search(browser, "widgets", 10)
138
+
139
+ assert browser.page.closed and browser.context.closed
140
+
141
+
142
+ async def test_context_is_closed_even_when_closing_the_page_raises():
143
+ """A crashed renderer can make page.close() throw; the context still has
144
+ to be released or it leaks for the process's lifetime."""
145
+ browser = _RecordingBrowser()
146
+
147
+ async def exploding_close():
148
+ raise RuntimeError("renderer gone")
149
+
150
+ browser.page.close = exploding_close
151
+
152
+ with pytest.raises(RuntimeError):
153
+ async with playwright_open_page(browser):
154
+ pass
155
+
156
+ assert browser.context.closed