Spaces:
Running
Move orchestration out of the route handlers into a service layer
Browse filesThe backend policy - which backend to try, in what order, when to fall
back, what counts as evidence a backend is unhealthy, when a patent is
absent rather than merely unreachable - was written inline in FastAPI route
handlers. It is not HTTP-specific: it would be equally true behind a CLI or
a queue consumer, and putting it there meant exercising any of it cost an
ASGI round-trip plus monkeypatching module globals.
It now lives in services.py as SearchService and PatentService, which take
their stateful collaborators (HTTP client, browser provider, circuit
breaker, OPS credentials) as constructor arguments. app.py keeps the
module-level instances and wires them in through `Depends`, so
`app.dependency_overrides` replaces them in a test.
The services raise domain errors - PatentNotFound, UpstreamUnavailable,
OPSUnconfigured - rather than HTTPException. Mapping those onto status
codes is the HTTP layer's job and is now three exception handlers instead
of the same branching repeated across two endpoints. app.py no longer
imports HTTPException at all.
What this buys, concretely:
- The autouse fixture that reset a process-wide circuit breaker between
tests is gone; each test builds a service with its own breaker.
- The pokes at ops token_manager._key / ._secret from two different test
files are gone; tests pass a small stand-in with a `configured` flag.
- 40 tests that needed an ASGI client are now plain unit tests. The
SearchService suite runs 22 tests in 0.21s.
- test_app.py is now about what app.py does: endpoint-to-service dispatch,
validation, dependency wiring, and the lifespan.
Also closes the coverage the refactor would otherwise have lost.
Splitting the scrapers left the four-line composition in each `query_*`
untested, so tests/test_serp_navigation.py covers it with a recording
stand-in browser - asserting each scraper navigates to the URL its own
builder produced. That is precisely the link the deleted goto-patching
harness could never test, since it accepted the URL and discarded it.
The full mutation battery from the review now stands at 18 of 18 killed,
against 6 of 13 when the review was written.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MSNSYnceqvdVz7Csis4K9e
- .coverage +0 -0
- CONTRIBUTING.md +29 -1
- app.py +101 -242
- services.py +337 -0
- tests/test_app.py +127 -406
- tests/test_auth.py +5 -4
- tests/test_input_bounds.py +3 -2
- tests/test_patent_fallback.py +124 -87
- tests/test_scrap.py +12 -0
- tests/test_search_service.py +390 -0
- tests/test_serp_navigation.py +156 -0
|
Binary file (53.2 kB). View file
|
|
|
|
@@ -30,7 +30,7 @@ ruff check . # same lint CI runs
|
|
| 30 |
pytest
|
| 31 |
```
|
| 32 |
|
| 33 |
-
|
| 34 |
`respx`; the Playwright-driven scrapers (Bing, Brave, Google Scholar,
|
| 35 |
Google Patents search) run against a real headless Chromium but navigate to
|
| 36 |
local fixture HTML instead of the live sites — see `tests/helpers.py` for
|
|
@@ -55,6 +55,34 @@ bug here — fix it the same way as anything else (see below): update the
|
|
| 55 |
fixture HTML to match the new real markup, confirm the test fails against
|
| 56 |
the current selectors, fix the selectors, confirm green.
|
| 57 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
## The development loop (TDD, not test-after)
|
| 59 |
|
| 60 |
Every change in this codebase's history followed the same loop, and new
|
|
|
|
| 30 |
pytest
|
| 31 |
```
|
| 32 |
|
| 33 |
+
238 tests, a few seconds, fully offline. All outbound HTTP is mocked with
|
| 34 |
`respx`; the Playwright-driven scrapers (Bing, Brave, Google Scholar,
|
| 35 |
Google Patents search) run against a real headless Chromium but navigate to
|
| 36 |
local fixture HTML instead of the live sites — see `tests/helpers.py` for
|
|
|
|
| 55 |
fixture HTML to match the new real markup, confirm the test fails against
|
| 56 |
the current selectors, fix the selectors, confirm green.
|
| 57 |
|
| 58 |
+
## How the code is laid out
|
| 59 |
+
|
| 60 |
+
- **`app.py`** is HTTP wiring only: request validation, dependency
|
| 61 |
+
providers, the lifespan, and mapping domain errors onto status codes.
|
| 62 |
+
- **`services.py`** holds the orchestration policy - which backend to try,
|
| 63 |
+
in what order, when to fall back, what counts as evidence a backend is
|
| 64 |
+
unhealthy, and when a patent is genuinely absent rather than merely
|
| 65 |
+
unreachable. None of that is HTTP-specific, and keeping it out of route
|
| 66 |
+
handlers is why most of its tests need no ASGI client.
|
| 67 |
+
- **`serp.py` / `scrap.py` / `ops.py`** are the backends. Each scraper is
|
| 68 |
+
split into a pure URL builder, a navigation step, and an `_extract_*`
|
| 69 |
+
function taking an already-loaded page, so all three can be tested
|
| 70 |
+
separately.
|
| 71 |
+
- **`utils.py` / `circuit_breaker.py`** are dependency-free helpers.
|
| 72 |
+
|
| 73 |
+
The services take their stateful collaborators (HTTP client, browser,
|
| 74 |
+
circuit breaker, OPS credentials) as constructor arguments. That is
|
| 75 |
+
deliberate and worth preserving: it is what lets a test build a service
|
| 76 |
+
with its own circuit breaker instead of resetting a process-wide singleton
|
| 77 |
+
between tests, and hand in a stand-in for the OPS credentials instead of
|
| 78 |
+
assigning to private attributes on the real token manager. `app.py`'s
|
| 79 |
+
`get_search_service` / `get_patent_service` providers exist so
|
| 80 |
+
`app.dependency_overrides` can replace them wholesale in a test.
|
| 81 |
+
|
| 82 |
+
The stateless backend functions stay module-level imports in
|
| 83 |
+
`services.py`; monkeypatching one of those in a test is fine, because they
|
| 84 |
+
hold no state to leak between tests.
|
| 85 |
+
|
| 86 |
## The development loop (TDD, not test-after)
|
| 87 |
|
| 88 |
Every change in this codebase's history followed the same loop, and new
|
|
@@ -1,24 +1,24 @@
|
|
| 1 |
-
import asyncio
|
| 2 |
import logging
|
| 3 |
import os
|
| 4 |
import secrets
|
| 5 |
from contextlib import asynccontextmanager
|
| 6 |
from typing import Annotated, Optional
|
| 7 |
-
|
|
|
|
|
|
|
|
|
|
| 8 |
from fastapi.responses import JSONResponse
|
| 9 |
from fastapi.routing import APIRouter
|
| 10 |
-
import
|
| 11 |
-
from httpx import HTTPStatusError
|
| 12 |
from pydantic import BaseModel, Field, StringConstraints
|
| 13 |
-
from playwright.async_api import async_playwright, Browser
|
| 14 |
-
import uvicorn
|
| 15 |
|
| 16 |
-
from circuit_breaker import CircuitBreaker
|
| 17 |
-
from scrap import PatentScrapBulkResponse, PatentScrapResult, scrap_patent_async, scrap_patent_bulk_async
|
| 18 |
-
from serp import EmptyResultsError, PATENT_ID_CORE, SerpQuery, SerpResults, query_arxiv, query_bing_search, query_brave_search, query_ddg_search, query_google_patents, query_google_scholar
|
| 19 |
-
from ops import OPSBulkResponse, OPSNotConfigured, ops_scrap_patent, ops_scrap_patent_bulk, ops_search, token_manager as ops_token_manager
|
| 20 |
-
from utils import log_gathered_exceptions
|
| 21 |
from mcp_server import mount_mcp_server
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
|
| 23 |
# Anchored version of serp.py's PATENT_ID_CORE: validates a whole
|
| 24 |
# user-supplied patent id rather than finding one inside free text. Rejects
|
|
@@ -26,6 +26,11 @@ from mcp_server import mount_mcp_server
|
|
| 26 |
# instead of it reaching Google Patents/OPS as a confusing request.
|
| 27 |
PatentId = Annotated[str, StringConstraints(pattern=rf"^{PATENT_ID_CORE}$")]
|
| 28 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
logging.basicConfig(
|
| 30 |
level=logging.INFO,
|
| 31 |
format='[%(asctime)s][%(levelname)s][%(filename)s:%(lineno)d]: %(message)s',
|
|
@@ -36,12 +41,19 @@ logging.basicConfig(
|
|
| 36 |
playwright = None
|
| 37 |
pw_browser: Optional[Browser] = None
|
| 38 |
|
| 39 |
-
# httpx client. Constructed at import so
|
| 40 |
-
# over it, but its lifetime is owned by `api_lifespan`, which closes
|
| 41 |
-
# shutdown alongside the browser.
|
| 42 |
httpx_client = httpx.AsyncClient(timeout=30, limits=httpx.Limits(
|
| 43 |
max_connections=30, max_keepalive_connections=20))
|
| 44 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
# ===================== Optional API key protection =====================
|
| 46 |
# Unset by default, which keeps the API exactly as open as before. Set
|
| 47 |
# SERPENT_API_KEY on the deployment to require every request (REST and MCP
|
|
@@ -97,6 +109,32 @@ app = FastAPI(lifespan=api_lifespan, docs_url="/",
|
|
| 97 |
title="SERPent", description=_load_docs())
|
| 98 |
|
| 99 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 100 |
@app.middleware("http")
|
| 101 |
async def api_key_guard(request: Request, call_next):
|
| 102 |
"""No-op unless SERPENT_API_KEY is set; then gates every path but the docs."""
|
|
@@ -114,261 +152,101 @@ async def api_key_guard(request: Request, call_next):
|
|
| 114 |
)
|
| 115 |
return await call_next(request)
|
| 116 |
|
| 117 |
-
# Router for scrapping related endpoints
|
| 118 |
-
scrap_router = APIRouter(prefix="/scrap", tags=["scrapping"])
|
| 119 |
-
# Router for SERP-scrapping related endpoints
|
| 120 |
-
serp_router = APIRouter(prefix="/serp", tags=["serp scrapping"])
|
| 121 |
-
# Router for EPO OPS (official patent API) endpoints
|
| 122 |
-
ops_router = APIRouter(prefix="/ops", tags=["EPO OPS"])
|
| 123 |
|
| 124 |
-
# =====================
|
|
|
|
|
|
|
|
|
|
|
|
|
| 125 |
|
| 126 |
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
last error only when every query failed.
|
| 131 |
-
"""
|
| 132 |
-
if not results:
|
| 133 |
-
# SerpQuery bounds the query list, so this is unreachable from the
|
| 134 |
-
# endpoints; keep the helper total anyway rather than indexing into
|
| 135 |
-
# an empty list, since every search endpoint shares it.
|
| 136 |
-
return SerpResults(results=[], error="No queries were provided.")
|
| 137 |
|
| 138 |
-
filtered_results = [r for r in results if not isinstance(r, Exception)]
|
| 139 |
-
flattened_results = [
|
| 140 |
-
item for sublist in filtered_results for item in sublist]
|
| 141 |
|
| 142 |
-
|
| 143 |
-
|
|
|
|
| 144 |
|
| 145 |
-
return SerpResults(results=flattened_results, error=None)
|
| 146 |
|
|
|
|
|
|
|
|
|
|
| 147 |
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
|
| 158 |
|
| 159 |
@serp_router.post("/search_scholar")
|
| 160 |
-
async def search_google_scholar(params: SerpQuery) -> SerpResults:
|
| 161 |
"""Queries google scholar for the specified query"""
|
| 162 |
logging.info(f"Searching Google Scholar for queries: {params.queries}")
|
| 163 |
-
return await
|
| 164 |
-
lambda q, n: query_google_scholar(pw_browser, q, n), params, "google scholar search")
|
| 165 |
|
| 166 |
|
| 167 |
@serp_router.post("/search_arxiv")
|
| 168 |
-
async def search_arxiv(params: SerpQuery) -> SerpResults:
|
| 169 |
"""Searches arxiv for the specified queries and returns the found documents."""
|
| 170 |
logging.info(f"Searching Arxiv for queries: {params.queries}")
|
| 171 |
-
return await
|
| 172 |
-
lambda q, n: query_arxiv(httpx_client, q, n), params, "arxiv search")
|
| 173 |
|
| 174 |
|
| 175 |
@serp_router.post("/search_patents")
|
| 176 |
-
async def search_patents(params: SerpQuery) -> SerpResults:
|
| 177 |
"""Searches google patents for the specified queries and returns the found documents.
|
| 178 |
|
| 179 |
Falls back to the EPO OPS API for any query Google Patents returns nothing
|
| 180 |
for, when OPS credentials are configured.
|
| 181 |
"""
|
| 182 |
logging.info(f"Searching Google Patents for queries: {params.queries}")
|
| 183 |
-
|
| 184 |
-
log_gathered_exceptions(results, "google patent search", params.queries)
|
| 185 |
-
|
| 186 |
-
# Fall back to OPS for queries that errored or returned no results.
|
| 187 |
-
# Gathered rather than awaited one at a time: this path exists for when
|
| 188 |
-
# Google Patents is unavailable, so it is exactly when a request is
|
| 189 |
-
# least able to afford N sequential round-trips to the slower backend.
|
| 190 |
-
if ops_token_manager.configured:
|
| 191 |
-
needs_fallback = [
|
| 192 |
-
i for i, res in enumerate(results)
|
| 193 |
-
if isinstance(res, Exception) or not res]
|
| 194 |
-
if needs_fallback:
|
| 195 |
-
logging.info(
|
| 196 |
-
f"Google Patents empty for {len(needs_fallback)} quer(y/ies), trying OPS.")
|
| 197 |
-
fallback_results = await asyncio.gather(
|
| 198 |
-
*[ops_search(httpx_client, params.queries[i], params.n_results)
|
| 199 |
-
for i in needs_fallback],
|
| 200 |
-
return_exceptions=True)
|
| 201 |
-
# strict: gather returns exactly one result per index, so a
|
| 202 |
-
# length mismatch here would be a bug worth surfacing.
|
| 203 |
-
for i, res in zip(needs_fallback, fallback_results, strict=True):
|
| 204 |
-
if isinstance(res, Exception):
|
| 205 |
-
logging.warning(f"OPS fallback failed for `{params.queries[i]}`: {res}")
|
| 206 |
-
else:
|
| 207 |
-
results[i] = res
|
| 208 |
-
|
| 209 |
-
return _shape_serp_results(results)
|
| 210 |
|
| 211 |
|
| 212 |
@serp_router.post("/search_brave")
|
| 213 |
-
async def search_brave(params: SerpQuery) -> SerpResults:
|
| 214 |
"""Searches brave search for the specified queries and returns the found documents."""
|
| 215 |
logging.info(f"Searching Brave Search for queries: {params.queries}")
|
| 216 |
-
return await
|
| 217 |
-
lambda q, n: query_brave_search(pw_browser, q, n), params, "brave search")
|
| 218 |
|
| 219 |
|
| 220 |
@serp_router.post("/search_bing")
|
| 221 |
-
async def search_bing(params: SerpQuery) -> SerpResults:
|
| 222 |
"""Searches Bing search for the specified queries and returns the found documents."""
|
| 223 |
logging.info(f"Searching Bing Search for queries: {params.queries}")
|
| 224 |
-
return await
|
| 225 |
-
lambda q, n: query_bing_search(pw_browser, q, n), params, "bing search")
|
| 226 |
|
| 227 |
|
| 228 |
@serp_router.post("/search_duck")
|
| 229 |
-
async def search_duck(params: SerpQuery) -> SerpResults:
|
| 230 |
"""Searches duckduckgo for the specified queries and returns the found documents"""
|
| 231 |
logging.info(f"Searching DuckDuckGo for queries: {params.queries}")
|
| 232 |
-
return await
|
| 233 |
-
lambda q, n: query_ddg_search(q, n), params, "duckduckgo search")
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
# Shared across all queries and requests: once a backend has failed
|
| 237 |
-
# `failure_threshold` times in a row, skip it for `cooldown_seconds` instead
|
| 238 |
-
# of attempting (and paying the timeout cost of) another call that's very
|
| 239 |
-
# likely to fail - and, more importantly, stop hammering a backend that may
|
| 240 |
-
# already be rate-limiting or blocking this deployment's IP.
|
| 241 |
-
_backend_circuit_breaker = CircuitBreaker(failure_threshold=3, cooldown_seconds=60.0)
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
async def _search_one(q: str, n_results: int) -> tuple[str, list[dict], Optional[str]]:
|
| 245 |
-
"""Try DDG, then Brave, then Bing for a single query; stop at the first success."""
|
| 246 |
-
backends = [
|
| 247 |
-
("DuckDuckGo", lambda: query_ddg_search(q, n_results)),
|
| 248 |
-
("Brave Search", lambda: query_brave_search(pw_browser, q, n_results)),
|
| 249 |
-
("Bing", lambda: query_bing_search(pw_browser, q, n_results)),
|
| 250 |
-
]
|
| 251 |
-
last_error: Optional[Exception] = None
|
| 252 |
-
for name, call in backends:
|
| 253 |
-
try:
|
| 254 |
-
_backend_circuit_breaker.before_call(name)
|
| 255 |
-
except CircuitOpenError as e:
|
| 256 |
-
logging.info(f"Skipping {name} for query `{q}`: {e}")
|
| 257 |
-
last_error = e
|
| 258 |
-
continue
|
| 259 |
-
|
| 260 |
-
try:
|
| 261 |
-
logging.info(f"Querying {name} with query: `{q}`")
|
| 262 |
-
result = await call()
|
| 263 |
-
_backend_circuit_breaker.record_success(name)
|
| 264 |
-
return q, result, None
|
| 265 |
-
except EmptyResultsError as e:
|
| 266 |
-
# Zero results is ambiguous - no hits, or a soft block that
|
| 267 |
-
# parsed to nothing. Move on to the next backend, but don't
|
| 268 |
-
# hold it against this one: the breaker is process-wide, so
|
| 269 |
-
# counting it would let a few unusual-but-legitimate queries
|
| 270 |
-
# disable the primary backend for every user. A hard block
|
| 271 |
-
# normally raises a real error, which is handled below.
|
| 272 |
-
logging.info(f"{name} returned no results for query `{q}`")
|
| 273 |
-
last_error = e
|
| 274 |
-
except Exception as e:
|
| 275 |
-
logging.error(f"Failed to query {name} with query `{q}`: {e}")
|
| 276 |
-
_backend_circuit_breaker.record_failure(name)
|
| 277 |
-
last_error = e
|
| 278 |
-
|
| 279 |
-
return q, [], f"All backends failed for query '{q}': {last_error}"
|
| 280 |
|
| 281 |
|
| 282 |
@serp_router.post("/search")
|
| 283 |
-
async def search(params: SerpQuery) -> SerpResults:
|
| 284 |
"""Attempts to search the specified queries using ALL backends"""
|
| 285 |
-
|
| 286 |
-
|
| 287 |
-
results: list[dict] = []
|
| 288 |
-
errors: list[str] = []
|
| 289 |
-
for _q, res, err in outcomes:
|
| 290 |
-
results.extend(res)
|
| 291 |
-
if err:
|
| 292 |
-
errors.append(err)
|
| 293 |
-
|
| 294 |
-
if len(results) == 0:
|
| 295 |
-
return SerpResults(results=[], error="; ".join(errors) if errors else "All backends are rate-limited.")
|
| 296 |
-
|
| 297 |
-
return SerpResults(results=results, error="; ".join(errors) if errors else None)
|
| 298 |
|
| 299 |
# =========================== Scrapping endpoints ===========================
|
| 300 |
|
| 301 |
|
| 302 |
-
def _is_not_found(exc: Exception) -> bool:
|
| 303 |
-
"""True only when a backend positively said the document is absent.
|
| 304 |
-
|
| 305 |
-
Anything else - a 5xx, a rate-limit, a timeout, an unparseable page -
|
| 306 |
-
is a failure to answer, not an answer.
|
| 307 |
-
"""
|
| 308 |
-
return isinstance(exc, HTTPStatusError) and exc.response.status_code == 404
|
| 309 |
-
|
| 310 |
-
|
| 311 |
-
def _upstream_status(exc: Exception) -> int:
|
| 312 |
-
"""The status to report for a failed upstream call. Never 404.
|
| 313 |
-
|
| 314 |
-
404 is a claim that the patent does not exist, and MCP_INSTRUCTIONS
|
| 315 |
-
tells agents to believe it and move on without retrying - so it has to
|
| 316 |
-
be reserved for a backend actually saying so. A transient outage
|
| 317 |
-
reported as 404 teaches an agent that a real patent is absent.
|
| 318 |
-
"""
|
| 319 |
-
return 504 if isinstance(exc, httpx.TimeoutException) else 502
|
| 320 |
-
|
| 321 |
-
|
| 322 |
@scrap_router.get("/scrap_patent/{patent_id}")
|
| 323 |
-
async def scrap_patent(patent_id: PatentId) -> PatentScrapResult:
|
| 324 |
"""Scraps the specified patent from Google Patents.
|
| 325 |
|
| 326 |
Falls back to the EPO OPS API (which covers patents missing from Google
|
| 327 |
Patents) when the scrape fails and OPS credentials are configured.
|
| 328 |
"""
|
| 329 |
-
|
| 330 |
-
try:
|
| 331 |
-
return await scrap_patent_async(httpx_client, f"https://patents.google.com/patent/{patent_id}/en")
|
| 332 |
-
except HTTPStatusError as e:
|
| 333 |
-
google_error = e
|
| 334 |
-
logging.warning(
|
| 335 |
-
f"Google Patents returned {e.response.status_code} for {patent_id}.")
|
| 336 |
-
except Exception as e:
|
| 337 |
-
google_error = e
|
| 338 |
-
logging.warning(f"Failed to scrap patent {patent_id}: {e}")
|
| 339 |
-
|
| 340 |
-
if not ops_token_manager.configured:
|
| 341 |
-
if _is_not_found(google_error):
|
| 342 |
-
raise HTTPException(
|
| 343 |
-
status_code=404,
|
| 344 |
-
detail=f"Patent '{patent_id}' not found on Google Patents (EPO OPS fallback not configured).")
|
| 345 |
-
raise HTTPException(
|
| 346 |
-
status_code=_upstream_status(google_error),
|
| 347 |
-
detail=f"Google Patents request failed for '{patent_id}' and the EPO OPS fallback is not configured: {google_error}")
|
| 348 |
-
|
| 349 |
-
try:
|
| 350 |
-
logging.info(f"Trying OPS for patent {patent_id}.")
|
| 351 |
-
return await ops_scrap_patent(httpx_client, patent_id)
|
| 352 |
-
except Exception as e:
|
| 353 |
-
logging.warning(f"OPS fallback failed for {patent_id}: {e}")
|
| 354 |
-
# Only claim the patent doesn't exist when *both* backends said so.
|
| 355 |
-
if _is_not_found(google_error) and _is_not_found(e):
|
| 356 |
-
raise HTTPException(
|
| 357 |
-
status_code=404,
|
| 358 |
-
detail=f"Patent '{patent_id}' not found on Google Patents or EPO OPS.") from e
|
| 359 |
-
if isinstance(e, HTTPStatusError):
|
| 360 |
-
raise HTTPException(
|
| 361 |
-
status_code=_upstream_status(e),
|
| 362 |
-
detail=f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
|
| 363 |
-
raise HTTPException(
|
| 364 |
-
status_code=_upstream_status(e),
|
| 365 |
-
detail=f"EPO OPS request failed for '{patent_id}': {e}") from e
|
| 366 |
-
|
| 367 |
-
|
| 368 |
-
# Same reasoning as MAX_QUERIES_PER_REQUEST in serp.py: one accepted
|
| 369 |
-
# request must not be able to queue unbounded outbound work. The OPS
|
| 370 |
-
# variant is the expensive one - three requests per id.
|
| 371 |
-
MAX_BULK_PATENT_IDS = 50
|
| 372 |
|
| 373 |
|
| 374 |
class ScrapPatentsRequest(BaseModel):
|
|
@@ -380,51 +258,32 @@ class ScrapPatentsRequest(BaseModel):
|
|
| 380 |
|
| 381 |
|
| 382 |
@scrap_router.post("/scrap_patents_bulk", response_model=PatentScrapBulkResponse)
|
| 383 |
-
async def scrap_patents(params: ScrapPatentsRequest
|
|
|
|
| 384 |
"""Scraps multiple patents from Google Patents."""
|
| 385 |
-
|
| 386 |
-
return patents
|
| 387 |
|
| 388 |
# =========================== EPO OPS endpoints ===========================
|
| 389 |
|
| 390 |
|
| 391 |
@ops_router.post("/search")
|
| 392 |
-
async def ops_keyword_search(params: SerpQuery) -> SerpResults:
|
| 393 |
"""Keyword-searches patents via the official EPO OPS API."""
|
| 394 |
logging.info(f"Searching EPO OPS for queries: {params.queries}")
|
| 395 |
-
return await
|
| 396 |
-
lambda q, n: ops_search(httpx_client, q, n), params, "OPS search")
|
| 397 |
|
| 398 |
|
| 399 |
@ops_router.get("/scrap_patent/{patent_id}")
|
| 400 |
-
async def ops_get_patent(patent_id: PatentId) -> PatentScrapResult:
|
| 401 |
"""Retrieves a patent (biblio + abstract + claims + description) via EPO OPS."""
|
| 402 |
-
|
| 403 |
-
raise HTTPException(
|
| 404 |
-
status_code=503,
|
| 405 |
-
detail="EPO OPS is not configured (OPS_CONSUMER_KEY / OPS_CONSUMER_SECRET missing).")
|
| 406 |
-
try:
|
| 407 |
-
return await ops_scrap_patent(httpx_client, patent_id)
|
| 408 |
-
except OPSNotConfigured as e:
|
| 409 |
-
raise HTTPException(status_code=503, detail="EPO OPS is not configured.") from e
|
| 410 |
-
except HTTPStatusError as e:
|
| 411 |
-
if e.response.status_code == 404:
|
| 412 |
-
raise HTTPException(
|
| 413 |
-
status_code=404, detail=f"Patent '{patent_id}' not found in EPO OPS.") from e
|
| 414 |
-
raise HTTPException(
|
| 415 |
-
status_code=_upstream_status(e),
|
| 416 |
-
detail=f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
|
| 417 |
-
except Exception as e:
|
| 418 |
-
logging.warning(f"Failed to retrieve patent {patent_id} from OPS: {e}")
|
| 419 |
-
raise HTTPException(
|
| 420 |
-
status_code=_upstream_status(e),
|
| 421 |
-
detail=f"EPO OPS request failed for '{patent_id}': {e}") from e
|
| 422 |
|
| 423 |
|
| 424 |
@ops_router.post("/scrap_patents_bulk", response_model=OPSBulkResponse)
|
| 425 |
-
async def ops_get_patents_bulk(params: ScrapPatentsRequest
|
|
|
|
| 426 |
"""Retrieves multiple patents via EPO OPS."""
|
| 427 |
-
return await
|
| 428 |
|
| 429 |
# ===============================================================================
|
| 430 |
|
|
|
|
|
|
|
| 1 |
import logging
|
| 2 |
import os
|
| 3 |
import secrets
|
| 4 |
from contextlib import asynccontextmanager
|
| 5 |
from typing import Annotated, Optional
|
| 6 |
+
|
| 7 |
+
import httpx
|
| 8 |
+
import uvicorn
|
| 9 |
+
from fastapi import Depends, FastAPI, Request
|
| 10 |
from fastapi.responses import JSONResponse
|
| 11 |
from fastapi.routing import APIRouter
|
| 12 |
+
from playwright.async_api import Browser, async_playwright
|
|
|
|
| 13 |
from pydantic import BaseModel, Field, StringConstraints
|
|
|
|
|
|
|
| 14 |
|
| 15 |
+
from circuit_breaker import CircuitBreaker
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
from mcp_server import mount_mcp_server
|
| 17 |
+
from ops import OPSBulkResponse, token_manager as ops_token_manager
|
| 18 |
+
from scrap import PatentScrapBulkResponse, PatentScrapResult
|
| 19 |
+
from serp import PATENT_ID_CORE, SerpQuery, SerpResults
|
| 20 |
+
from services import (OPSUnconfigured, PatentNotFound, PatentService,
|
| 21 |
+
SearchService, UpstreamUnavailable)
|
| 22 |
|
| 23 |
# Anchored version of serp.py's PATENT_ID_CORE: validates a whole
|
| 24 |
# user-supplied patent id rather than finding one inside free text. Rejects
|
|
|
|
| 26 |
# instead of it reaching Google Patents/OPS as a confusing request.
|
| 27 |
PatentId = Annotated[str, StringConstraints(pattern=rf"^{PATENT_ID_CORE}$")]
|
| 28 |
|
| 29 |
+
# Same reasoning as MAX_QUERIES_PER_REQUEST in serp.py: one accepted
|
| 30 |
+
# request must not be able to queue unbounded outbound work. The OPS
|
| 31 |
+
# variant is the expensive one - three requests per id.
|
| 32 |
+
MAX_BULK_PATENT_IDS = 50
|
| 33 |
+
|
| 34 |
logging.basicConfig(
|
| 35 |
level=logging.INFO,
|
| 36 |
format='[%(asctime)s][%(levelname)s][%(filename)s:%(lineno)d]: %(message)s',
|
|
|
|
| 41 |
playwright = None
|
| 42 |
pw_browser: Optional[Browser] = None
|
| 43 |
|
| 44 |
+
# httpx client. Constructed at import so the dependency providers below can
|
| 45 |
+
# close over it, but its lifetime is owned by `api_lifespan`, which closes
|
| 46 |
+
# it on shutdown alongside the browser.
|
| 47 |
httpx_client = httpx.AsyncClient(timeout=30, limits=httpx.Limits(
|
| 48 |
max_connections=30, max_keepalive_connections=20))
|
| 49 |
|
| 50 |
+
# Shared across all queries and requests: once a backend has failed
|
| 51 |
+
# `failure_threshold` times in a row, skip it for `cooldown_seconds` instead
|
| 52 |
+
# of attempting (and paying the timeout cost of) another call that's very
|
| 53 |
+
# likely to fail - and, more importantly, stop hammering a backend that may
|
| 54 |
+
# already be rate-limiting or blocking this deployment's IP.
|
| 55 |
+
_backend_circuit_breaker = CircuitBreaker(failure_threshold=3, cooldown_seconds=60.0)
|
| 56 |
+
|
| 57 |
# ===================== Optional API key protection =====================
|
| 58 |
# Unset by default, which keeps the API exactly as open as before. Set
|
| 59 |
# SERPENT_API_KEY on the deployment to require every request (REST and MCP
|
|
|
|
| 109 |
title="SERPent", description=_load_docs())
|
| 110 |
|
| 111 |
|
| 112 |
+
# ============================ service dependencies ============================
|
| 113 |
+
# The orchestration itself lives in services.py; these just wire the app's
|
| 114 |
+
# collaborators into it. Overriding one of these via
|
| 115 |
+
# `app.dependency_overrides` is how a test supplies its own circuit breaker
|
| 116 |
+
# or OPS credentials, rather than reassigning module globals.
|
| 117 |
+
|
| 118 |
+
|
| 119 |
+
def get_search_service() -> SearchService:
|
| 120 |
+
return SearchService(
|
| 121 |
+
http_client=httpx_client,
|
| 122 |
+
# A provider, not the browser itself: the lifespan starts it after
|
| 123 |
+
# import, and leaves it None when Playwright fails to start.
|
| 124 |
+
browser_provider=lambda: pw_browser,
|
| 125 |
+
circuit_breaker=_backend_circuit_breaker,
|
| 126 |
+
ops_tokens=ops_token_manager,
|
| 127 |
+
)
|
| 128 |
+
|
| 129 |
+
|
| 130 |
+
def get_patent_service() -> PatentService:
|
| 131 |
+
return PatentService(http_client=httpx_client, ops_tokens=ops_token_manager)
|
| 132 |
+
|
| 133 |
+
|
| 134 |
+
SearchServiceDep = Annotated[SearchService, Depends(get_search_service)]
|
| 135 |
+
PatentServiceDep = Annotated[PatentService, Depends(get_patent_service)]
|
| 136 |
+
|
| 137 |
+
|
| 138 |
@app.middleware("http")
|
| 139 |
async def api_key_guard(request: Request, call_next):
|
| 140 |
"""No-op unless SERPENT_API_KEY is set; then gates every path but the docs."""
|
|
|
|
| 152 |
)
|
| 153 |
return await call_next(request)
|
| 154 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 155 |
|
| 156 |
+
# ========================= domain errors -> status codes =========================
|
| 157 |
+
# 404 is a claim that the patent does not exist, and MCP_INSTRUCTIONS tells
|
| 158 |
+
# agents to believe it and move on without retrying - so the services raise
|
| 159 |
+
# PatentNotFound only when a backend positively said so, and everything else
|
| 160 |
+
# arrives here as UpstreamUnavailable.
|
| 161 |
|
| 162 |
|
| 163 |
+
@app.exception_handler(PatentNotFound)
|
| 164 |
+
async def _patent_not_found_handler(request: Request, exc: PatentNotFound):
|
| 165 |
+
return JSONResponse({"detail": str(exc)}, status_code=404)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 166 |
|
|
|
|
|
|
|
|
|
|
| 167 |
|
| 168 |
+
@app.exception_handler(UpstreamUnavailable)
|
| 169 |
+
async def _upstream_unavailable_handler(request: Request, exc: UpstreamUnavailable):
|
| 170 |
+
return JSONResponse({"detail": str(exc)}, status_code=504 if exc.timeout else 502)
|
| 171 |
|
|
|
|
| 172 |
|
| 173 |
+
@app.exception_handler(OPSUnconfigured)
|
| 174 |
+
async def _ops_unconfigured_handler(request: Request, exc: OPSUnconfigured):
|
| 175 |
+
return JSONResponse({"detail": str(exc)}, status_code=503)
|
| 176 |
|
| 177 |
+
|
| 178 |
+
# Router for scrapping related endpoints
|
| 179 |
+
scrap_router = APIRouter(prefix="/scrap", tags=["scrapping"])
|
| 180 |
+
# Router for SERP-scrapping related endpoints
|
| 181 |
+
serp_router = APIRouter(prefix="/serp", tags=["serp scrapping"])
|
| 182 |
+
# Router for EPO OPS (official patent API) endpoints
|
| 183 |
+
ops_router = APIRouter(prefix="/ops", tags=["EPO OPS"])
|
| 184 |
+
|
| 185 |
+
# ===================== Search endpoints =====================
|
| 186 |
|
| 187 |
|
| 188 |
@serp_router.post("/search_scholar")
|
| 189 |
+
async def search_google_scholar(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 190 |
"""Queries google scholar for the specified query"""
|
| 191 |
logging.info(f"Searching Google Scholar for queries: {params.queries}")
|
| 192 |
+
return await service.google_scholar(params)
|
|
|
|
| 193 |
|
| 194 |
|
| 195 |
@serp_router.post("/search_arxiv")
|
| 196 |
+
async def search_arxiv(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 197 |
"""Searches arxiv for the specified queries and returns the found documents."""
|
| 198 |
logging.info(f"Searching Arxiv for queries: {params.queries}")
|
| 199 |
+
return await service.arxiv(params)
|
|
|
|
| 200 |
|
| 201 |
|
| 202 |
@serp_router.post("/search_patents")
|
| 203 |
+
async def search_patents(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 204 |
"""Searches google patents for the specified queries and returns the found documents.
|
| 205 |
|
| 206 |
Falls back to the EPO OPS API for any query Google Patents returns nothing
|
| 207 |
for, when OPS credentials are configured.
|
| 208 |
"""
|
| 209 |
logging.info(f"Searching Google Patents for queries: {params.queries}")
|
| 210 |
+
return await service.patents(params)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 211 |
|
| 212 |
|
| 213 |
@serp_router.post("/search_brave")
|
| 214 |
+
async def search_brave(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 215 |
"""Searches brave search for the specified queries and returns the found documents."""
|
| 216 |
logging.info(f"Searching Brave Search for queries: {params.queries}")
|
| 217 |
+
return await service.brave(params)
|
|
|
|
| 218 |
|
| 219 |
|
| 220 |
@serp_router.post("/search_bing")
|
| 221 |
+
async def search_bing(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 222 |
"""Searches Bing search for the specified queries and returns the found documents."""
|
| 223 |
logging.info(f"Searching Bing Search for queries: {params.queries}")
|
| 224 |
+
return await service.bing(params)
|
|
|
|
| 225 |
|
| 226 |
|
| 227 |
@serp_router.post("/search_duck")
|
| 228 |
+
async def search_duck(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 229 |
"""Searches duckduckgo for the specified queries and returns the found documents"""
|
| 230 |
logging.info(f"Searching DuckDuckGo for queries: {params.queries}")
|
| 231 |
+
return await service.duckduckgo(params)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 232 |
|
| 233 |
|
| 234 |
@serp_router.post("/search")
|
| 235 |
+
async def search(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 236 |
"""Attempts to search the specified queries using ALL backends"""
|
| 237 |
+
return await service.search(params)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 238 |
|
| 239 |
# =========================== Scrapping endpoints ===========================
|
| 240 |
|
| 241 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 242 |
@scrap_router.get("/scrap_patent/{patent_id}")
|
| 243 |
+
async def scrap_patent(patent_id: PatentId, service: PatentServiceDep) -> PatentScrapResult:
|
| 244 |
"""Scraps the specified patent from Google Patents.
|
| 245 |
|
| 246 |
Falls back to the EPO OPS API (which covers patents missing from Google
|
| 247 |
Patents) when the scrape fails and OPS credentials are configured.
|
| 248 |
"""
|
| 249 |
+
return await service.scrap(patent_id)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 250 |
|
| 251 |
|
| 252 |
class ScrapPatentsRequest(BaseModel):
|
|
|
|
| 258 |
|
| 259 |
|
| 260 |
@scrap_router.post("/scrap_patents_bulk", response_model=PatentScrapBulkResponse)
|
| 261 |
+
async def scrap_patents(params: ScrapPatentsRequest,
|
| 262 |
+
service: PatentServiceDep) -> PatentScrapBulkResponse:
|
| 263 |
"""Scraps multiple patents from Google Patents."""
|
| 264 |
+
return await service.scrap_bulk(params.patent_ids)
|
|
|
|
| 265 |
|
| 266 |
# =========================== EPO OPS endpoints ===========================
|
| 267 |
|
| 268 |
|
| 269 |
@ops_router.post("/search")
|
| 270 |
+
async def ops_keyword_search(params: SerpQuery, service: SearchServiceDep) -> SerpResults:
|
| 271 |
"""Keyword-searches patents via the official EPO OPS API."""
|
| 272 |
logging.info(f"Searching EPO OPS for queries: {params.queries}")
|
| 273 |
+
return await service.ops_keyword_search(params)
|
|
|
|
| 274 |
|
| 275 |
|
| 276 |
@ops_router.get("/scrap_patent/{patent_id}")
|
| 277 |
+
async def ops_get_patent(patent_id: PatentId, service: PatentServiceDep) -> PatentScrapResult:
|
| 278 |
"""Retrieves a patent (biblio + abstract + claims + description) via EPO OPS."""
|
| 279 |
+
return await service.ops_scrap(patent_id)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 280 |
|
| 281 |
|
| 282 |
@ops_router.post("/scrap_patents_bulk", response_model=OPSBulkResponse)
|
| 283 |
+
async def ops_get_patents_bulk(params: ScrapPatentsRequest,
|
| 284 |
+
service: PatentServiceDep) -> OPSBulkResponse:
|
| 285 |
"""Retrieves multiple patents via EPO OPS."""
|
| 286 |
+
return await service.ops_scrap_bulk(params.patent_ids)
|
| 287 |
|
| 288 |
# ===============================================================================
|
| 289 |
|
|
@@ -0,0 +1,337 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Backend orchestration, separated from HTTP wiring.
|
| 2 |
+
|
| 3 |
+
The policy in here - which backend to try, in what order, when to fall back
|
| 4 |
+
to another, what counts as evidence a backend is unhealthy, and when a
|
| 5 |
+
patent is genuinely absent rather than merely unreachable - is domain
|
| 6 |
+
policy. It would be equally true behind a CLI or a queue consumer, and it
|
| 7 |
+
was previously written inline in FastAPI route handlers, which meant
|
| 8 |
+
exercising any of it cost an ASGI round-trip.
|
| 9 |
+
|
| 10 |
+
The stateful collaborators (HTTP client, browser, circuit breaker, OPS
|
| 11 |
+
credentials) are constructor arguments rather than module globals reached
|
| 12 |
+
into at call time. That is what lets a test build a service with its own
|
| 13 |
+
circuit breaker instead of resetting a process-wide singleton between
|
| 14 |
+
tests, and hand in a stand-in for the OPS credentials instead of assigning
|
| 15 |
+
to private attributes on the real token manager.
|
| 16 |
+
|
| 17 |
+
The stateless backend functions (`query_arxiv` and friends) stay
|
| 18 |
+
module-level imports: they hold no state, so a test that wants to replace
|
| 19 |
+
one can monkeypatch it here without any of the problems shared mutable
|
| 20 |
+
state causes.
|
| 21 |
+
|
| 22 |
+
Services raise domain errors - `PatentNotFound`, `UpstreamUnavailable` -
|
| 23 |
+
rather than `HTTPException`. Mapping those onto status codes is the HTTP
|
| 24 |
+
layer's job, and lives in app.py.
|
| 25 |
+
"""
|
| 26 |
+
|
| 27 |
+
import asyncio
|
| 28 |
+
import logging
|
| 29 |
+
from typing import Callable, Optional
|
| 30 |
+
|
| 31 |
+
import httpx
|
| 32 |
+
from httpx import AsyncClient, HTTPStatusError
|
| 33 |
+
from playwright.async_api import Browser
|
| 34 |
+
|
| 35 |
+
from circuit_breaker import CircuitBreaker, CircuitOpenError
|
| 36 |
+
from ops import (OPSBulkResponse, ops_scrap_patent, ops_scrap_patent_bulk,
|
| 37 |
+
ops_search)
|
| 38 |
+
from scrap import (PatentScrapBulkResponse, PatentScrapResult,
|
| 39 |
+
scrap_patent_async, scrap_patent_bulk_async)
|
| 40 |
+
from serp import (EmptyResultsError, SerpQuery, SerpResults, query_arxiv,
|
| 41 |
+
query_bing_search, query_brave_search, query_ddg_search,
|
| 42 |
+
query_google_patents, query_google_scholar)
|
| 43 |
+
from utils import log_gathered_exceptions
|
| 44 |
+
|
| 45 |
+
logger = logging.getLogger(__name__)
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
# ================================ domain errors ================================
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
class PatentNotFound(Exception):
|
| 52 |
+
"""Every backend consulted positively said the document is absent.
|
| 53 |
+
|
| 54 |
+
Distinct from UpstreamUnavailable on purpose: MCP_INSTRUCTIONS tells an
|
| 55 |
+
agent that a not-found patent is genuinely absent everywhere and to
|
| 56 |
+
move on without retrying, so this must never stand in for "we couldn't
|
| 57 |
+
reach anyone".
|
| 58 |
+
"""
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
class UpstreamUnavailable(Exception):
|
| 62 |
+
"""A backend failed to answer. Says nothing about whether the document
|
| 63 |
+
exists."""
|
| 64 |
+
|
| 65 |
+
def __init__(self, message: str, *, timeout: bool = False):
|
| 66 |
+
super().__init__(message)
|
| 67 |
+
self.timeout = timeout
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
class OPSUnconfigured(Exception):
|
| 71 |
+
"""An OPS-only operation was requested without credentials configured."""
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def _is_not_found(exc: BaseException) -> bool:
|
| 75 |
+
"""True only when a backend positively said the document is absent."""
|
| 76 |
+
return isinstance(exc, HTTPStatusError) and exc.response.status_code == 404
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
def _as_upstream(exc: BaseException, message: str) -> UpstreamUnavailable:
|
| 80 |
+
return UpstreamUnavailable(message, timeout=isinstance(exc, httpx.TimeoutException))
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
# ================================ result shaping ================================
|
| 84 |
+
|
| 85 |
+
|
| 86 |
+
def shape_serp_results(results: list) -> SerpResults:
|
| 87 |
+
"""Flatten a list of per-query result-lists (or Exceptions, from a
|
| 88 |
+
`return_exceptions=True` gather) into one SerpResults, surfacing the
|
| 89 |
+
last error only when every query failed.
|
| 90 |
+
"""
|
| 91 |
+
if not results:
|
| 92 |
+
# The request models bound the query list, so this is unreachable
|
| 93 |
+
# from the endpoints; keep the helper total anyway rather than
|
| 94 |
+
# indexing into an empty list, since every search path shares it.
|
| 95 |
+
return SerpResults(results=[], error="No queries were provided.")
|
| 96 |
+
|
| 97 |
+
filtered_results = [r for r in results if not isinstance(r, Exception)]
|
| 98 |
+
flattened_results = [
|
| 99 |
+
item for sublist in filtered_results for item in sublist]
|
| 100 |
+
|
| 101 |
+
if len(filtered_results) == 0:
|
| 102 |
+
return SerpResults(results=[], error=str(results[-1]))
|
| 103 |
+
|
| 104 |
+
return SerpResults(results=flattened_results, error=None)
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
# ================================= search service =================================
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
class SearchService:
|
| 111 |
+
"""Runs SERP queries across the available backends."""
|
| 112 |
+
|
| 113 |
+
def __init__(self, *, http_client: AsyncClient,
|
| 114 |
+
browser_provider: Callable[[], Optional[Browser]],
|
| 115 |
+
circuit_breaker: CircuitBreaker, ops_tokens):
|
| 116 |
+
self._http = http_client
|
| 117 |
+
# A provider rather than the browser itself: the browser is started
|
| 118 |
+
# by the app's lifespan, after this service may already exist, and
|
| 119 |
+
# is None when Playwright failed to start.
|
| 120 |
+
self._browser_provider = browser_provider
|
| 121 |
+
self._breaker = circuit_breaker
|
| 122 |
+
self._ops_tokens = ops_tokens
|
| 123 |
+
|
| 124 |
+
@property
|
| 125 |
+
def browser(self) -> Optional[Browser]:
|
| 126 |
+
return self._browser_provider()
|
| 127 |
+
|
| 128 |
+
async def _run_queries(self, fn, params: SerpQuery, context: str) -> SerpResults:
|
| 129 |
+
"""Run `fn(query, n_results)` concurrently for every query, log any
|
| 130 |
+
failures, and shape the result the way every simple search path
|
| 131 |
+
expects.
|
| 132 |
+
"""
|
| 133 |
+
results = await asyncio.gather(
|
| 134 |
+
*[fn(q, params.n_results) for q in params.queries], return_exceptions=True)
|
| 135 |
+
log_gathered_exceptions(results, context, params.queries)
|
| 136 |
+
return shape_serp_results(results)
|
| 137 |
+
|
| 138 |
+
# ------------------------------ single backend ------------------------------
|
| 139 |
+
|
| 140 |
+
async def google_scholar(self, params: SerpQuery) -> SerpResults:
|
| 141 |
+
return await self._run_queries(
|
| 142 |
+
lambda q, n: query_google_scholar(self.browser, q, n), params,
|
| 143 |
+
"google scholar search")
|
| 144 |
+
|
| 145 |
+
async def arxiv(self, params: SerpQuery) -> SerpResults:
|
| 146 |
+
return await self._run_queries(
|
| 147 |
+
lambda q, n: query_arxiv(self._http, q, n), params, "arxiv search")
|
| 148 |
+
|
| 149 |
+
async def brave(self, params: SerpQuery) -> SerpResults:
|
| 150 |
+
return await self._run_queries(
|
| 151 |
+
lambda q, n: query_brave_search(self.browser, q, n), params, "brave search")
|
| 152 |
+
|
| 153 |
+
async def bing(self, params: SerpQuery) -> SerpResults:
|
| 154 |
+
return await self._run_queries(
|
| 155 |
+
lambda q, n: query_bing_search(self.browser, q, n), params, "bing search")
|
| 156 |
+
|
| 157 |
+
async def duckduckgo(self, params: SerpQuery) -> SerpResults:
|
| 158 |
+
return await self._run_queries(
|
| 159 |
+
lambda q, n: query_ddg_search(q, n), params, "duckduckgo search")
|
| 160 |
+
|
| 161 |
+
async def ops_keyword_search(self, params: SerpQuery) -> SerpResults:
|
| 162 |
+
return await self._run_queries(
|
| 163 |
+
lambda q, n: ops_search(self._http, q, n), params, "OPS search")
|
| 164 |
+
|
| 165 |
+
# -------------------------------- patents --------------------------------
|
| 166 |
+
|
| 167 |
+
async def patents(self, params: SerpQuery) -> SerpResults:
|
| 168 |
+
"""Google Patents, falling back to EPO OPS for any query it returns
|
| 169 |
+
nothing for (when OPS credentials are configured)."""
|
| 170 |
+
results = await asyncio.gather(
|
| 171 |
+
*[query_google_patents(self.browser, q, params.n_results)
|
| 172 |
+
for q in params.queries],
|
| 173 |
+
return_exceptions=True)
|
| 174 |
+
log_gathered_exceptions(results, "google patent search", params.queries)
|
| 175 |
+
|
| 176 |
+
# Gathered rather than awaited one at a time: this path exists for
|
| 177 |
+
# when Google Patents is unavailable, so it is exactly when a
|
| 178 |
+
# request can least afford N sequential round-trips to the slower
|
| 179 |
+
# backend.
|
| 180 |
+
if self._ops_tokens.configured:
|
| 181 |
+
needs_fallback = [
|
| 182 |
+
i for i, res in enumerate(results)
|
| 183 |
+
if isinstance(res, Exception) or not res]
|
| 184 |
+
if needs_fallback:
|
| 185 |
+
logger.info(
|
| 186 |
+
f"Google Patents empty for {len(needs_fallback)} quer(y/ies), trying OPS.")
|
| 187 |
+
fallback = await asyncio.gather(
|
| 188 |
+
*[ops_search(self._http, params.queries[i], params.n_results)
|
| 189 |
+
for i in needs_fallback],
|
| 190 |
+
return_exceptions=True)
|
| 191 |
+
# strict: gather returns exactly one result per index, so a
|
| 192 |
+
# length mismatch here would be a bug worth surfacing.
|
| 193 |
+
for i, res in zip(needs_fallback, fallback, strict=True):
|
| 194 |
+
if isinstance(res, Exception):
|
| 195 |
+
logger.warning(f"OPS fallback failed for `{params.queries[i]}`: {res}")
|
| 196 |
+
else:
|
| 197 |
+
results[i] = res
|
| 198 |
+
|
| 199 |
+
return shape_serp_results(results)
|
| 200 |
+
|
| 201 |
+
# ----------------------------- all backends -----------------------------
|
| 202 |
+
|
| 203 |
+
async def _search_one(self, q: str, n_results: int) -> tuple[str, list[dict], Optional[str]]:
|
| 204 |
+
"""Try DDG, then Brave, then Bing for a single query; stop at the
|
| 205 |
+
first success."""
|
| 206 |
+
backends = [
|
| 207 |
+
("DuckDuckGo", lambda: query_ddg_search(q, n_results)),
|
| 208 |
+
("Brave Search", lambda: query_brave_search(self.browser, q, n_results)),
|
| 209 |
+
("Bing", lambda: query_bing_search(self.browser, q, n_results)),
|
| 210 |
+
]
|
| 211 |
+
last_error: Optional[Exception] = None
|
| 212 |
+
for name, call in backends:
|
| 213 |
+
try:
|
| 214 |
+
self._breaker.before_call(name)
|
| 215 |
+
except CircuitOpenError as e:
|
| 216 |
+
logger.info(f"Skipping {name} for query `{q}`: {e}")
|
| 217 |
+
last_error = e
|
| 218 |
+
continue
|
| 219 |
+
|
| 220 |
+
try:
|
| 221 |
+
logger.info(f"Querying {name} with query: `{q}`")
|
| 222 |
+
result = await call()
|
| 223 |
+
self._breaker.record_success(name)
|
| 224 |
+
return q, result, None
|
| 225 |
+
except EmptyResultsError as e:
|
| 226 |
+
# Zero results is ambiguous - no hits, or a soft block that
|
| 227 |
+
# parsed to nothing. Move on to the next backend, but don't
|
| 228 |
+
# hold it against this one: the breaker is shared across
|
| 229 |
+
# every request, so counting it would let a few
|
| 230 |
+
# unusual-but-legitimate queries disable the primary backend
|
| 231 |
+
# for everyone. A hard block normally raises a real error,
|
| 232 |
+
# which is handled below.
|
| 233 |
+
logger.info(f"{name} returned no results for query `{q}`")
|
| 234 |
+
last_error = e
|
| 235 |
+
except Exception as e:
|
| 236 |
+
logger.error(f"Failed to query {name} with query `{q}`: {e}")
|
| 237 |
+
self._breaker.record_failure(name)
|
| 238 |
+
last_error = e
|
| 239 |
+
|
| 240 |
+
return q, [], f"All backends failed for query '{q}': {last_error}"
|
| 241 |
+
|
| 242 |
+
async def search(self, params: SerpQuery) -> SerpResults:
|
| 243 |
+
"""Search every query across the full backend fallback chain."""
|
| 244 |
+
outcomes = await asyncio.gather(
|
| 245 |
+
*[self._search_one(q, params.n_results) for q in params.queries])
|
| 246 |
+
|
| 247 |
+
results: list[dict] = []
|
| 248 |
+
errors: list[str] = []
|
| 249 |
+
for _q, res, err in outcomes:
|
| 250 |
+
results.extend(res)
|
| 251 |
+
if err:
|
| 252 |
+
errors.append(err)
|
| 253 |
+
|
| 254 |
+
if len(results) == 0:
|
| 255 |
+
return SerpResults(
|
| 256 |
+
results=[],
|
| 257 |
+
error="; ".join(errors) if errors else "All backends are rate-limited.")
|
| 258 |
+
|
| 259 |
+
return SerpResults(results=results, error="; ".join(errors) if errors else None)
|
| 260 |
+
|
| 261 |
+
|
| 262 |
+
# ================================= patent service =================================
|
| 263 |
+
|
| 264 |
+
|
| 265 |
+
class PatentService:
|
| 266 |
+
"""Retrieves patent full text, from Google Patents or EPO OPS."""
|
| 267 |
+
|
| 268 |
+
def __init__(self, *, http_client: AsyncClient, ops_tokens):
|
| 269 |
+
self._http = http_client
|
| 270 |
+
self._ops_tokens = ops_tokens
|
| 271 |
+
|
| 272 |
+
async def scrap(self, patent_id: str) -> PatentScrapResult:
|
| 273 |
+
"""Scrape from Google Patents, falling back to EPO OPS.
|
| 274 |
+
|
| 275 |
+
Raises PatentNotFound only when every backend consulted said the
|
| 276 |
+
document is absent; anything else raises UpstreamUnavailable.
|
| 277 |
+
"""
|
| 278 |
+
google_error: BaseException
|
| 279 |
+
try:
|
| 280 |
+
return await scrap_patent_async(
|
| 281 |
+
self._http, f"https://patents.google.com/patent/{patent_id}/en")
|
| 282 |
+
except HTTPStatusError as e:
|
| 283 |
+
google_error = e
|
| 284 |
+
logger.warning(
|
| 285 |
+
f"Google Patents returned {e.response.status_code} for {patent_id}.")
|
| 286 |
+
except Exception as e:
|
| 287 |
+
google_error = e
|
| 288 |
+
logger.warning(f"Failed to scrap patent {patent_id}: {e}")
|
| 289 |
+
|
| 290 |
+
if not self._ops_tokens.configured:
|
| 291 |
+
if _is_not_found(google_error):
|
| 292 |
+
raise PatentNotFound(
|
| 293 |
+
f"Patent '{patent_id}' not found on Google Patents "
|
| 294 |
+
"(EPO OPS fallback not configured).")
|
| 295 |
+
raise _as_upstream(
|
| 296 |
+
google_error,
|
| 297 |
+
f"Google Patents request failed for '{patent_id}' and the EPO OPS "
|
| 298 |
+
f"fallback is not configured: {google_error}")
|
| 299 |
+
|
| 300 |
+
try:
|
| 301 |
+
logger.info(f"Trying OPS for patent {patent_id}.")
|
| 302 |
+
return await ops_scrap_patent(self._http, patent_id)
|
| 303 |
+
except Exception as e:
|
| 304 |
+
logger.warning(f"OPS fallback failed for {patent_id}: {e}")
|
| 305 |
+
# Only claim the patent doesn't exist when *both* backends said so.
|
| 306 |
+
if _is_not_found(google_error) and _is_not_found(e):
|
| 307 |
+
raise PatentNotFound(
|
| 308 |
+
f"Patent '{patent_id}' not found on Google Patents or EPO OPS.") from e
|
| 309 |
+
if isinstance(e, HTTPStatusError):
|
| 310 |
+
raise _as_upstream(
|
| 311 |
+
e, f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
|
| 312 |
+
raise _as_upstream(e, f"EPO OPS request failed for '{patent_id}': {e}") from e
|
| 313 |
+
|
| 314 |
+
async def scrap_bulk(self, patent_ids: list[str]) -> PatentScrapBulkResponse:
|
| 315 |
+
return await scrap_patent_bulk_async(self._http, patent_ids)
|
| 316 |
+
|
| 317 |
+
async def ops_scrap(self, patent_id: str) -> PatentScrapResult:
|
| 318 |
+
"""Retrieve via EPO OPS only, with no Google Patents fallback."""
|
| 319 |
+
if not self._ops_tokens.configured:
|
| 320 |
+
raise OPSUnconfigured(
|
| 321 |
+
"EPO OPS is not configured (OPS_CONSUMER_KEY / OPS_CONSUMER_SECRET missing).")
|
| 322 |
+
try:
|
| 323 |
+
return await ops_scrap_patent(self._http, patent_id)
|
| 324 |
+
except HTTPStatusError as e:
|
| 325 |
+
if e.response.status_code == 404:
|
| 326 |
+
# OPS is the only backend asked here, so its 404 is the
|
| 327 |
+
# whole answer.
|
| 328 |
+
raise PatentNotFound(
|
| 329 |
+
f"Patent '{patent_id}' not found in EPO OPS.") from e
|
| 330 |
+
raise _as_upstream(
|
| 331 |
+
e, f"EPO OPS returned {e.response.status_code} for '{patent_id}'.") from e
|
| 332 |
+
except Exception as e:
|
| 333 |
+
logger.warning(f"Failed to retrieve patent {patent_id} from OPS: {e}")
|
| 334 |
+
raise _as_upstream(e, f"EPO OPS request failed for '{patent_id}': {e}") from e
|
| 335 |
+
|
| 336 |
+
async def ops_scrap_bulk(self, patent_ids: list[str]) -> OPSBulkResponse:
|
| 337 |
+
return await ops_scrap_patent_bulk(self._http, patent_ids)
|
|
@@ -1,12 +1,10 @@
|
|
| 1 |
-
"""
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
doesn't run it, and these tests don't need it since pw_browser is never
|
| 9 |
-
actually used once the query functions are patched out.
|
| 10 |
"""
|
| 11 |
|
| 12 |
import httpx
|
|
@@ -14,9 +12,8 @@ import pytest
|
|
| 14 |
from httpx import ASGITransport
|
| 15 |
|
| 16 |
import app as app_module
|
| 17 |
-
from
|
| 18 |
-
from
|
| 19 |
-
from serp import SerpQuery
|
| 20 |
|
| 21 |
|
| 22 |
@pytest.fixture
|
|
@@ -26,151 +23,134 @@ async def client():
|
|
| 26 |
yield c
|
| 27 |
|
| 28 |
|
| 29 |
-
@pytest.fixture
|
| 30 |
-
def
|
| 31 |
-
"""
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
# ------------------------- _shape_serp_results / _run_serp_queries -------------------------
|
| 40 |
|
|
|
|
|
|
|
| 41 |
|
| 42 |
-
def test_shape_serp_results_flattens_successful_lists():
|
| 43 |
-
result = app_module._shape_serp_results([[{"a": 1}], [{"b": 2}, {"c": 3}]])
|
| 44 |
-
assert result.results == [{"a": 1}, {"b": 2}, {"c": 3}]
|
| 45 |
-
assert result.error is None
|
| 46 |
|
|
|
|
|
|
|
| 47 |
|
| 48 |
-
def
|
| 49 |
-
|
| 50 |
-
assert result.results == []
|
| 51 |
-
assert "second" in result.error
|
| 52 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
|
| 54 |
-
def
|
| 55 |
-
|
| 56 |
-
assert result.results == [{"a": 1}]
|
| 57 |
-
assert result.error is None
|
| 58 |
|
| 59 |
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
|
| 64 |
-
async def
|
| 65 |
-
calls.append((
|
| 66 |
-
return
|
| 67 |
|
| 68 |
-
|
|
|
|
|
|
|
| 69 |
|
| 70 |
-
|
| 71 |
-
|
|
|
|
| 72 |
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
-
# ------------------------------------- endpoints -------------------------------------
|
| 75 |
|
|
|
|
| 76 |
|
| 77 |
-
async def test_search_arxiv_flattens_results_from_all_queries(client, monkeypatch):
|
| 78 |
-
async def fake_query_arxiv(client_arg, q, n):
|
| 79 |
-
return [{"title": f"paper about {q}", "href": "x", "body": "y", "id": "1"}]
|
| 80 |
|
| 81 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
-
resp = await client.post(
|
| 84 |
|
| 85 |
assert resp.status_code == 200
|
| 86 |
-
|
| 87 |
-
assert
|
| 88 |
-
assert {r["title"] for r in data["results"]} == {"paper about a", "paper about b"}
|
| 89 |
-
|
| 90 |
|
| 91 |
-
async def test_search_arxiv_all_queries_failing_returns_error(client, monkeypatch):
|
| 92 |
-
async def failing_query_arxiv(client_arg, q, n):
|
| 93 |
-
raise RuntimeError(f"boom for {q}")
|
| 94 |
|
| 95 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
|
| 97 |
-
resp = await
|
| 98 |
|
| 99 |
assert resp.status_code == 200
|
| 100 |
-
|
| 101 |
-
assert data["results"] == []
|
| 102 |
-
assert "boom for a" in data["error"]
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
async def test_search_patents_does_not_fall_back_when_google_patents_succeeds(client, monkeypatch):
|
| 106 |
-
async def fake_google_patents(browser, q, n):
|
| 107 |
-
return [{"id": "US1", "href": "h", "title": "t", "body": "b"}]
|
| 108 |
-
|
| 109 |
-
ops_search_called = False
|
| 110 |
-
|
| 111 |
-
async def fake_ops_search(client_arg, q, n):
|
| 112 |
-
nonlocal ops_search_called
|
| 113 |
-
ops_search_called = True
|
| 114 |
-
return []
|
| 115 |
-
|
| 116 |
-
monkeypatch.setattr(app_module, "query_google_patents", fake_google_patents)
|
| 117 |
-
monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
|
| 118 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
|
| 119 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
|
| 120 |
-
|
| 121 |
-
resp = await client.post("/serp/search_patents", json={"queries": ["widget"], "n_results": 5})
|
| 122 |
-
|
| 123 |
-
data = resp.json()
|
| 124 |
-
assert data["results"] == [{"id": "US1", "href": "h", "title": "t", "body": "b"}]
|
| 125 |
-
assert ops_search_called is False
|
| 126 |
-
|
| 127 |
|
| 128 |
-
async def test_search_patents_falls_back_to_ops_when_google_patents_empty(client, monkeypatch):
|
| 129 |
-
async def fake_google_patents(browser, q, n):
|
| 130 |
-
return []
|
| 131 |
|
| 132 |
-
|
| 133 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 134 |
|
| 135 |
-
|
| 136 |
-
monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
|
| 137 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
|
| 138 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
|
| 139 |
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
data = resp.json()
|
| 143 |
-
assert data["results"] == [{"title": "OPS result", "body": "", "href": "", "id": "US123"}]
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
async def test_search_patents_no_fallback_when_ops_not_configured(client, monkeypatch):
|
| 147 |
-
async def fake_google_patents(browser, q, n):
|
| 148 |
-
return []
|
| 149 |
-
|
| 150 |
-
monkeypatch.setattr(app_module, "query_google_patents", fake_google_patents)
|
| 151 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_key", None)
|
| 152 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_secret", None)
|
| 153 |
-
|
| 154 |
-
resp = await client.post("/serp/search_patents", json={"queries": ["widget"], "n_results": 5})
|
| 155 |
-
|
| 156 |
-
data = resp.json()
|
| 157 |
-
assert data["results"] == []
|
| 158 |
-
assert data["error"] is None
|
| 159 |
|
| 160 |
|
| 161 |
# ------------------------------- patent_id validation -------------------------------
|
| 162 |
|
| 163 |
|
| 164 |
-
|
| 165 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 166 |
assert resp.status_code == 422
|
| 167 |
|
| 168 |
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
|
|
|
|
|
|
| 172 |
|
| 173 |
-
|
|
|
|
|
|
|
| 174 |
|
| 175 |
resp = await client.get("/scrap/scrap_patent/US11930446B2")
|
| 176 |
|
|
@@ -178,38 +158,47 @@ async def test_scrap_patent_accepts_well_formed_patent_id(client, monkeypatch):
|
|
| 178 |
assert resp.json()["title"] == "Widget apparatus"
|
| 179 |
|
| 180 |
|
| 181 |
-
|
| 182 |
-
resp = await client.get("/ops/scrap_patent/not-a-patent-id")
|
| 183 |
-
assert resp.status_code == 422
|
| 184 |
|
| 185 |
|
| 186 |
-
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
|
| 190 |
|
|
|
|
|
|
|
|
|
|
| 191 |
|
| 192 |
-
async def test_ops_scrap_patents_bulk_rejects_malformed_id_in_list(client):
|
| 193 |
-
resp = await client.post("/ops/scrap_patents_bulk", json={"patent_ids": ["garbage!!"]})
|
| 194 |
-
assert resp.status_code == 422
|
| 195 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 196 |
|
| 197 |
-
|
| 198 |
-
async def fake_ops_search(client_arg, q, n):
|
| 199 |
-
return [{"title": f"OPS {q}", "body": "", "href": "", "id": "1"}]
|
| 200 |
|
| 201 |
-
monkeypatch.setattr(app_module, "ops_search", fake_ops_search)
|
| 202 |
|
| 203 |
-
|
|
|
|
| 204 |
|
| 205 |
-
|
| 206 |
-
assert
|
| 207 |
-
assert {r["title"] for r in data["results"]} == {"OPS a", "OPS b"}
|
| 208 |
|
| 209 |
|
| 210 |
# ------------------------------------- api_lifespan -------------------------------------
|
| 211 |
|
| 212 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 213 |
async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
|
| 214 |
"""The container runs as a non-root user (see the Dockerfile), and
|
| 215 |
Chromium's own internal sandbox needs privileges a non-root container
|
|
@@ -255,24 +244,10 @@ async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
|
|
| 255 |
assert "--no-sandbox" in launch_calls[0].get("args", [])
|
| 256 |
|
| 257 |
|
| 258 |
-
class _RecordingClient:
|
| 259 |
-
def __init__(self):
|
| 260 |
-
self.closed = False
|
| 261 |
-
|
| 262 |
-
async def aclose(self):
|
| 263 |
-
self.closed = True
|
| 264 |
-
|
| 265 |
-
|
| 266 |
async def test_api_lifespan_closes_the_http_client_on_shutdown(monkeypatch):
|
| 267 |
"""The client was created at import and never closed - the browser was
|
| 268 |
torn down on shutdown but its connection pool was not.
|
| 269 |
"""
|
| 270 |
-
class FakePlaywright:
|
| 271 |
-
chromium = None
|
| 272 |
-
|
| 273 |
-
async def stop(self):
|
| 274 |
-
pass
|
| 275 |
-
|
| 276 |
recording_client = _RecordingClient()
|
| 277 |
monkeypatch.setattr(app_module, "pw_browser", None)
|
| 278 |
monkeypatch.setattr(app_module, "playwright", None)
|
|
@@ -301,257 +276,3 @@ async def test_api_lifespan_survives_playwright_failing_to_start(monkeypatch):
|
|
| 301 |
|
| 302 |
async with app_module.api_lifespan(app_module.app):
|
| 303 |
assert app_module.pw_browser is None
|
| 304 |
-
|
| 305 |
-
|
| 306 |
-
# --------------------------------- _search_one fallback ---------------------------------
|
| 307 |
-
|
| 308 |
-
|
| 309 |
-
async def test_search_one_returns_first_backend_that_succeeds(monkeypatch):
|
| 310 |
-
calls = []
|
| 311 |
-
|
| 312 |
-
async def fake_ddg(q, n):
|
| 313 |
-
calls.append("ddg")
|
| 314 |
-
return [{"title": "from ddg"}]
|
| 315 |
-
|
| 316 |
-
async def fake_brave(browser, q, n):
|
| 317 |
-
calls.append("brave")
|
| 318 |
-
return [{"title": "from brave"}]
|
| 319 |
-
|
| 320 |
-
monkeypatch.setattr(app_module, "query_ddg_search", fake_ddg)
|
| 321 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 322 |
-
|
| 323 |
-
q, results, err = await app_module._search_one("widget", 5)
|
| 324 |
-
|
| 325 |
-
assert calls == ["ddg"]
|
| 326 |
-
assert results == [{"title": "from ddg"}]
|
| 327 |
-
assert err is None
|
| 328 |
-
|
| 329 |
-
|
| 330 |
-
async def test_search_one_falls_through_to_the_next_backend_on_failure(monkeypatch):
|
| 331 |
-
calls = []
|
| 332 |
-
|
| 333 |
-
async def failing_ddg(q, n):
|
| 334 |
-
calls.append("ddg")
|
| 335 |
-
raise RuntimeError("ddg blocked")
|
| 336 |
-
|
| 337 |
-
async def fake_brave(browser, q, n):
|
| 338 |
-
calls.append("brave")
|
| 339 |
-
return [{"title": "from brave"}]
|
| 340 |
-
|
| 341 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
|
| 342 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 343 |
-
|
| 344 |
-
q, results, err = await app_module._search_one("widget", 5)
|
| 345 |
-
|
| 346 |
-
assert calls == ["ddg", "brave"]
|
| 347 |
-
assert results == [{"title": "from brave"}]
|
| 348 |
-
assert err is None
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
async def test_search_one_reports_an_error_when_every_backend_fails(monkeypatch):
|
| 352 |
-
async def failing(*args):
|
| 353 |
-
raise RuntimeError("blocked")
|
| 354 |
-
|
| 355 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing)
|
| 356 |
-
monkeypatch.setattr(app_module, "query_brave_search", failing)
|
| 357 |
-
monkeypatch.setattr(app_module, "query_bing_search", failing)
|
| 358 |
-
|
| 359 |
-
q, results, err = await app_module._search_one("widget", 5)
|
| 360 |
-
|
| 361 |
-
assert results == []
|
| 362 |
-
assert "All backends failed for query 'widget'" in err
|
| 363 |
-
assert "blocked" in err
|
| 364 |
-
|
| 365 |
-
|
| 366 |
-
async def test_search_one_skips_a_backend_whose_circuit_is_open(monkeypatch):
|
| 367 |
-
"""A backend that's failed `failure_threshold` times in a row shouldn't
|
| 368 |
-
be attempted again until its cooldown elapses - repeatedly hammering a
|
| 369 |
-
backend that's already rate-limiting us only makes that worse.
|
| 370 |
-
"""
|
| 371 |
-
from circuit_breaker import CircuitBreaker
|
| 372 |
-
|
| 373 |
-
fresh_breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
|
| 374 |
-
monkeypatch.setattr(app_module, "_backend_circuit_breaker", fresh_breaker)
|
| 375 |
-
|
| 376 |
-
ddg_calls = []
|
| 377 |
-
|
| 378 |
-
async def failing_ddg(q, n):
|
| 379 |
-
ddg_calls.append(q)
|
| 380 |
-
raise RuntimeError("ddg blocked")
|
| 381 |
-
|
| 382 |
-
async def fake_brave(browser, q, n):
|
| 383 |
-
return [{"title": "from brave"}]
|
| 384 |
-
|
| 385 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
|
| 386 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 387 |
-
|
| 388 |
-
# First call: DDG fails, opening its circuit (threshold=1); Brave picks it up.
|
| 389 |
-
await app_module._search_one("first query", 5)
|
| 390 |
-
assert ddg_calls == ["first query"]
|
| 391 |
-
|
| 392 |
-
# Second call: DDG's circuit is now open, so it should be skipped
|
| 393 |
-
# entirely (not called again) and go straight to Brave.
|
| 394 |
-
q, results, err = await app_module._search_one("second query", 5)
|
| 395 |
-
assert ddg_calls == ["first query"] # unchanged - DDG was skipped
|
| 396 |
-
assert results == [{"title": "from brave"}]
|
| 397 |
-
|
| 398 |
-
|
| 399 |
-
# ------------------------------- /serp/search endpoint -------------------------------
|
| 400 |
-
# `_search_one` is covered above; this is the aggregation wrapped around it,
|
| 401 |
-
# which had no coverage - dropping every per-query error left the suite green.
|
| 402 |
-
|
| 403 |
-
|
| 404 |
-
async def test_search_endpoint_merges_results_from_every_query(client, monkeypatch):
|
| 405 |
-
async def fake_ddg(q, n):
|
| 406 |
-
return [{"title": f"result for {q}"}]
|
| 407 |
-
|
| 408 |
-
monkeypatch.setattr(app_module, "query_ddg_search", fake_ddg)
|
| 409 |
-
|
| 410 |
-
resp = await client.post("/serp/search", json={"queries": ["a", "b"], "n_results": 5})
|
| 411 |
-
|
| 412 |
-
data = resp.json()
|
| 413 |
-
assert {r["title"] for r in data["results"]} == {"result for a", "result for b"}
|
| 414 |
-
assert data["error"] is None
|
| 415 |
-
|
| 416 |
-
|
| 417 |
-
async def test_search_endpoint_reports_partial_failure_alongside_results(client, monkeypatch):
|
| 418 |
-
"""A query that no backend could answer must be reported even when
|
| 419 |
-
another query succeeded - otherwise the caller silently receives fewer
|
| 420 |
-
results than they asked for with no indication why.
|
| 421 |
-
"""
|
| 422 |
-
async def selective_ddg(q, n):
|
| 423 |
-
if q == "bad":
|
| 424 |
-
raise RuntimeError("blocked")
|
| 425 |
-
return [{"title": f"result for {q}"}]
|
| 426 |
-
|
| 427 |
-
async def failing(*args):
|
| 428 |
-
raise RuntimeError("blocked")
|
| 429 |
-
|
| 430 |
-
monkeypatch.setattr(app_module, "query_ddg_search", selective_ddg)
|
| 431 |
-
monkeypatch.setattr(app_module, "query_brave_search", failing)
|
| 432 |
-
monkeypatch.setattr(app_module, "query_bing_search", failing)
|
| 433 |
-
|
| 434 |
-
resp = await client.post("/serp/search", json={"queries": ["good", "bad"], "n_results": 5})
|
| 435 |
-
|
| 436 |
-
data = resp.json()
|
| 437 |
-
assert [r["title"] for r in data["results"]] == ["result for good"]
|
| 438 |
-
assert "bad" in data["error"]
|
| 439 |
-
|
| 440 |
-
|
| 441 |
-
async def test_search_endpoint_reports_an_error_when_everything_fails(client, monkeypatch):
|
| 442 |
-
async def failing(*args):
|
| 443 |
-
raise RuntimeError("blocked")
|
| 444 |
-
|
| 445 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing)
|
| 446 |
-
monkeypatch.setattr(app_module, "query_brave_search", failing)
|
| 447 |
-
monkeypatch.setattr(app_module, "query_bing_search", failing)
|
| 448 |
-
|
| 449 |
-
resp = await client.post("/serp/search", json={"queries": ["a", "b"], "n_results": 5})
|
| 450 |
-
|
| 451 |
-
data = resp.json()
|
| 452 |
-
assert data["results"] == []
|
| 453 |
-
# Both failed queries are named, not just the last one.
|
| 454 |
-
assert "'a'" in data["error"] and "'b'" in data["error"]
|
| 455 |
-
|
| 456 |
-
|
| 457 |
-
async def test_search_endpoint_falls_through_backends_per_query(client, monkeypatch):
|
| 458 |
-
"""The fallback chain is per query, so one blocked backend shouldn't
|
| 459 |
-
cost results for queries the next backend can answer."""
|
| 460 |
-
async def failing_ddg(q, n):
|
| 461 |
-
raise RuntimeError("ddg blocked")
|
| 462 |
-
|
| 463 |
-
async def fake_brave(browser, q, n):
|
| 464 |
-
return [{"title": f"brave has {q}"}]
|
| 465 |
-
|
| 466 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
|
| 467 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 468 |
-
|
| 469 |
-
resp = await client.post("/serp/search", json={"queries": ["a"], "n_results": 5})
|
| 470 |
-
|
| 471 |
-
data = resp.json()
|
| 472 |
-
assert data["results"] == [{"title": "brave has a"}]
|
| 473 |
-
assert data["error"] is None
|
| 474 |
-
|
| 475 |
-
|
| 476 |
-
# --------------------- what counts as a backend failure ---------------------
|
| 477 |
-
|
| 478 |
-
|
| 479 |
-
async def test_a_zero_result_query_does_not_open_the_backends_circuit(monkeypatch):
|
| 480 |
-
"""An empty result is ambiguous - it can mean the query genuinely has
|
| 481 |
-
no hits, or that the backend served an anti-bot page that parsed to
|
| 482 |
-
nothing. The circuit breaker is a process-wide singleton, so counting
|
| 483 |
-
the ambiguous case meant three unusual-but-legitimate queries in a row
|
| 484 |
-
could disable the primary backend for every user of the deployment for
|
| 485 |
-
a full cooldown.
|
| 486 |
-
"""
|
| 487 |
-
from circuit_breaker import CircuitBreaker
|
| 488 |
-
from serp import DuckDuckGoBlockedException
|
| 489 |
-
|
| 490 |
-
breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
|
| 491 |
-
monkeypatch.setattr(app_module, "_backend_circuit_breaker", breaker)
|
| 492 |
-
|
| 493 |
-
ddg_calls = []
|
| 494 |
-
|
| 495 |
-
async def empty_ddg(q, n):
|
| 496 |
-
ddg_calls.append(q)
|
| 497 |
-
raise DuckDuckGoBlockedException()
|
| 498 |
-
|
| 499 |
-
async def fake_brave(browser, q, n):
|
| 500 |
-
return [{"title": "from brave"}]
|
| 501 |
-
|
| 502 |
-
monkeypatch.setattr(app_module, "query_ddg_search", empty_ddg)
|
| 503 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 504 |
-
|
| 505 |
-
await app_module._search_one("obscure query", 5)
|
| 506 |
-
await app_module._search_one("another obscure query", 5)
|
| 507 |
-
|
| 508 |
-
# Still being tried: an empty result is not evidence the backend is down.
|
| 509 |
-
assert ddg_calls == ["obscure query", "another obscure query"]
|
| 510 |
-
|
| 511 |
-
|
| 512 |
-
async def test_an_empty_result_still_falls_through_to_the_next_backend(monkeypatch):
|
| 513 |
-
"""Not counting it against the backend must not stop the fallback: the
|
| 514 |
-
caller still needs results from somewhere."""
|
| 515 |
-
from serp import DuckDuckGoBlockedException
|
| 516 |
-
|
| 517 |
-
async def empty_ddg(q, n):
|
| 518 |
-
raise DuckDuckGoBlockedException()
|
| 519 |
-
|
| 520 |
-
async def fake_brave(browser, q, n):
|
| 521 |
-
return [{"title": "from brave"}]
|
| 522 |
-
|
| 523 |
-
monkeypatch.setattr(app_module, "query_ddg_search", empty_ddg)
|
| 524 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 525 |
-
|
| 526 |
-
_q, results, err = await app_module._search_one("widget", 5)
|
| 527 |
-
|
| 528 |
-
assert results == [{"title": "from brave"}]
|
| 529 |
-
assert err is None
|
| 530 |
-
|
| 531 |
-
|
| 532 |
-
async def test_a_real_backend_error_still_opens_the_circuit(monkeypatch):
|
| 533 |
-
"""The distinction has to cut both ways: a backend that raises a
|
| 534 |
-
genuine error (a rate limit, a transport failure) is still evidence,
|
| 535 |
-
and must still trip the breaker.
|
| 536 |
-
"""
|
| 537 |
-
from circuit_breaker import CircuitBreaker
|
| 538 |
-
|
| 539 |
-
breaker = CircuitBreaker(failure_threshold=1, cooldown_seconds=60)
|
| 540 |
-
monkeypatch.setattr(app_module, "_backend_circuit_breaker", breaker)
|
| 541 |
-
|
| 542 |
-
ddg_calls = []
|
| 543 |
-
|
| 544 |
-
async def failing_ddg(q, n):
|
| 545 |
-
ddg_calls.append(q)
|
| 546 |
-
raise RuntimeError("rate limited")
|
| 547 |
-
|
| 548 |
-
async def fake_brave(browser, q, n):
|
| 549 |
-
return [{"title": "from brave"}]
|
| 550 |
-
|
| 551 |
-
monkeypatch.setattr(app_module, "query_ddg_search", failing_ddg)
|
| 552 |
-
monkeypatch.setattr(app_module, "query_brave_search", fake_brave)
|
| 553 |
-
|
| 554 |
-
await app_module._search_one("first", 5)
|
| 555 |
-
await app_module._search_one("second", 5)
|
| 556 |
-
|
| 557 |
-
assert ddg_calls == ["first"] # skipped the second time
|
|
|
|
| 1 |
+
"""app.py: HTTP wiring only.
|
| 2 |
+
|
| 3 |
+
The orchestration these endpoints used to contain now lives in
|
| 4 |
+
services.py and is tested in test_search_service.py and
|
| 5 |
+
test_patent_fallback.py without an HTTP layer. What's left here is what
|
| 6 |
+
app.py is actually responsible for: request validation, dependency wiring,
|
| 7 |
+
the lifespan, and turning domain results into responses.
|
|
|
|
|
|
|
| 8 |
"""
|
| 9 |
|
| 10 |
import httpx
|
|
|
|
| 12 |
from httpx import ASGITransport
|
| 13 |
|
| 14 |
import app as app_module
|
| 15 |
+
from scrap import PatentScrapBulkResponse, PatentScrapResult
|
| 16 |
+
from serp import SerpResults
|
|
|
|
| 17 |
|
| 18 |
|
| 19 |
@pytest.fixture
|
|
|
|
| 23 |
yield c
|
| 24 |
|
| 25 |
|
| 26 |
+
@pytest.fixture
|
| 27 |
+
def override_services():
|
| 28 |
+
"""Install stand-in services for the duration of one test."""
|
| 29 |
+
def _install(*, search=None, patent=None):
|
| 30 |
+
if search is not None:
|
| 31 |
+
app_module.app.dependency_overrides[app_module.get_search_service] = lambda: search
|
| 32 |
+
if patent is not None:
|
| 33 |
+
app_module.app.dependency_overrides[app_module.get_patent_service] = lambda: patent
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
+
yield _install
|
| 36 |
+
app_module.app.dependency_overrides.clear()
|
| 37 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
+
class _RecordingSearchService:
|
| 40 |
+
"""Records which service method each endpoint dispatched to."""
|
| 41 |
|
| 42 |
+
def __init__(self):
|
| 43 |
+
self.calls = []
|
|
|
|
|
|
|
| 44 |
|
| 45 |
+
def _record(self, name):
|
| 46 |
+
async def _fn(params):
|
| 47 |
+
self.calls.append((name, list(params.queries), params.n_results))
|
| 48 |
+
return SerpResults(results=[{"title": name}], error=None)
|
| 49 |
+
return _fn
|
| 50 |
|
| 51 |
+
def __getattr__(self, name):
|
| 52 |
+
return self._record(name)
|
|
|
|
|
|
|
| 53 |
|
| 54 |
|
| 55 |
+
class _StubPatentService:
|
| 56 |
+
def __init__(self):
|
| 57 |
+
self.calls = []
|
| 58 |
|
| 59 |
+
async def scrap(self, patent_id):
|
| 60 |
+
self.calls.append(("scrap", patent_id))
|
| 61 |
+
return PatentScrapResult(title="Widget apparatus")
|
| 62 |
|
| 63 |
+
async def scrap_bulk(self, patent_ids):
|
| 64 |
+
self.calls.append(("scrap_bulk", patent_ids))
|
| 65 |
+
return PatentScrapBulkResponse(patents=[], failed_ids=list(patent_ids))
|
| 66 |
|
| 67 |
+
async def ops_scrap(self, patent_id):
|
| 68 |
+
self.calls.append(("ops_scrap", patent_id))
|
| 69 |
+
return PatentScrapResult(title="From OPS")
|
| 70 |
|
| 71 |
+
async def ops_scrap_bulk(self, patent_ids):
|
| 72 |
+
self.calls.append(("ops_scrap_bulk", patent_ids))
|
| 73 |
+
return PatentScrapBulkResponse(patents=[], failed_ids=list(patent_ids))
|
| 74 |
|
|
|
|
| 75 |
|
| 76 |
+
# ------------------------------- endpoint dispatch -------------------------------
|
| 77 |
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
+
@pytest.mark.parametrize("path, expected_method", [
|
| 80 |
+
("/serp/search_scholar", "google_scholar"),
|
| 81 |
+
("/serp/search_arxiv", "arxiv"),
|
| 82 |
+
("/serp/search_patents", "patents"),
|
| 83 |
+
("/serp/search_brave", "brave"),
|
| 84 |
+
("/serp/search_bing", "bing"),
|
| 85 |
+
("/serp/search_duck", "duckduckgo"),
|
| 86 |
+
("/serp/search", "search"),
|
| 87 |
+
("/ops/search", "ops_keyword_search"),
|
| 88 |
+
])
|
| 89 |
+
async def test_each_search_endpoint_dispatches_to_its_service_method(
|
| 90 |
+
client, override_services, path, expected_method):
|
| 91 |
+
"""Guards the wiring itself: every endpoint reaching the right method
|
| 92 |
+
with the request's queries and result count intact."""
|
| 93 |
+
service = _RecordingSearchService()
|
| 94 |
+
override_services(search=service)
|
| 95 |
|
| 96 |
+
resp = await client.post(path, json={"queries": ["a", "b"], "n_results": 25})
|
| 97 |
|
| 98 |
assert resp.status_code == 200
|
| 99 |
+
assert service.calls == [(expected_method, ["a", "b"], 25)]
|
| 100 |
+
assert resp.json()["results"] == [{"title": expected_method}]
|
|
|
|
|
|
|
| 101 |
|
|
|
|
|
|
|
|
|
|
| 102 |
|
| 103 |
+
@pytest.mark.parametrize("method, path, expected_call", [
|
| 104 |
+
("get", "/scrap/scrap_patent/US11930446B2", ("scrap", "US11930446B2")),
|
| 105 |
+
("get", "/ops/scrap_patent/US11930446B2", ("ops_scrap", "US11930446B2")),
|
| 106 |
+
])
|
| 107 |
+
async def test_patent_endpoints_dispatch_to_their_service_method(
|
| 108 |
+
client, override_services, method, path, expected_call):
|
| 109 |
+
service = _StubPatentService()
|
| 110 |
+
override_services(patent=service)
|
| 111 |
|
| 112 |
+
resp = await getattr(client, method)(path)
|
| 113 |
|
| 114 |
assert resp.status_code == 200
|
| 115 |
+
assert service.calls == [expected_call]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 116 |
|
|
|
|
|
|
|
|
|
|
| 117 |
|
| 118 |
+
@pytest.mark.parametrize("path, expected_method", [
|
| 119 |
+
("/scrap/scrap_patents_bulk", "scrap_bulk"),
|
| 120 |
+
("/ops/scrap_patents_bulk", "ops_scrap_bulk"),
|
| 121 |
+
])
|
| 122 |
+
async def test_bulk_endpoints_dispatch_to_their_service_method(
|
| 123 |
+
client, override_services, path, expected_method):
|
| 124 |
+
service = _StubPatentService()
|
| 125 |
+
override_services(patent=service)
|
| 126 |
|
| 127 |
+
resp = await client.post(path, json={"patent_ids": ["US11930446B2"]})
|
|
|
|
|
|
|
|
|
|
| 128 |
|
| 129 |
+
assert resp.status_code == 200
|
| 130 |
+
assert service.calls == [(expected_method, ["US11930446B2"])]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
|
| 132 |
|
| 133 |
# ------------------------------- patent_id validation -------------------------------
|
| 134 |
|
| 135 |
|
| 136 |
+
@pytest.mark.parametrize("path", [
|
| 137 |
+
"/scrap/scrap_patent/not-a-patent-id",
|
| 138 |
+
"/ops/scrap_patent/not-a-patent-id",
|
| 139 |
+
])
|
| 140 |
+
async def test_malformed_patent_id_is_rejected(client, path):
|
| 141 |
+
resp = await client.get(path)
|
| 142 |
assert resp.status_code == 422
|
| 143 |
|
| 144 |
|
| 145 |
+
@pytest.mark.parametrize("path", [
|
| 146 |
+
"/scrap/scrap_patents_bulk", "/ops/scrap_patents_bulk"])
|
| 147 |
+
async def test_malformed_patent_id_in_a_bulk_list_is_rejected(client, path):
|
| 148 |
+
resp = await client.post(path, json={"patent_ids": ["US11930446B2", "garbage!!"]})
|
| 149 |
+
assert resp.status_code == 422
|
| 150 |
|
| 151 |
+
|
| 152 |
+
async def test_well_formed_patent_id_is_accepted(client, override_services):
|
| 153 |
+
override_services(patent=_StubPatentService())
|
| 154 |
|
| 155 |
resp = await client.get("/scrap/scrap_patent/US11930446B2")
|
| 156 |
|
|
|
|
| 158 |
assert resp.json()["title"] == "Widget apparatus"
|
| 159 |
|
| 160 |
|
| 161 |
+
# ---------------------------- default dependency wiring ----------------------------
|
|
|
|
|
|
|
| 162 |
|
| 163 |
|
| 164 |
+
def test_the_default_search_service_gets_the_apps_collaborators():
|
| 165 |
+
"""The providers exist so tests can override them; this pins that the
|
| 166 |
+
un-overridden ones still hand the service the real collaborators."""
|
| 167 |
+
service = app_module.get_search_service()
|
| 168 |
|
| 169 |
+
assert service._http is app_module.httpx_client
|
| 170 |
+
assert service._breaker is app_module._backend_circuit_breaker
|
| 171 |
+
assert service._ops_tokens is app_module.ops_token_manager
|
| 172 |
|
|
|
|
|
|
|
|
|
|
| 173 |
|
| 174 |
+
def test_the_search_service_reads_the_browser_lazily(monkeypatch):
|
| 175 |
+
"""The browser is started by the lifespan, after the service may already
|
| 176 |
+
exist, and stays None when Playwright fails to start - so the service
|
| 177 |
+
has to read it per call rather than capture it at construction."""
|
| 178 |
+
service = app_module.get_search_service()
|
| 179 |
+
monkeypatch.setattr(app_module, "pw_browser", "a-browser")
|
| 180 |
|
| 181 |
+
assert service.browser == "a-browser"
|
|
|
|
|
|
|
| 182 |
|
|
|
|
| 183 |
|
| 184 |
+
def test_the_default_patent_service_gets_the_apps_collaborators():
|
| 185 |
+
service = app_module.get_patent_service()
|
| 186 |
|
| 187 |
+
assert service._http is app_module.httpx_client
|
| 188 |
+
assert service._ops_tokens is app_module.ops_token_manager
|
|
|
|
| 189 |
|
| 190 |
|
| 191 |
# ------------------------------------- api_lifespan -------------------------------------
|
| 192 |
|
| 193 |
|
| 194 |
+
class _RecordingClient:
|
| 195 |
+
def __init__(self):
|
| 196 |
+
self.closed = False
|
| 197 |
+
|
| 198 |
+
async def aclose(self):
|
| 199 |
+
self.closed = True
|
| 200 |
+
|
| 201 |
+
|
| 202 |
async def test_api_lifespan_launches_chromium_without_a_sandbox(monkeypatch):
|
| 203 |
"""The container runs as a non-root user (see the Dockerfile), and
|
| 204 |
Chromium's own internal sandbox needs privileges a non-root container
|
|
|
|
| 244 |
assert "--no-sandbox" in launch_calls[0].get("args", [])
|
| 245 |
|
| 246 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 247 |
async def test_api_lifespan_closes_the_http_client_on_shutdown(monkeypatch):
|
| 248 |
"""The client was created at import and never closed - the browser was
|
| 249 |
torn down on shutdown but its connection pool was not.
|
| 250 |
"""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 251 |
recording_client = _RecordingClient()
|
| 252 |
monkeypatch.setattr(app_module, "pw_browser", None)
|
| 253 |
monkeypatch.setattr(app_module, "playwright", None)
|
|
|
|
| 276 |
|
| 277 |
async with app_module.api_lifespan(app_module.app):
|
| 278 |
assert app_module.pw_browser is None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@@ -15,6 +15,7 @@ import pytest
|
|
| 15 |
from httpx import ASGITransport
|
| 16 |
|
| 17 |
import app as app_module
|
|
|
|
| 18 |
|
| 19 |
|
| 20 |
@pytest.fixture
|
|
@@ -47,7 +48,7 @@ async def test_no_key_configured_leaves_the_api_open(client, monkeypatch):
|
|
| 47 |
"""The middleware is a no-op unless SERPENT_API_KEY is set - this keeps
|
| 48 |
the default deployment exactly as open as it was before it existed.
|
| 49 |
"""
|
| 50 |
-
monkeypatch.setattr(
|
| 51 |
|
| 52 |
resp = await client.post("/serp/search_arxiv", json={"queries": ["a"]})
|
| 53 |
|
|
@@ -76,7 +77,7 @@ async def test_wrong_key_is_rejected(client, api_key):
|
|
| 76 |
|
| 77 |
|
| 78 |
async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, monkeypatch):
|
| 79 |
-
monkeypatch.setattr(
|
| 80 |
|
| 81 |
resp = await client.post(
|
| 82 |
"/serp/search_arxiv", json={"queries": ["a"]}, headers={"X-API-Key": api_key})
|
|
@@ -85,7 +86,7 @@ async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, mon
|
|
| 85 |
|
| 86 |
|
| 87 |
async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkeypatch):
|
| 88 |
-
monkeypatch.setattr(
|
| 89 |
|
| 90 |
resp = await client.post(
|
| 91 |
"/serp/search_arxiv", json={"queries": ["a"]},
|
|
@@ -95,7 +96,7 @@ async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkey
|
|
| 95 |
|
| 96 |
|
| 97 |
async def test_bearer_scheme_match_is_case_insensitive(client, api_key, monkeypatch):
|
| 98 |
-
monkeypatch.setattr(
|
| 99 |
|
| 100 |
resp = await client.post(
|
| 101 |
"/serp/search_arxiv", json={"queries": ["a"]},
|
|
|
|
| 15 |
from httpx import ASGITransport
|
| 16 |
|
| 17 |
import app as app_module
|
| 18 |
+
import services as services_module
|
| 19 |
|
| 20 |
|
| 21 |
@pytest.fixture
|
|
|
|
| 48 |
"""The middleware is a no-op unless SERPENT_API_KEY is set - this keeps
|
| 49 |
the default deployment exactly as open as it was before it existed.
|
| 50 |
"""
|
| 51 |
+
monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
|
| 52 |
|
| 53 |
resp = await client.post("/serp/search_arxiv", json={"queries": ["a"]})
|
| 54 |
|
|
|
|
| 77 |
|
| 78 |
|
| 79 |
async def test_correct_key_via_x_api_key_header_is_accepted(client, api_key, monkeypatch):
|
| 80 |
+
monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
|
| 81 |
|
| 82 |
resp = await client.post(
|
| 83 |
"/serp/search_arxiv", json={"queries": ["a"]}, headers={"X-API-Key": api_key})
|
|
|
|
| 86 |
|
| 87 |
|
| 88 |
async def test_correct_key_via_bearer_header_is_accepted(client, api_key, monkeypatch):
|
| 89 |
+
monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
|
| 90 |
|
| 91 |
resp = await client.post(
|
| 92 |
"/serp/search_arxiv", json={"queries": ["a"]},
|
|
|
|
| 96 |
|
| 97 |
|
| 98 |
async def test_bearer_scheme_match_is_case_insensitive(client, api_key, monkeypatch):
|
| 99 |
+
monkeypatch.setattr(services_module, "query_arxiv", _stub_arxiv)
|
| 100 |
|
| 101 |
resp = await client.post(
|
| 102 |
"/serp/search_arxiv", json={"queries": ["a"]},
|
|
@@ -16,6 +16,7 @@ import app as app_module
|
|
| 16 |
import scrap as scrap_module
|
| 17 |
from app import MAX_BULK_PATENT_IDS, ScrapPatentsRequest
|
| 18 |
from serp import MAX_QUERIES_PER_REQUEST, SerpQuery
|
|
|
|
| 19 |
|
| 20 |
|
| 21 |
@pytest.fixture
|
|
@@ -57,9 +58,9 @@ async def test_oversized_query_list_returns_422(client):
|
|
| 57 |
|
| 58 |
def test_shape_serp_results_is_total_for_an_empty_list():
|
| 59 |
"""Defence in depth: the helper must not depend on its caller having
|
| 60 |
-
validated the query list, since it is shared by every search
|
| 61 |
"""
|
| 62 |
-
result =
|
| 63 |
|
| 64 |
assert result.results == []
|
| 65 |
assert result.error is not None
|
|
|
|
| 16 |
import scrap as scrap_module
|
| 17 |
from app import MAX_BULK_PATENT_IDS, ScrapPatentsRequest
|
| 18 |
from serp import MAX_QUERIES_PER_REQUEST, SerpQuery
|
| 19 |
+
from services import shape_serp_results
|
| 20 |
|
| 21 |
|
| 22 |
@pytest.fixture
|
|
|
|
| 58 |
|
| 59 |
def test_shape_serp_results_is_total_for_an_empty_list():
|
| 60 |
"""Defence in depth: the helper must not depend on its caller having
|
| 61 |
+
validated the query list, since it is shared by every search path.
|
| 62 |
"""
|
| 63 |
+
result = shape_serp_results([])
|
| 64 |
|
| 65 |
assert result.results == []
|
| 66 |
assert result.error is not None
|
|
@@ -1,13 +1,14 @@
|
|
| 1 |
-
"""The Google Patents -> EPO OPS fallback chain
|
| 2 |
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
|
|
|
| 6 |
|
| 7 |
-
The distinction these tests exist to protect: a 404 means "this patent
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
"""
|
| 12 |
|
| 13 |
import httpx
|
|
@@ -15,28 +16,26 @@ import pytest
|
|
| 15 |
from httpx import ASGITransport, HTTPStatusError, Request, Response
|
| 16 |
|
| 17 |
import app as app_module
|
|
|
|
| 18 |
from scrap import PatentScrapResult
|
|
|
|
|
|
|
| 19 |
|
| 20 |
PATENT_ID = "US11930446B2"
|
| 21 |
|
| 22 |
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_key", "k")
|
| 33 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_secret", "s")
|
| 34 |
|
| 35 |
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_key", None)
|
| 39 |
-
monkeypatch.setattr(app_module.ops_token_manager, "_secret", None)
|
| 40 |
|
| 41 |
|
| 42 |
def _http_error(status: int) -> HTTPStatusError:
|
|
@@ -51,10 +50,16 @@ def _raises(exc):
|
|
| 51 |
return _fn
|
| 52 |
|
| 53 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
# ------------------------------- the happy path -------------------------------
|
| 55 |
|
| 56 |
|
| 57 |
-
async def test_successful_scrape_never_touches_ops(
|
| 58 |
ops_called = False
|
| 59 |
|
| 60 |
async def fake_ops(*args, **kwargs):
|
|
@@ -63,133 +68,165 @@ async def test_successful_scrape_never_touches_ops(client, ops_configured, monke
|
|
| 63 |
return PatentScrapResult(title="from OPS")
|
| 64 |
|
| 65 |
monkeypatch.setattr(
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
monkeypatch.setattr(
|
| 69 |
|
| 70 |
-
|
| 71 |
|
| 72 |
-
assert
|
| 73 |
-
assert resp.json()["title"] == "from Google Patents"
|
| 74 |
assert ops_called is False
|
| 75 |
|
| 76 |
|
| 77 |
-
async def _ok(value):
|
| 78 |
-
return value
|
| 79 |
-
|
| 80 |
-
|
| 81 |
# ------------------------- genuine miss vs upstream failure -------------------------
|
| 82 |
|
| 83 |
|
| 84 |
-
async def
|
| 85 |
"""The one case where "not found" is the truth."""
|
| 86 |
-
monkeypatch.setattr(
|
| 87 |
-
|
| 88 |
-
resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
|
| 89 |
|
| 90 |
-
|
|
|
|
| 91 |
|
| 92 |
|
| 93 |
@pytest.mark.parametrize("status", [429, 500, 502, 503])
|
| 94 |
-
async def test_an_upstream_failure_is_not_reported_as_not_found(
|
| 95 |
-
client, ops_unconfigured, monkeypatch, status):
|
| 96 |
"""MCP_INSTRUCTIONS tells the model a "not found" patent is genuinely
|
| 97 |
absent everywhere and to move on rather than retrying. Reporting a
|
| 98 |
transient upstream failure that way teaches an agent - with the
|
| 99 |
server's explicit encouragement - that a real patent does not exist.
|
| 100 |
"""
|
| 101 |
-
monkeypatch.setattr(
|
| 102 |
|
| 103 |
-
|
|
|
|
| 104 |
|
| 105 |
-
assert
|
| 106 |
|
| 107 |
|
| 108 |
-
async def
|
| 109 |
monkeypatch.setattr(
|
| 110 |
-
|
| 111 |
|
| 112 |
-
|
|
|
|
| 113 |
|
| 114 |
-
assert
|
| 115 |
|
| 116 |
|
| 117 |
-
async def
|
| 118 |
-
client, ops_unconfigured, monkeypatch):
|
| 119 |
"""parse_patent_html raises ValueError when the page isn't a patent page
|
| 120 |
(an interstitial, or a markup change). That is our problem or theirs,
|
| 121 |
but it is not evidence the patent doesn't exist.
|
| 122 |
"""
|
| 123 |
monkeypatch.setattr(
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
|
| 127 |
|
| 128 |
-
|
|
|
|
| 129 |
|
| 130 |
|
| 131 |
# --------------------------------- the OPS fallback ---------------------------------
|
| 132 |
|
| 133 |
|
| 134 |
-
async def test_falls_back_to_ops_when_google_patents_fails(
|
| 135 |
-
monkeypatch.setattr(
|
| 136 |
monkeypatch.setattr(
|
| 137 |
-
|
| 138 |
-
lambda c, pid: _ok(PatentScrapResult(title="from OPS")))
|
| 139 |
|
| 140 |
-
|
| 141 |
|
| 142 |
-
assert
|
| 143 |
-
assert resp.json()["title"] == "from OPS"
|
| 144 |
|
| 145 |
|
| 146 |
-
async def
|
| 147 |
-
monkeypatch.setattr(
|
| 148 |
-
monkeypatch.setattr(
|
| 149 |
|
| 150 |
-
|
|
|
|
| 151 |
|
| 152 |
-
assert resp.status_code == 404
|
| 153 |
-
assert "not found" in resp.json()["detail"].lower()
|
| 154 |
|
| 155 |
-
|
| 156 |
-
async def test_ops_failing_after_a_google_patents_404_is_not_a_404(
|
| 157 |
-
client, ops_configured, monkeypatch):
|
| 158 |
"""Google Patents says the patent is missing, but OPS - the backend that
|
| 159 |
covers what Google Patents doesn't - never answered. That is unknown,
|
| 160 |
not absent.
|
| 161 |
"""
|
| 162 |
-
monkeypatch.setattr(
|
| 163 |
-
monkeypatch.setattr(
|
| 164 |
|
| 165 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 166 |
|
| 167 |
-
assert resp.status_code == 502
|
| 168 |
|
|
|
|
|
|
|
|
|
|
| 169 |
|
| 170 |
-
|
|
|
|
| 171 |
|
| 172 |
|
| 173 |
-
async def
|
| 174 |
-
client, ops_configured, monkeypatch):
|
| 175 |
monkeypatch.setattr(
|
| 176 |
-
|
| 177 |
|
| 178 |
-
|
|
|
|
| 179 |
|
| 180 |
-
assert
|
| 181 |
|
| 182 |
|
| 183 |
-
|
| 184 |
-
"""Here OPS is the only backend asked, so its 404 is the whole answer."""
|
| 185 |
-
monkeypatch.setattr(app_module, "ops_scrap_patent", _raises(_http_error(404)))
|
| 186 |
|
| 187 |
-
resp = await client.get(f"/ops/scrap_patent/{PATENT_ID}")
|
| 188 |
|
| 189 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 190 |
|
|
|
|
|
|
|
|
|
|
| 191 |
|
| 192 |
-
async def
|
| 193 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 194 |
|
| 195 |
-
assert resp.status_code ==
|
|
|
|
|
|
| 1 |
+
"""The Google Patents -> EPO OPS fallback chain.
|
| 2 |
|
| 3 |
+
The policy lives in services.PatentService, so most of this is plain unit
|
| 4 |
+
testing with no ASGI client involved. Only the last section goes through
|
| 5 |
+
the app, and only to check that the domain errors the service raises map
|
| 6 |
+
onto the right status codes.
|
| 7 |
|
| 8 |
+
The distinction these tests exist to protect: a 404 means "this patent does
|
| 9 |
+
not exist", and an agent is explicitly told by MCP_INSTRUCTIONS to believe
|
| 10 |
+
it and move on without retrying. Anything else - an upstream outage, a
|
| 11 |
+
timeout, a blocked scrape - must not be reported that way.
|
| 12 |
"""
|
| 13 |
|
| 14 |
import httpx
|
|
|
|
| 16 |
from httpx import ASGITransport, HTTPStatusError, Request, Response
|
| 17 |
|
| 18 |
import app as app_module
|
| 19 |
+
import services as services_module
|
| 20 |
from scrap import PatentScrapResult
|
| 21 |
+
from services import (OPSUnconfigured, PatentNotFound, PatentService,
|
| 22 |
+
UpstreamUnavailable)
|
| 23 |
|
| 24 |
PATENT_ID = "US11930446B2"
|
| 25 |
|
| 26 |
|
| 27 |
+
class FakeOPSTokens:
|
| 28 |
+
"""Stand-in for the OPS token manager. Tests only care whether
|
| 29 |
+
credentials are configured, and this avoids assigning to private
|
| 30 |
+
attributes on the real module-level singleton.
|
| 31 |
+
"""
|
|
|
|
| 32 |
|
| 33 |
+
def __init__(self, configured: bool):
|
| 34 |
+
self.configured = configured
|
|
|
|
|
|
|
| 35 |
|
| 36 |
|
| 37 |
+
def make_service(*, ops_configured: bool) -> PatentService:
|
| 38 |
+
return PatentService(http_client=None, ops_tokens=FakeOPSTokens(ops_configured))
|
|
|
|
|
|
|
| 39 |
|
| 40 |
|
| 41 |
def _http_error(status: int) -> HTTPStatusError:
|
|
|
|
| 50 |
return _fn
|
| 51 |
|
| 52 |
|
| 53 |
+
def _returns(value):
|
| 54 |
+
async def _fn(*args, **kwargs):
|
| 55 |
+
return value
|
| 56 |
+
return _fn
|
| 57 |
+
|
| 58 |
+
|
| 59 |
# ------------------------------- the happy path -------------------------------
|
| 60 |
|
| 61 |
|
| 62 |
+
async def test_successful_scrape_never_touches_ops(monkeypatch):
|
| 63 |
ops_called = False
|
| 64 |
|
| 65 |
async def fake_ops(*args, **kwargs):
|
|
|
|
| 68 |
return PatentScrapResult(title="from OPS")
|
| 69 |
|
| 70 |
monkeypatch.setattr(
|
| 71 |
+
services_module, "scrap_patent_async",
|
| 72 |
+
_returns(PatentScrapResult(title="from Google Patents")))
|
| 73 |
+
monkeypatch.setattr(services_module, "ops_scrap_patent", fake_ops)
|
| 74 |
|
| 75 |
+
result = await make_service(ops_configured=True).scrap(PATENT_ID)
|
| 76 |
|
| 77 |
+
assert result.title == "from Google Patents"
|
|
|
|
| 78 |
assert ops_called is False
|
| 79 |
|
| 80 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 81 |
# ------------------------- genuine miss vs upstream failure -------------------------
|
| 82 |
|
| 83 |
|
| 84 |
+
async def test_a_real_404_from_google_patents_is_not_found(monkeypatch):
|
| 85 |
"""The one case where "not found" is the truth."""
|
| 86 |
+
monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
|
|
|
|
|
|
|
| 87 |
|
| 88 |
+
with pytest.raises(PatentNotFound):
|
| 89 |
+
await make_service(ops_configured=False).scrap(PATENT_ID)
|
| 90 |
|
| 91 |
|
| 92 |
@pytest.mark.parametrize("status", [429, 500, 502, 503])
|
| 93 |
+
async def test_an_upstream_failure_is_not_reported_as_not_found(monkeypatch, status):
|
|
|
|
| 94 |
"""MCP_INSTRUCTIONS tells the model a "not found" patent is genuinely
|
| 95 |
absent everywhere and to move on rather than retrying. Reporting a
|
| 96 |
transient upstream failure that way teaches an agent - with the
|
| 97 |
server's explicit encouragement - that a real patent does not exist.
|
| 98 |
"""
|
| 99 |
+
monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(status)))
|
| 100 |
|
| 101 |
+
with pytest.raises(UpstreamUnavailable) as exc_info:
|
| 102 |
+
await make_service(ops_configured=False).scrap(PATENT_ID)
|
| 103 |
|
| 104 |
+
assert exc_info.value.timeout is False
|
| 105 |
|
| 106 |
|
| 107 |
+
async def test_a_timeout_is_flagged_as_a_timeout(monkeypatch):
|
| 108 |
monkeypatch.setattr(
|
| 109 |
+
services_module, "scrap_patent_async", _raises(httpx.ConnectTimeout("timed out")))
|
| 110 |
|
| 111 |
+
with pytest.raises(UpstreamUnavailable) as exc_info:
|
| 112 |
+
await make_service(ops_configured=False).scrap(PATENT_ID)
|
| 113 |
|
| 114 |
+
assert exc_info.value.timeout is True
|
| 115 |
|
| 116 |
|
| 117 |
+
async def test_an_unparseable_page_is_an_upstream_failure(monkeypatch):
|
|
|
|
| 118 |
"""parse_patent_html raises ValueError when the page isn't a patent page
|
| 119 |
(an interstitial, or a markup change). That is our problem or theirs,
|
| 120 |
but it is not evidence the patent doesn't exist.
|
| 121 |
"""
|
| 122 |
monkeypatch.setattr(
|
| 123 |
+
services_module, "scrap_patent_async", _raises(ValueError("no DC.title")))
|
|
|
|
|
|
|
| 124 |
|
| 125 |
+
with pytest.raises(UpstreamUnavailable):
|
| 126 |
+
await make_service(ops_configured=False).scrap(PATENT_ID)
|
| 127 |
|
| 128 |
|
| 129 |
# --------------------------------- the OPS fallback ---------------------------------
|
| 130 |
|
| 131 |
|
| 132 |
+
async def test_falls_back_to_ops_when_google_patents_fails(monkeypatch):
|
| 133 |
+
monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(503)))
|
| 134 |
monkeypatch.setattr(
|
| 135 |
+
services_module, "ops_scrap_patent", _returns(PatentScrapResult(title="from OPS")))
|
|
|
|
| 136 |
|
| 137 |
+
result = await make_service(ops_configured=True).scrap(PATENT_ID)
|
| 138 |
|
| 139 |
+
assert result.title == "from OPS"
|
|
|
|
| 140 |
|
| 141 |
|
| 142 |
+
async def test_missing_from_both_backends_is_not_found(monkeypatch):
|
| 143 |
+
monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
|
| 144 |
+
monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(404)))
|
| 145 |
|
| 146 |
+
with pytest.raises(PatentNotFound):
|
| 147 |
+
await make_service(ops_configured=True).scrap(PATENT_ID)
|
| 148 |
|
|
|
|
|
|
|
| 149 |
|
| 150 |
+
async def test_ops_failing_after_a_google_patents_404_is_not_a_miss(monkeypatch):
|
|
|
|
|
|
|
| 151 |
"""Google Patents says the patent is missing, but OPS - the backend that
|
| 152 |
covers what Google Patents doesn't - never answered. That is unknown,
|
| 153 |
not absent.
|
| 154 |
"""
|
| 155 |
+
monkeypatch.setattr(services_module, "scrap_patent_async", _raises(_http_error(404)))
|
| 156 |
+
monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(500)))
|
| 157 |
|
| 158 |
+
with pytest.raises(UpstreamUnavailable):
|
| 159 |
+
await make_service(ops_configured=True).scrap(PATENT_ID)
|
| 160 |
+
|
| 161 |
+
|
| 162 |
+
# ------------------------------ the OPS-only path ------------------------------
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
async def test_ops_only_retrieval_requires_credentials():
|
| 166 |
+
with pytest.raises(OPSUnconfigured):
|
| 167 |
+
await make_service(ops_configured=False).ops_scrap(PATENT_ID)
|
| 168 |
|
|
|
|
| 169 |
|
| 170 |
+
async def test_ops_only_retrieval_reports_a_real_miss(monkeypatch):
|
| 171 |
+
"""Here OPS is the only backend asked, so its 404 is the whole answer."""
|
| 172 |
+
monkeypatch.setattr(services_module, "ops_scrap_patent", _raises(_http_error(404)))
|
| 173 |
|
| 174 |
+
with pytest.raises(PatentNotFound):
|
| 175 |
+
await make_service(ops_configured=True).ops_scrap(PATENT_ID)
|
| 176 |
|
| 177 |
|
| 178 |
+
async def test_ops_only_retrieval_flags_a_timeout(monkeypatch):
|
|
|
|
| 179 |
monkeypatch.setattr(
|
| 180 |
+
services_module, "ops_scrap_patent", _raises(httpx.ConnectTimeout("timed out")))
|
| 181 |
|
| 182 |
+
with pytest.raises(UpstreamUnavailable) as exc_info:
|
| 183 |
+
await make_service(ops_configured=True).ops_scrap(PATENT_ID)
|
| 184 |
|
| 185 |
+
assert exc_info.value.timeout is True
|
| 186 |
|
| 187 |
|
| 188 |
+
# ---------------------------- domain error -> status code ----------------------------
|
|
|
|
|
|
|
| 189 |
|
|
|
|
| 190 |
|
| 191 |
+
@pytest.fixture
|
| 192 |
+
async def client():
|
| 193 |
+
transport = ASGITransport(app=app_module.app)
|
| 194 |
+
async with httpx.AsyncClient(transport=transport, base_url="http://test") as c:
|
| 195 |
+
yield c
|
| 196 |
+
|
| 197 |
+
|
| 198 |
+
@pytest.fixture
|
| 199 |
+
def override_patent_service():
|
| 200 |
+
"""Swap the app's patent service for one the test controls, then put the
|
| 201 |
+
real provider back."""
|
| 202 |
+
def _install(service):
|
| 203 |
+
app_module.app.dependency_overrides[app_module.get_patent_service] = lambda: service
|
| 204 |
+
|
| 205 |
+
yield _install
|
| 206 |
+
app_module.app.dependency_overrides.clear()
|
| 207 |
+
|
| 208 |
|
| 209 |
+
class _StubService:
|
| 210 |
+
def __init__(self, error):
|
| 211 |
+
self._error = error
|
| 212 |
|
| 213 |
+
async def scrap(self, patent_id):
|
| 214 |
+
raise self._error
|
| 215 |
+
|
| 216 |
+
ops_scrap = scrap
|
| 217 |
+
|
| 218 |
+
|
| 219 |
+
@pytest.mark.parametrize("error, expected_status", [
|
| 220 |
+
(PatentNotFound("gone"), 404),
|
| 221 |
+
(UpstreamUnavailable("upstream broke"), 502),
|
| 222 |
+
(UpstreamUnavailable("slow", timeout=True), 504),
|
| 223 |
+
(OPSUnconfigured("no credentials"), 503),
|
| 224 |
+
])
|
| 225 |
+
async def test_domain_errors_map_onto_status_codes(
|
| 226 |
+
client, override_patent_service, error, expected_status):
|
| 227 |
+
override_patent_service(_StubService(error))
|
| 228 |
+
|
| 229 |
+
resp = await client.get(f"/scrap/scrap_patent/{PATENT_ID}")
|
| 230 |
|
| 231 |
+
assert resp.status_code == expected_status
|
| 232 |
+
assert resp.json()["detail"] == str(error)
|
|
@@ -89,3 +89,15 @@ async def test_scrap_patent_bulk_async_separates_successes_from_failures():
|
|
| 89 |
assert len(result.patents) == 1
|
| 90 |
assert result.patents[0].title == "Widget with improved gadget mechanism"
|
| 91 |
assert result.failed_ids == ["US_FAIL"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 89 |
assert len(result.patents) == 1
|
| 90 |
assert result.patents[0].title == "Widget with improved gadget mechanism"
|
| 91 |
assert result.failed_ids == ["US_FAIL"]
|
| 92 |
+
|
| 93 |
+
|
| 94 |
+
def test_parse_patent_html_strips_whitespace_around_the_title():
|
| 95 |
+
"""Google Patents' DC.title meta content is frequently padded with
|
| 96 |
+
newlines and indentation from the surrounding markup, which would
|
| 97 |
+
otherwise end up in the API response and in every downstream citation.
|
| 98 |
+
"""
|
| 99 |
+
html = '<html><head><meta name="DC.title" content=" Widget apparatus\n "></head><body></body></html>'
|
| 100 |
+
|
| 101 |
+
result = parse_patent_html(html, "https://patents.google.com/patent/US1/en")
|
| 102 |
+
|
| 103 |
+
assert result.title == "Widget apparatus"
|
|
@@ -0,0 +1,390 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""SearchService: the backend orchestration policy.
|
| 2 |
+
|
| 3 |
+
Which backend to try, in what order, when to fall back, and what counts as
|
| 4 |
+
evidence a backend is unhealthy. None of this needs an HTTP layer, so none
|
| 5 |
+
of these tests use one.
|
| 6 |
+
|
| 7 |
+
Each test builds a service with its own circuit breaker. That replaces the
|
| 8 |
+
autouse fixture that used to reset a process-wide singleton between tests -
|
| 9 |
+
the breaker being shared across every request is correct in production and
|
| 10 |
+
was only a problem because tests had no way to opt out of it.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
|
| 14 |
+
import services as services_module
|
| 15 |
+
from circuit_breaker import CircuitBreaker
|
| 16 |
+
from serp import DuckDuckGoBlockedException, SerpQuery
|
| 17 |
+
from services import SearchService, shape_serp_results
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
class FakeOPSTokens:
|
| 21 |
+
def __init__(self, configured: bool):
|
| 22 |
+
self.configured = configured
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def make_service(*, ops_configured: bool = False, breaker: CircuitBreaker = None,
|
| 26 |
+
browser=None) -> SearchService:
|
| 27 |
+
return SearchService(
|
| 28 |
+
http_client=None,
|
| 29 |
+
browser_provider=lambda: browser,
|
| 30 |
+
circuit_breaker=breaker or CircuitBreaker(),
|
| 31 |
+
ops_tokens=FakeOPSTokens(ops_configured),
|
| 32 |
+
)
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
# --------------------------------- result shaping ---------------------------------
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
class TestShapeSerpResults:
|
| 39 |
+
def test_flattens_successful_lists(self):
|
| 40 |
+
result = shape_serp_results([[{"a": 1}], [{"b": 2}, {"c": 3}]])
|
| 41 |
+
assert result.results == [{"a": 1}, {"b": 2}, {"c": 3}]
|
| 42 |
+
assert result.error is None
|
| 43 |
+
|
| 44 |
+
def test_surfaces_last_error_when_all_failed(self):
|
| 45 |
+
result = shape_serp_results([ValueError("first"), RuntimeError("second")])
|
| 46 |
+
assert result.results == []
|
| 47 |
+
assert "second" in result.error
|
| 48 |
+
|
| 49 |
+
def test_ignores_failures_when_some_succeed(self):
|
| 50 |
+
result = shape_serp_results([[{"a": 1}], RuntimeError("boom")])
|
| 51 |
+
assert result.results == [{"a": 1}]
|
| 52 |
+
assert result.error is None
|
| 53 |
+
|
| 54 |
+
def test_is_total_for_an_empty_list(self):
|
| 55 |
+
"""Defence in depth: the request models bound the query list, but
|
| 56 |
+
this helper is shared by every search path and must not depend on
|
| 57 |
+
its caller having validated anything."""
|
| 58 |
+
result = shape_serp_results([])
|
| 59 |
+
assert result.results == []
|
| 60 |
+
assert result.error is not None
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
# ------------------------------- per-query dispatch -------------------------------
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
async def test_every_query_is_dispatched_with_the_result_count():
|
| 67 |
+
calls = []
|
| 68 |
+
|
| 69 |
+
async def fn(q, n):
|
| 70 |
+
calls.append((q, n))
|
| 71 |
+
return [{"q": q}]
|
| 72 |
+
|
| 73 |
+
result = await make_service()._run_queries(
|
| 74 |
+
fn, SerpQuery(queries=["x", "y"], n_results=5), "test context")
|
| 75 |
+
|
| 76 |
+
assert sorted(calls) == [("x", 5), ("y", 5)]
|
| 77 |
+
assert sorted(result.results, key=str) == sorted([{"q": "x"}, {"q": "y"}], key=str)
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
async def test_arxiv_flattens_results_from_all_queries(monkeypatch):
|
| 81 |
+
async def fake_query_arxiv(client, q, n):
|
| 82 |
+
return [{"title": f"paper about {q}"}]
|
| 83 |
+
|
| 84 |
+
monkeypatch.setattr(services_module, "query_arxiv", fake_query_arxiv)
|
| 85 |
+
|
| 86 |
+
result = await make_service().arxiv(SerpQuery(queries=["a", "b"], n_results=5))
|
| 87 |
+
|
| 88 |
+
assert {r["title"] for r in result.results} == {"paper about a", "paper about b"}
|
| 89 |
+
assert result.error is None
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
async def test_arxiv_reports_an_error_when_every_query_fails(monkeypatch):
|
| 93 |
+
async def failing(client, q, n):
|
| 94 |
+
raise RuntimeError(f"boom for {q}")
|
| 95 |
+
|
| 96 |
+
monkeypatch.setattr(services_module, "query_arxiv", failing)
|
| 97 |
+
|
| 98 |
+
result = await make_service().arxiv(SerpQuery(queries=["a"], n_results=5))
|
| 99 |
+
|
| 100 |
+
assert result.results == []
|
| 101 |
+
assert "boom for a" in result.error
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
# ------------------------------ patents + OPS fallback ------------------------------
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
async def test_patents_does_not_fall_back_when_google_patents_succeeds(monkeypatch):
|
| 108 |
+
ops_called = False
|
| 109 |
+
|
| 110 |
+
async def fake_ops(client, q, n):
|
| 111 |
+
nonlocal ops_called
|
| 112 |
+
ops_called = True
|
| 113 |
+
return []
|
| 114 |
+
|
| 115 |
+
monkeypatch.setattr(
|
| 116 |
+
services_module, "query_google_patents",
|
| 117 |
+
lambda b, q, n: _ok([{"id": "US1", "title": "t"}]))
|
| 118 |
+
monkeypatch.setattr(services_module, "ops_search", fake_ops)
|
| 119 |
+
|
| 120 |
+
result = await make_service(ops_configured=True).patents(
|
| 121 |
+
SerpQuery(queries=["widget"], n_results=5))
|
| 122 |
+
|
| 123 |
+
assert result.results == [{"id": "US1", "title": "t"}]
|
| 124 |
+
assert ops_called is False
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
async def _ok(value):
|
| 128 |
+
return value
|
| 129 |
+
|
| 130 |
+
|
| 131 |
+
async def test_patents_falls_back_to_ops_when_google_patents_is_empty(monkeypatch):
|
| 132 |
+
monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
|
| 133 |
+
monkeypatch.setattr(
|
| 134 |
+
services_module, "ops_search", lambda c, q, n: _ok([{"id": "US123", "title": "OPS"}]))
|
| 135 |
+
|
| 136 |
+
result = await make_service(ops_configured=True).patents(
|
| 137 |
+
SerpQuery(queries=["widget"], n_results=5))
|
| 138 |
+
|
| 139 |
+
assert result.results == [{"id": "US123", "title": "OPS"}]
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
async def test_patents_does_not_fall_back_when_ops_is_unconfigured(monkeypatch):
|
| 143 |
+
monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
|
| 144 |
+
|
| 145 |
+
result = await make_service(ops_configured=False).patents(
|
| 146 |
+
SerpQuery(queries=["widget"], n_results=5))
|
| 147 |
+
|
| 148 |
+
assert result.results == []
|
| 149 |
+
assert result.error is None
|
| 150 |
+
|
| 151 |
+
|
| 152 |
+
async def test_ops_fallback_runs_concurrently_not_one_query_at_a_time(monkeypatch):
|
| 153 |
+
"""This path exists for when Google Patents is unavailable, which is
|
| 154 |
+
exactly when a request can least afford N sequential round-trips to the
|
| 155 |
+
slower backend."""
|
| 156 |
+
import asyncio
|
| 157 |
+
|
| 158 |
+
in_flight = 0
|
| 159 |
+
peak = 0
|
| 160 |
+
|
| 161 |
+
async def slow_ops(client, q, n):
|
| 162 |
+
nonlocal in_flight, peak
|
| 163 |
+
in_flight += 1
|
| 164 |
+
peak = max(peak, in_flight)
|
| 165 |
+
await asyncio.sleep(0.02)
|
| 166 |
+
in_flight -= 1
|
| 167 |
+
return [{"id": q}]
|
| 168 |
+
|
| 169 |
+
monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
|
| 170 |
+
monkeypatch.setattr(services_module, "ops_search", slow_ops)
|
| 171 |
+
|
| 172 |
+
await make_service(ops_configured=True).patents(
|
| 173 |
+
SerpQuery(queries=["a", "b", "c", "d"], n_results=5))
|
| 174 |
+
|
| 175 |
+
assert peak > 1, "OPS fallback queries were issued one at a time"
|
| 176 |
+
|
| 177 |
+
|
| 178 |
+
async def test_a_failing_ops_fallback_does_not_lose_other_queries_results(monkeypatch):
|
| 179 |
+
async def selective_ops(client, q, n):
|
| 180 |
+
if q == "bad":
|
| 181 |
+
raise RuntimeError("OPS down")
|
| 182 |
+
return [{"id": q}]
|
| 183 |
+
|
| 184 |
+
monkeypatch.setattr(services_module, "query_google_patents", lambda b, q, n: _ok([]))
|
| 185 |
+
monkeypatch.setattr(services_module, "ops_search", selective_ops)
|
| 186 |
+
|
| 187 |
+
result = await make_service(ops_configured=True).patents(
|
| 188 |
+
SerpQuery(queries=["good", "bad"], n_results=5))
|
| 189 |
+
|
| 190 |
+
assert result.results == [{"id": "good"}]
|
| 191 |
+
|
| 192 |
+
|
| 193 |
+
# -------------------------------- the fallback chain --------------------------------
|
| 194 |
+
|
| 195 |
+
|
| 196 |
+
async def test_returns_the_first_backend_that_succeeds(monkeypatch):
|
| 197 |
+
calls = []
|
| 198 |
+
|
| 199 |
+
async def fake_ddg(q, n):
|
| 200 |
+
calls.append("ddg")
|
| 201 |
+
return [{"title": "from ddg"}]
|
| 202 |
+
|
| 203 |
+
async def fake_brave(browser, q, n):
|
| 204 |
+
calls.append("brave")
|
| 205 |
+
return [{"title": "from brave"}]
|
| 206 |
+
|
| 207 |
+
monkeypatch.setattr(services_module, "query_ddg_search", fake_ddg)
|
| 208 |
+
monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
|
| 209 |
+
|
| 210 |
+
q, results, err = await make_service()._search_one("widget", 5)
|
| 211 |
+
|
| 212 |
+
assert calls == ["ddg"]
|
| 213 |
+
assert results == [{"title": "from ddg"}]
|
| 214 |
+
assert err is None
|
| 215 |
+
|
| 216 |
+
|
| 217 |
+
async def test_falls_through_to_the_next_backend_on_failure(monkeypatch):
|
| 218 |
+
calls = []
|
| 219 |
+
|
| 220 |
+
async def failing_ddg(q, n):
|
| 221 |
+
calls.append("ddg")
|
| 222 |
+
raise RuntimeError("ddg blocked")
|
| 223 |
+
|
| 224 |
+
async def fake_brave(browser, q, n):
|
| 225 |
+
calls.append("brave")
|
| 226 |
+
return [{"title": "from brave"}]
|
| 227 |
+
|
| 228 |
+
monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
|
| 229 |
+
monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
|
| 230 |
+
|
| 231 |
+
q, results, err = await make_service()._search_one("widget", 5)
|
| 232 |
+
|
| 233 |
+
assert calls == ["ddg", "brave"]
|
| 234 |
+
assert results == [{"title": "from brave"}]
|
| 235 |
+
assert err is None
|
| 236 |
+
|
| 237 |
+
|
| 238 |
+
async def test_reports_an_error_when_every_backend_fails(monkeypatch):
|
| 239 |
+
async def failing(*args):
|
| 240 |
+
raise RuntimeError("blocked")
|
| 241 |
+
|
| 242 |
+
for name in ("query_ddg_search", "query_brave_search", "query_bing_search"):
|
| 243 |
+
monkeypatch.setattr(services_module, name, failing)
|
| 244 |
+
|
| 245 |
+
q, results, err = await make_service()._search_one("widget", 5)
|
| 246 |
+
|
| 247 |
+
assert results == []
|
| 248 |
+
assert "All backends failed for query 'widget'" in err
|
| 249 |
+
assert "blocked" in err
|
| 250 |
+
|
| 251 |
+
|
| 252 |
+
# ------------------------------ the circuit breaker ------------------------------
|
| 253 |
+
|
| 254 |
+
|
| 255 |
+
async def test_skips_a_backend_whose_circuit_is_open(monkeypatch):
|
| 256 |
+
"""A backend that's failed `failure_threshold` times in a row shouldn't
|
| 257 |
+
be attempted again until its cooldown elapses - repeatedly hammering a
|
| 258 |
+
backend that's already rate-limiting us only makes that worse.
|
| 259 |
+
"""
|
| 260 |
+
ddg_calls = []
|
| 261 |
+
|
| 262 |
+
async def failing_ddg(q, n):
|
| 263 |
+
ddg_calls.append(q)
|
| 264 |
+
raise RuntimeError("ddg blocked")
|
| 265 |
+
|
| 266 |
+
async def fake_brave(browser, q, n):
|
| 267 |
+
return [{"title": "from brave"}]
|
| 268 |
+
|
| 269 |
+
monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
|
| 270 |
+
monkeypatch.setattr(services_module, "query_brave_search", fake_brave)
|
| 271 |
+
|
| 272 |
+
service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
|
| 273 |
+
|
| 274 |
+
await service._search_one("first query", 5)
|
| 275 |
+
assert ddg_calls == ["first query"]
|
| 276 |
+
|
| 277 |
+
q, results, err = await service._search_one("second query", 5)
|
| 278 |
+
|
| 279 |
+
assert ddg_calls == ["first query"] # unchanged - DDG was skipped
|
| 280 |
+
assert results == [{"title": "from brave"}]
|
| 281 |
+
|
| 282 |
+
|
| 283 |
+
async def test_a_zero_result_query_does_not_open_the_circuit(monkeypatch):
|
| 284 |
+
"""An empty result is ambiguous - it can mean the query genuinely has no
|
| 285 |
+
hits, or that the backend served an anti-bot page that parsed to
|
| 286 |
+
nothing. The breaker is shared across every request, so counting the
|
| 287 |
+
ambiguous case meant a few unusual-but-legitimate queries could disable
|
| 288 |
+
the primary backend for everyone for a full cooldown.
|
| 289 |
+
"""
|
| 290 |
+
ddg_calls = []
|
| 291 |
+
|
| 292 |
+
async def empty_ddg(q, n):
|
| 293 |
+
ddg_calls.append(q)
|
| 294 |
+
raise DuckDuckGoBlockedException()
|
| 295 |
+
|
| 296 |
+
monkeypatch.setattr(services_module, "query_ddg_search", empty_ddg)
|
| 297 |
+
monkeypatch.setattr(
|
| 298 |
+
services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "brave"}]))
|
| 299 |
+
|
| 300 |
+
service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
|
| 301 |
+
|
| 302 |
+
await service._search_one("obscure", 5)
|
| 303 |
+
await service._search_one("another obscure", 5)
|
| 304 |
+
|
| 305 |
+
assert ddg_calls == ["obscure", "another obscure"]
|
| 306 |
+
|
| 307 |
+
|
| 308 |
+
async def test_an_empty_result_still_falls_through_to_the_next_backend(monkeypatch):
|
| 309 |
+
"""Not counting it against the backend must not stop the fallback: the
|
| 310 |
+
caller still needs results from somewhere."""
|
| 311 |
+
async def empty_ddg(q, n):
|
| 312 |
+
raise DuckDuckGoBlockedException()
|
| 313 |
+
|
| 314 |
+
monkeypatch.setattr(services_module, "query_ddg_search", empty_ddg)
|
| 315 |
+
monkeypatch.setattr(
|
| 316 |
+
services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "from brave"}]))
|
| 317 |
+
|
| 318 |
+
_q, results, err = await make_service()._search_one("widget", 5)
|
| 319 |
+
|
| 320 |
+
assert results == [{"title": "from brave"}]
|
| 321 |
+
assert err is None
|
| 322 |
+
|
| 323 |
+
|
| 324 |
+
async def test_a_real_backend_error_still_opens_the_circuit(monkeypatch):
|
| 325 |
+
"""The distinction has to cut both ways: a backend raising a genuine
|
| 326 |
+
error (a rate limit, a transport failure) is still evidence."""
|
| 327 |
+
ddg_calls = []
|
| 328 |
+
|
| 329 |
+
async def failing_ddg(q, n):
|
| 330 |
+
ddg_calls.append(q)
|
| 331 |
+
raise RuntimeError("rate limited")
|
| 332 |
+
|
| 333 |
+
monkeypatch.setattr(services_module, "query_ddg_search", failing_ddg)
|
| 334 |
+
monkeypatch.setattr(
|
| 335 |
+
services_module, "query_brave_search", lambda b, q, n: _ok([{"title": "brave"}]))
|
| 336 |
+
|
| 337 |
+
service = make_service(breaker=CircuitBreaker(failure_threshold=1, cooldown_seconds=60))
|
| 338 |
+
|
| 339 |
+
await service._search_one("first", 5)
|
| 340 |
+
await service._search_one("second", 5)
|
| 341 |
+
|
| 342 |
+
assert ddg_calls == ["first"]
|
| 343 |
+
|
| 344 |
+
|
| 345 |
+
# ------------------------------- search aggregation -------------------------------
|
| 346 |
+
|
| 347 |
+
|
| 348 |
+
async def test_search_merges_results_from_every_query(monkeypatch):
|
| 349 |
+
monkeypatch.setattr(
|
| 350 |
+
services_module, "query_ddg_search", lambda q, n: _ok([{"title": f"result for {q}"}]))
|
| 351 |
+
|
| 352 |
+
result = await make_service().search(SerpQuery(queries=["a", "b"], n_results=5))
|
| 353 |
+
|
| 354 |
+
assert {r["title"] for r in result.results} == {"result for a", "result for b"}
|
| 355 |
+
assert result.error is None
|
| 356 |
+
|
| 357 |
+
|
| 358 |
+
async def test_search_reports_partial_failure_alongside_results(monkeypatch):
|
| 359 |
+
"""A query no backend could answer must be reported even when another
|
| 360 |
+
query succeeded - otherwise the caller silently receives fewer results
|
| 361 |
+
than they asked for with no indication why."""
|
| 362 |
+
async def selective_ddg(q, n):
|
| 363 |
+
if q == "bad":
|
| 364 |
+
raise RuntimeError("blocked")
|
| 365 |
+
return [{"title": f"result for {q}"}]
|
| 366 |
+
|
| 367 |
+
async def failing(*args):
|
| 368 |
+
raise RuntimeError("blocked")
|
| 369 |
+
|
| 370 |
+
monkeypatch.setattr(services_module, "query_ddg_search", selective_ddg)
|
| 371 |
+
monkeypatch.setattr(services_module, "query_brave_search", failing)
|
| 372 |
+
monkeypatch.setattr(services_module, "query_bing_search", failing)
|
| 373 |
+
|
| 374 |
+
result = await make_service().search(SerpQuery(queries=["good", "bad"], n_results=5))
|
| 375 |
+
|
| 376 |
+
assert [r["title"] for r in result.results] == ["result for good"]
|
| 377 |
+
assert "bad" in result.error
|
| 378 |
+
|
| 379 |
+
|
| 380 |
+
async def test_search_names_every_failed_query_not_just_the_last(monkeypatch):
|
| 381 |
+
async def failing(*args):
|
| 382 |
+
raise RuntimeError("blocked")
|
| 383 |
+
|
| 384 |
+
for name in ("query_ddg_search", "query_brave_search", "query_bing_search"):
|
| 385 |
+
monkeypatch.setattr(services_module, name, failing)
|
| 386 |
+
|
| 387 |
+
result = await make_service().search(SerpQuery(queries=["a", "b"], n_results=5))
|
| 388 |
+
|
| 389 |
+
assert result.results == []
|
| 390 |
+
assert "'a'" in result.error and "'b'" in result.error
|
|
@@ -0,0 +1,156 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The glue holding each scraper together.
|
| 2 |
+
|
| 3 |
+
Every Playwright scraper is now: build a URL, open a page, block
|
| 4 |
+
decorative resources, navigate, extract. The builders are tested in
|
| 5 |
+
test_serp_urls.py and the extractors in test_serp_playwright.py; this
|
| 6 |
+
covers the composition - specifically that each scraper navigates to the
|
| 7 |
+
URL its own builder produced.
|
| 8 |
+
|
| 9 |
+
That link is the one the old goto-patching harness could not test, because
|
| 10 |
+
the stand-in it installed accepted the navigation URL and discarded it.
|
| 11 |
+
These use a recording stand-in browser instead: no real Chromium, and the
|
| 12 |
+
URL is the thing being asserted on.
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
import pytest
|
| 16 |
+
|
| 17 |
+
import serp
|
| 18 |
+
from serp import (bing_search_url, brave_search_url, google_patents_search_url,
|
| 19 |
+
google_scholar_url, playwright_open_page,
|
| 20 |
+
query_bing_search, query_brave_search, query_google_patents,
|
| 21 |
+
query_google_scholar)
|
| 22 |
+
|
| 23 |
+
SENTINEL = [{"title": "extracted"}]
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
class _RecordingPage:
|
| 27 |
+
def __init__(self):
|
| 28 |
+
self.goto_urls = []
|
| 29 |
+
self.route_patterns = []
|
| 30 |
+
self.closed = False
|
| 31 |
+
|
| 32 |
+
async def route(self, pattern, handler):
|
| 33 |
+
self.route_patterns.append(pattern)
|
| 34 |
+
|
| 35 |
+
async def goto(self, url, **kwargs):
|
| 36 |
+
self.goto_urls.append(url)
|
| 37 |
+
return None
|
| 38 |
+
|
| 39 |
+
async def close(self):
|
| 40 |
+
self.closed = True
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
class _RecordingContext:
|
| 44 |
+
def __init__(self, page):
|
| 45 |
+
self._page = page
|
| 46 |
+
self.closed = False
|
| 47 |
+
|
| 48 |
+
async def new_page(self):
|
| 49 |
+
return self._page
|
| 50 |
+
|
| 51 |
+
async def close(self):
|
| 52 |
+
self.closed = True
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
class _RecordingBrowser:
|
| 56 |
+
def __init__(self):
|
| 57 |
+
self.page = _RecordingPage()
|
| 58 |
+
self.context = _RecordingContext(self.page)
|
| 59 |
+
|
| 60 |
+
async def new_context(self, **kwargs):
|
| 61 |
+
return self.context
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
async def _sentinel_extractor(page, n_results):
|
| 65 |
+
return SENTINEL
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
SCRAPERS = [
|
| 69 |
+
("query_google_scholar", query_google_scholar,
|
| 70 |
+
"_extract_google_scholar_results", google_scholar_url),
|
| 71 |
+
("query_google_patents", query_google_patents,
|
| 72 |
+
"_extract_google_patents_results", google_patents_search_url),
|
| 73 |
+
("query_brave_search", query_brave_search,
|
| 74 |
+
"_extract_brave_results", brave_search_url),
|
| 75 |
+
("query_bing_search", query_bing_search,
|
| 76 |
+
"_extract_bing_results", bing_search_url),
|
| 77 |
+
]
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
@pytest.mark.parametrize("name, scraper, extractor_attr, builder",
|
| 81 |
+
SCRAPERS, ids=[s[0] for s in SCRAPERS])
|
| 82 |
+
async def test_scraper_navigates_to_the_url_its_builder_produced(
|
| 83 |
+
monkeypatch, name, scraper, extractor_attr, builder):
|
| 84 |
+
monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
|
| 85 |
+
browser = _RecordingBrowser()
|
| 86 |
+
|
| 87 |
+
await scraper(browser, "agentic ai", 25)
|
| 88 |
+
|
| 89 |
+
assert browser.page.goto_urls == [builder("agentic ai", 25)]
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
@pytest.mark.parametrize("name, scraper, extractor_attr, builder",
|
| 93 |
+
SCRAPERS, ids=[s[0] for s in SCRAPERS])
|
| 94 |
+
async def test_scraper_returns_what_its_extractor_produced(
|
| 95 |
+
monkeypatch, name, scraper, extractor_attr, builder):
|
| 96 |
+
monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
|
| 97 |
+
|
| 98 |
+
results = await scraper(_RecordingBrowser(), "widgets", 10)
|
| 99 |
+
|
| 100 |
+
assert results == SENTINEL
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
@pytest.mark.parametrize("name, scraper, extractor_attr, builder",
|
| 104 |
+
SCRAPERS, ids=[s[0] for s in SCRAPERS])
|
| 105 |
+
async def test_scraper_blocks_decorative_resources(
|
| 106 |
+
monkeypatch, name, scraper, extractor_attr, builder):
|
| 107 |
+
"""Skipping stylesheets and images meaningfully speeds up navigation,
|
| 108 |
+
and every scraper is supposed to do it."""
|
| 109 |
+
monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
|
| 110 |
+
browser = _RecordingBrowser()
|
| 111 |
+
|
| 112 |
+
await scraper(browser, "widgets", 10)
|
| 113 |
+
|
| 114 |
+
assert browser.page.route_patterns == ["**/*"]
|
| 115 |
+
|
| 116 |
+
|
| 117 |
+
@pytest.mark.parametrize("name, scraper, extractor_attr, builder",
|
| 118 |
+
SCRAPERS, ids=[s[0] for s in SCRAPERS])
|
| 119 |
+
async def test_scraper_closes_its_page_and_context(
|
| 120 |
+
monkeypatch, name, scraper, extractor_attr, builder):
|
| 121 |
+
monkeypatch.setattr(serp, extractor_attr, _sentinel_extractor)
|
| 122 |
+
browser = _RecordingBrowser()
|
| 123 |
+
|
| 124 |
+
await scraper(browser, "widgets", 10)
|
| 125 |
+
|
| 126 |
+
assert browser.page.closed and browser.context.closed
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
async def test_page_and_context_are_closed_even_when_extraction_raises(monkeypatch):
|
| 130 |
+
async def exploding_extractor(page, n_results):
|
| 131 |
+
raise RuntimeError("selectors changed")
|
| 132 |
+
|
| 133 |
+
monkeypatch.setattr(serp, "_extract_bing_results", exploding_extractor)
|
| 134 |
+
browser = _RecordingBrowser()
|
| 135 |
+
|
| 136 |
+
with pytest.raises(RuntimeError):
|
| 137 |
+
await query_bing_search(browser, "widgets", 10)
|
| 138 |
+
|
| 139 |
+
assert browser.page.closed and browser.context.closed
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
async def test_context_is_closed_even_when_closing_the_page_raises():
|
| 143 |
+
"""A crashed renderer can make page.close() throw; the context still has
|
| 144 |
+
to be released or it leaks for the process's lifetime."""
|
| 145 |
+
browser = _RecordingBrowser()
|
| 146 |
+
|
| 147 |
+
async def exploding_close():
|
| 148 |
+
raise RuntimeError("renderer gone")
|
| 149 |
+
|
| 150 |
+
browser.page.close = exploding_close
|
| 151 |
+
|
| 152 |
+
with pytest.raises(RuntimeError):
|
| 153 |
+
async with playwright_open_page(browser):
|
| 154 |
+
pass
|
| 155 |
+
|
| 156 |
+
assert browser.context.closed
|