File size: 10,220 Bytes
724b09e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
# Architecture

**Status vocabulary used throughout these docs:** `IMPLEMENTED` (code exists and runs) ·
`VERIFIED` (checked against evidence) · `MEASURED` (a number was produced) · `ATTEMPTED`
(tried, outcome recorded) · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` (tried and
explicitly not accepted). Every architectural claim below carries one of these tags.

---

## 1. What the system is

SatQuery AI answers natural-language questions about satellite imagery. It is a **router +
specialists** system: a small intent router reads the question, dispatches it to one of six
task specialists, and assembles the specialist's output into a single structured
`ResultEnvelope` that the browser renders as an answer, an evidence list, and a confidence.

It is not a single end-to-end vision-language model. The only learned components are:

- a **50,822-parameter router adapter** over a frozen `all-MiniLM-L6-v2` sentence encoder;
- four trained **task heads/adapters** (change, change_vqa, optical-SAR fusion, grounding);
- one **LoRA adapter** over a frozen `SmolVLM-500M-Instruct` for caption/VQA.

Everything else is frozen, publicly-pinned backbone weights resolved at run time.

## 2. Deployment topology (IMPLEMENTED, VERIFIED)

The shipped system is a four-tier chain. **Status: VERIFIED live on 2026-09-25** — the health,
capabilities and infer endpoints were probed against the running deployment.

```
Browser
  │ HTTPS
  ▼
Cloudflare Pages  —  satquery.pages.dev          (static frontend, 11 pages)
  │ HTTPS/JSON
  ▼
Render            —  satquery-backend-m4yv.onrender.com   (orchestrator / gateway)
  │ outbound long-poll tunnel  POST /tunnel/agent
  ▼
GitHub Codespace  —  FastAPI inference, port 8000 (CPU)
  │ build_space_app()
  ▼
Specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
```

```mermaid
flowchart LR
  U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"]
  CF -->|"HTTPS JSON<br/>/api/health · /api/capabilities · /api/infer"| R["Render<br/>orchestrator / gateway"]
  R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"]
  C --> S[(SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet)]
  C -->|ResultEnvelope| R
  R -->|"envelope + error translation"| CF
```

### 2.1 Why a gateway and a tunnel

- **The gateway (Render)** is deliberately thin and stateless: schema validation, size limits,
  per-IP rate limiting, a CORS allowlist (**never** `*`), request IDs, timeouts, secret custody,
  and translation of upstream failures into the documented error envelope. It is **not** a model
  host and has no database, auth, or queue.
- **The tunnel** exists because the inference host (a Codespace) is not publicly addressable for
  this private repository — a forwarded port returns `302`. Instead the Codespace runs
  `tunnel_agent.py`, which **dials out** to `POST /tunnel/agent` and long-polls. Work is executed
  against `http://127.0.0.1:8000` locally. This inverts the usual direction: the inference host
  needs no inbound firewall hole.

> **Status:** IMPLEMENTED and VERIFIED. `GET /api/health` reported `tunnel.agent_connected:true`
> with a non-zero `completed` counter on 2026-09-25.

### 2.2 Cold start and the `transport_mode: auto` fallthrough (KNOWN ISSUE, OPEN)

The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore
tens of seconds and is **documented, not hidden**. A timeout on the tunnel path is surfaced to the
client as a recoverable error.

`SATQUERY_TRANSPORT=auto` means: try the tunnel; if it does not answer within
`SATQUERY_TUNNEL_TIMEOUT_S` (150 s), **fall through to the forward path**. The forward path to a
private repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S`
(120 s) first — so a worst-case failed request can take ≈ 249 s. This is the root of the observed
"transient tunnel gap" and is recorded as **OPEN** (not fixed) in
[`LIMITATIONS.md`](LIMITATIONS.md). A deployed fix (`codespace_name` newline handling on the
wake path) was authored and pushed separately; the fallthrough itself is by design.

## 3. Request lifecycle (IMPLEMENTED, VERIFIED)

The controller is a nine-state machine, declared in `configs/base.yaml`:

```
RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND
```

1. **RECEIVE / PARSE** — the query and 1–2 image assets are decoded.
2. **VALIDATE** — modality inference from band count, dimension equality for pairs, byte cap
   (`4,194,304` bytes/file), GeoTIFF acceptance for the optical-SAR path.
3. **PLAN** — the router embeds the query with the frozen MiniLM encoder and the trained adapter
   predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change
   task.
4. **PREPROCESS** — percentile normalisation for optical (2/98), dB clip for SAR (−30…+5 dB),
   tiling for large scenes.
5. **EXECUTE** — the chosen specialist runs.
6. **AGGREGATE** — the specialist's raw output becomes `Evidence` items
   (`Region` / `ChangeRegion`).
7. **VERIFY** — confidence is computed; temperature scaling is applied from
   `calibration_v001.json`.
8. **RESPOND** — a `ResultEnvelope` is assembled and returned.

### 3.1 The eight execution events

The frontend renders a live trace from eight named events (source of truth:
`frontend/assets/js/core.js`, `SQ.EVENT_NAMES`):

| # | Event | Emitted when |
|---|---|---|
| 1 | `QUERY_RECEIVED` | the request is accepted |
| 2 | `QUERY_UNDERSTOOD` | the router produces task + modality |
| 3 | `ROUTE_SELECTED` | the specialist is chosen |
| 4 | `SPECIALIST_STARTED` | specialist execution begins |
| 5 | `SPECIALIST_COMPLETED` | specialist returns |
| 6 | `EVIDENCE_GENERATED` | evidence items are built |
| 7 | `CONFIDENCE_COMPUTED` | confidence (and calibration) is computed |
| 8 | `RESULT_ASSEMBLED` | the envelope is finalised |

**Live runs emit all eight; the offline preview path emits none of the specialist events and is
labelled as a preview in the UI.** Measured trace fill on every live run: **94.4444 %**.

## 4. Two reading paths: `interpret()` vs `chooseTask()` (IMPLEMENTED, VERIFIED)

There are two separate functions, and they see different information:

- **`interpret()`** — produces the *human-readable* interpretation shown in the console. It is
  **asset-count-blind**: it only sees the query text.
- **`chooseTask()`** — performs the *dispatch*. It is **asset-count-aware**: it knows how many
  assets are attached and will not route a single-image query to a two-image task.

This asymmetry is real and explains a documented behaviour: for the query *"What changed between
the earlier and later image?"* with a single asset attached, the console can *read* `change` while
dispatch correctly falls back to `change_vqa`. This is intentional, not a bug, and is covered in
[`RESEARCH_NOTES.md`](RESEARCH_NOTES.md).

## 5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED)

The Codespace exposes exactly four routes; the gateway mirrors each under `/api/*`:

| Codespace | Gateway | Cost | Purpose |
|---|---|---|---|
| `GET /v1/health` | `GET /api/health` | cheap | liveness, tunnel state, counters |
| `GET /v1/capabilities` | `GET /api/capabilities` | cheap | per-task availability |
| `POST /v1/analyze` | `POST /api/infer` | **costly** | run a query |
| `POST /v1/assets` | `POST /api/assets` | **costly** | stage an asset handle |

The gateway must **not** retry `POST /api/infer` on its own — a retry would consume inference a
second time. The client decides on retry. Errors are translated into a stable envelope with a
machine code (`invalid_request`, `tunnel_offline`, …) and a `recoverable` flag.

**Verified:** `POST /api/infer {}` returns `422 invalid_request` with header
`x-satquery-transport: tunnel`.

## 6. The frozen configuration (VERIFIED)

All tunables live in `configs/base.yaml`; **no magic numbers in Python**. The loader validates
invariants and computes a hash. The frozen hash is:

```
78f1e3700da15aa1
```

The loader **refuses to run** a config that violates a recorded invariant — for example
`fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels`
(2318) and `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (2048). A mismatch
in the latter is a *silent* shape error otherwise (torch raises only later, after features are
cached), so it is enforced at load time.

Editing `configs/base.yaml` moves the hash, which invalidates every artifact keyed to it. This is
why the stale `configs/deploy.yaml` (which still describes an HF-Space/ZeroGPU target) is left
**undisturbed as frozen paperwork** — it is never read at run time.

## 7. Component map

| Layer | Path (in the source repo) | Notes |
|---|---|---|
| Frontend | `frontend/` | static, staged by `scripts/stage_pages.mjs` |
| Gateway | `deploy/render/` | `main.py` exposes `app`; `render.yaml` blueprint |
| Inference entry | `deploy/codespace/serve.py` | serves `build_space_app()` on `$PORT` |
| Inference app | `app/space_app.py` | `build_space_app()`; 4-endpoint contract |
| Composition root | `app/serving.py` | reuses the local serving path |
| Specialists | `specialists/` | vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA) |
| Router | `router/` | frozen MiniLM + trained adapter |
| Config | `configs/base.yaml` | authoritative registry; hashed |
| Artifacts | `artifacts/` | weights, metrics, calibration, reports |

> **Caveat.** `deploy/` inside the monorepo is **stale and untracked**; the deployed sources live in
> three separate private repositories. See [`DEPLOYMENT.md`](DEPLOYMENT.md).

## 8. What is deliberately absent

- **No database, no auth, no queue.** The gateway is stateless by design.
- **No GPU requirement.** Device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`,
  never `.cuda()`. The shipped deployment runs CPU-only.
- **No system-level end-to-end benchmark.** There is no measured end-to-end accuracy number, and
  none is claimed. Per-specialist metrics exist and are reported in [`BENCHMARKS.md`](BENCHMARKS.md).