File size: 26,767 Bytes
b10b360
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bf13032
b10b360
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bf13032
b10b360
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
# SatQuery AI

**An interactive vision-language assistant for multimodal remote-sensing image analysis.**

Ask a natural-language question about a satellite or aerial image β€” or a pair of images β€” and SatQuery
routes it to the right specialist models, collects evidence, and returns a single confidence-scored
result envelope. It runs on CPU, is served from a static frontend, and is live at
**https://satquery.pages.dev**.

<p align="center">
  <img src="screenshots/analyze-grounding.png" alt="The SatQuery Analyze console answering a grounding question against live inference" width="820">
</p>

> **Status: research prototype, pre-1.0.** The architecture is frozen. This repository is the
> documented public release of the system, its trained artifacts, and its measured results β€”
> **including the negative ones.**

---

## Table of contents

- [Motivation](#motivation)
- [What the system supports](#what-the-system-supports)
- [Supported inputs](#supported-inputs)
- [Architecture](#architecture)
- [Routing and the execution trace](#routing-and-the-execution-trace)
- [Real inference vs. the preview path](#real-inference-vs-the-preview-path)
- [Models](#models)
- [Measured results](#measured-results)
- [Live validation](#live-validation)
- [Installation](#installation)
- [Local development](#local-development)
- [Deployment](#deployment)
- [Reproducibility](#reproducibility)
- [Known limitations](#known-limitations)
- [Links](#links)
- [Citation](#citation)

---

## Motivation

Remote-sensing analysis is fragmented. Detecting change between two acquisitions, localising an
object, captioning a scene, answering a question about it, and fusing optical with SAR each live in a
different model, a different preprocessing convention, and a different output schema. Assembling them
into one answer means re-solving the same problems β€” tiling, band handling, coordinate systems,
confidence β€” every time.

SatQuery AI explores a single hypothesis: **a small deterministic router plus a shared evidence
contract can make a heterogeneous specialist ensemble behave like one system**, without a large
language model in the control path. The router *understands* the query. A deterministic policy
*decides* which specialists run. The specialists *compute*. The evidence engine *proves* the answer.

Two design rules follow from that, and they are non-negotiable in the codebase:

- **No LLM-generated coordinates. No LLM-generated confidence.** Coordinates come from detection and
  segmentation heads; confidence comes from a calibrated scoring path.
- **Every result carries an observable execution trace** β€” never chain-of-thought.

## What the system supports

Six specialist tasks. All six are reported `available: true` by the live capability contract
(`GET /api/capabilities`, probed 2026-09-25).

| Task | What it answers | Assets |
|---|---|---|
| `vqa` | A free-form question about a single scene | 1 |
| `caption` | A description of a single scene | 1 |
| `grounding` | *Where* is a described object or region β€” returns boxes | 1 |
| `change` | *What changed* between two co-registered acquisitions β€” returns change regions | 2 |
| `change_vqa` | A yes/no or short question about a detected change | 2 |
| `optical_sar` | Joint scene classification from an optical + SAR pair | 2 |

## Supported inputs

Confirmed by the implementation, not assumed:

| Modality | Task(s) | Format |
|---|---|---|
| Optical, single image | `vqa`, `caption`, `grounding` | JPEG, PNG, TIFF |
| Temporal optical pair | `change`, `change_vqa` | Two images of **identical dimensions** |
| Optical + SAR pair | `optical_sar` | GeoTIFF/TIFF preferred |

**Modality is inferred server-side from band count**, not from the file extension: `{1, 2}` bands β‡’
SAR, `{3, 4, 8, 11, 12, 13}` bands β‡’ optical. The browser cannot read band count, so the console warns
when a submitted pair looks like two ordinary photographs rather than an optical/SAR pair.

**Per-file upload limit: 4,194,304 bytes (4 MiB).** Larger files are refused with HTTP 413 β€” imagery
must be downscaled first.

## Architecture

This is the **actually deployed** topology. An older direct-client-to-inference design is superseded.

```mermaid
flowchart TD
  B["Browser<br/>(static console)"] -->|HTTPS| CF["Cloudflare Pages<br/>satquery.pages.dev"]
  CF -->|"HTTPS JSON Β· /api/*"| R["Render<br/>satquery-orchestrator"]
  R -->|"outbound long-poll<br/>POST /tunnel/agent"| T{{"outbound tunnel"}}
  T --> C["GitHub Codespace<br/>FastAPI inference Β· CPU Β· :8000"]
  C --> S["Specialists"]
  S --> M["SmolVLM Β· RemoteCLIP Β· STANet-change<br/>CROMA-fusion Β· MiniLM router"]
  M --> E["Evidence engine<br/>+ temperature scaling"]
  E --> RE["ResultEnvelope"]
  RE -->|"tunnel β†’ Render"| B
```

**Why a tunnel.** The inference host runs in a GitHub Codespace. The forwarded-port path is not
reachable for a private repo (it returns HTTP 302), so the orchestrator keeps a **long-poll tunnel**:
the Codespace dials out to `POST /tunnel/agent` and holds the connection; Render queues work onto it.
`transport_mode` is `auto`, and the tunnel is the live transport.

### Repository map

| Path | Contents |
|---|---|
| `app/` | FastAPI inference service and its composition root |
| `core/` | Config registry, evidence engine, contracts |
| `specialists/` | One module per specialist (vqa, caption, grounding, change, optical_sar) |
| `router/` | MiniLM intent router |
| `gateway/` | Render orchestration hub (`/api/*`, CORS, wake flow) |
| `frontend/` | The static console (HTML/CSS/JS) |
| `configs/base.yaml` | The frozen configuration registry β€” single source of truth |
| `evaluation/`, `training/` | Evaluation harnesses and training entry points |
| `artifacts/` | Trained heads, checkpoints, evaluation outputs, provenance |
| `docs/` | Architecture, models, benchmarks, deployment, limitations |

## Routing and the execution trace

Routing is deliberately two-stage, and the split matters:

1. **`interpret()` β€” the reading.** A lexical pass over the query produces a *reading*: task intent,
   modality, temporal requirement, spatial scope, and expected evidence kind. It is
   **asset-count-blind**.
2. **`chooseTask()` β€” the dispatch.** The reading is combined with the number of attached assets to
   decide the task actually dispatched. This is why a reading of `change` with **one** asset dispatches
   `change_vqa` β€” the documented quantifier upgrade.

> **A router bug worth recording.** An earlier revision evaluated the temporal rule before the
> location rule, so *"Where are the built-up areas in this image?"* matched `\bbuilt\b` as a *change*
> marker and `area` inside *"areas"* as a quantifier. With one asset it collapsed to `vqa` and
> answered "River". Fixed on 2026-09-25 in `frontend/assets/js/mission.js`; the fix is covered by
> regression tests and verified live. The same defect existed on a second surface
> (`SQ.policy` in `core.js`) and was fixed the same day.

### The eight execution events

The console renders an execution trace built from **eight events**, emitted by the frontend around
real network calls (`SQ.EVENT_NAMES` in `frontend/assets/js/core.js`):

| # | Event | Emitted when |
|---|---|---|
| 1 | `QUERY_RECEIVED` | The query and assets are accepted |
| 2 | `QUERY_UNDERSTOOD` | `interpret()` has produced the reading |
| 3 | `ROUTE_SELECTED` | `chooseTask()` has selected the dispatched task |
| 4 | `SPECIALIST_STARTED` | The inference request has been issued |
| 5 | `SPECIALIST_COMPLETED` | The specialist has returned |
| 6 | `EVIDENCE_GENERATED` | Evidence items are available |
| 7 | `CONFIDENCE_COMPUTED` | The calibrated confidence is available |
| 8 | `RESULT_ASSEMBLED` | The `ResultEnvelope` is complete |

These are a **frontend** vocabulary driven by observable events β€” not a backend protocol and not a
model's reasoning trace. On live runs the trace bar reaches **94.4444 %** (17/18) and every node is
marked live; the preview path is the only source of mock-marked nodes.

## Real inference vs. the preview path

- **Real path (production).** With assets attached, the console calls
  `POST /api/infer` on the Render orchestrator. Every run returns a real `run_*` identifier from the
  inference service. **Live validation recorded 0 mock nodes across 24 live runs.**
- **Preview path.** With *no* files selected, the console renders a labelled illustrative preview so
  the interface is explorable without the stack awake. Preview nodes are explicitly marked `is-mock`
  and never appear in a live run.

The distinction is observable, not asserted: a live run shows `live Β· N evidence Β· transport …` and
zero `.trace__node.is-mock` elements.

## Models

Six trained artifacts are released. **Four are task heads and two are adapters** β€” none is a complete
standalone model, and each documents its backbone dependency. Full detail: [`docs/MODELS.md`](docs/MODELS.md),
[`MODEL_CARD.md`](MODEL_CARD.md), and the generated [`models/manifest.json`](models/manifest.json).

| Task | Backbone (pinned) | Custom component | Artifact | Size | Eval data | Metric | Status |
|---|---|---|---|---|---|---|---|
| `change` | STANet-style, ResNet-18 encoder, PAM | change head | `head.pt` | 63,231,009 B | LEVIR-CD-256, test n=2048 | pooled IoU **0.8122** Β· macro IoU **0.8457** Β· pooled F1 **0.8964** | **VERIFIED** |
| `grounding` | `chendelong/RemoteCLIP` ViT-B/32 @ `bf1d8a3ccf2d` (frozen) | trainable head (2048β†’512) | `head.pt` | 12,639,041 B | VRSBench, n=16159 | mean_best_IoU **0.2838** Β· recall@0.5 **0.2198** (canonical) | measured β€” two protocols |
| `optical_sar` | `antofuller/CROMA` base @ `0dd28e3d633b` | fusion head (2318β†’512β†’19) | `head.pt` | 14,427,457 B | BigEarthNet, 19 CLC classes, test n=4000 | accuracy **0.931** Β· macro_F1 **0.434161** | measured β€” **ruling OPEN** |
| `change_vqa` | as `change` | change-VQA head | `head.pt` | 5,822,809 B | test n=39686 | accuracy **0.697626** Β· macro_F1 **0.378373** | measured β€” **ruling OPEN** |
| `router` | `sentence-transformers/all-MiniLM-L6-v2` @ `1110a243fdf4` | intent adapter | `adapter.pt` | 211,961 B | val n=86 | accuracy **0.965116** | **TEST NOT RUN** |
| `vqa` / `caption` | `HuggingFaceTB/SmolVLM-500M-Instruct` @ `a7da5b986cb5` | **LoRA** (r=16, Ξ±=32, dropout 0.05) | `adapter_model.safetensors` | 34,798,048 B | frozen 1000-Q subset | exact_match **0.963** Β· F1 **0.96432** | **ACCEPTANCE-REJECTED** |

Backbones are third-party and pinned by `repo_id` + `revision` in `configs/base.yaml`; they are
fetched from the Hugging Face Hub, not redistributed here.

## Measured results

Every number below traces to an artifact, a test, or a live run. **Nothing here is a system-level
benchmark β€” no such benchmark exists** (see [Known limitations](#known-limitations)).

| Metric | Value | Split / protocol | Source key | Status |
|---|---|---|---|---|
| Change pooled IoU | 0.8122 | LEVIR-CD-256 test, n=2048, thr 0.50 | `metrics.pooled.iou` | **VERIFIED** |
| Change macro IoU | 0.8457 | same | `metrics.macro.miou` | **VERIFIED** |
| Change pooled F1 | 0.8964 | same | `metrics.pooled.f1` | **VERIFIED** |
| Grounding mean_best_IoU (canonical, `head_threshold`) | 0.2838 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
| Grounding recall@0.5 (canonical, `head_threshold`) | 0.2198 | same | `results.head_threshold.recall.0.50` | measured |
| Grounding mean_best_IoU (matched6, `head_threshold`) | 0.2566 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
| Grounding recall@0.5 (matched6, `head_threshold`) | 0.1938 | same | `results.head_threshold.recall.0.50` | measured |
| Grounding `head_argmax` decode (both protocols) | 0.1215 | same | `results.head_argmax.mean_best_iou` | measured β€” **worse** |
| Grounding zero-shot baseline | 0.0972 | same | `results.zero_shot_matched.mean_best_iou` | measured |
| Optical-SAR accuracy | 0.931 | BigEarthNet, held-out test n=4000 | `accuracy` | measured β€” ruling **OPEN** |
| Optical-SAR macro_F1 | 0.434161 | same | `macro_f1` | measured β€” ruling **OPEN** |
| Change-VQA accuracy (**test**) | 0.697626 | test n=39686 | `verification.test_accuracy` | measured β€” ruling **OPEN** |
| Change-VQA macro_F1 (**test**) | 0.378373 | same | `verification.test_macro_f1` | measured β€” ruling **OPEN** |
| Change-VQA accuracy (**test2**) | 0.651469 | second test set | `verification.test2_accuracy` | measured β€” **lower** |
| Change-VQA macro_F1 (**test2**) | 0.372309 | second test set | `verification.test2_macro_f1` | measured β€” **lower** |
| VLM adapter exact_match | 0.963 | frozen 1000-Q subset | `artifacts/vlm/phase6_closure.json` | USABLE_VERIFIED β€” **ACCEPTANCE-REJECTED** |
| VLM adapter F1 | 0.96432 | same | same | USABLE_VERIFIED β€” **ACCEPTANCE-REJECTED** |
| Router **overall ungated** accuracy | 0.965116 | val, n=86, corpus-limited | `overall_ungated_accuracy` | **TEST NOT RUN** |
| System-level end-to-end benchmark | β€” | β€” | β€” | **NOT RUN β€” none exists** |

### Grounding: three decode variants, two protocols

The grounding head is evaluated under **two matching protocols** (canonical, matched6) and **three
decode variants**. Quoting a single number would misrepresent the result, so all of them are listed:

| Decode | canonical mean_best_IoU | matched6 mean_best_IoU |
|---|---|---|
| `head_threshold` (the headline number) | **0.2838** | **0.2566** |
| `head_argmax` | 0.1215 | 0.1215 |
| `zero_shot_matched` (baseline, no head) | 0.0972 | 0.0972 |

The head clears the zero-shot baseline, but only the threshold decode is meaningfully above it β€” the
argmax decode (0.1215) is barely better than zero-shot. The absolute level is modest either way:
**grounding is useful, not solved.**

### Fusion: measured but the ruling is open

Optical-SAR fusion reaches **0.931 accuracy** on a 19-class held-out set of 4,000 β€” but only
**0.434 macro_F1**. Those two numbers describe very different things: the model is accurate on
frequent classes and weak on rare ones. The acceptance ruling for this head is **OPEN**, and the
headline accuracy must never be quoted without the macro_F1 beside it.

### Change-VQA: two test sets, and they disagree

`artifacts/change_vqa/run/PROMOTION.json` records **two** test evaluations:

| Split | accuracy | macro_F1 |
|---|---|---|
| `test` | 0.697626 | 0.378373 |
| `test2` | **0.651469** | 0.372309 |

The `test` numbers are the higher pair. Both are reported here; quoting only `test` would overstate
the result. The acceptance ruling is **OPEN**.

### Phase 6 / VLM: deployment success β‰  model acceptance

The Phase-6 SmolVLM LoRA adapter reaches **exact_match 0.963** and **F1 0.96432** on a frozen
1,000-question subset. It is marked **USABLE_VERIFIED** and **ACCEPTANCE-REJECTED**.

Those two verdicts are not in conflict, and the distinction is the point:

- **USABLE_VERIFIED** β€” the adapter loads, runs, and produces the measured numbers in the deployed
  pipeline.
- **ACCEPTANCE-REJECTED** β€” the change did not clear the project's own pre-registered acceptance bar.

A model can be a working engineering artifact and a rejected research result at the same time.
This release keeps both labels.

### Calibration: it got worse, and we say so

The `change_vqa` confidence path applies temperature scaling (`T = 0.9773`). Measured on the
validation split (n=16441):

| | ECE |
|---|---|
| Before temperature scaling | **0.013755** |
| After temperature scaling | **0.014929** |

**Calibration did not improve β€” it moved slightly worse.** The scaling is retained because it is part
of the frozen configuration, not because it helped. The reliability curve plotted on the Benchmark page
is explicitly labelled as the **pre-scaling** diagram so a reader cannot mistake it for the calibrated
result.

## Live validation

Validation drove the **production site** in a headed browser, one upload per case, with per-case
screenshots and recorded run identifiers.

| Property | Result |
|---|---|
| Independent full passes | **3** |
| Cases per pass | 8 (6 regression + 2 defect) |
| Passes at 8/8 | **3 of 3** |
| Live runs executed | **24** |
| Correct dispatches | **24** |
| Mock-node contamination | **0** on every live run |
| Trace fill | 94.4444 % on every live run |
| Frontend regression suite | **106 passed** (`tests/unit/test_frontend_live_wiring.py`) |

Each pass produced **fresh run identifiers** β€” no run id is shared between passes.

The harness asserts the form state **before** dispatch β€” that the query box really holds the intended
query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. This matters:
an earlier harness revision typed with synthetic key events that Chrome silently drops when the window
lacks OS focus, so it dispatched the page's *default* query and still recorded a "result". The
assertions exist because of that failure.

### Representative real run IDs

One full pass (pass 3 of 3) β€” the same pass the screenshots below are drawn from. Run identifiers
are fresh on every pass; the other two passes recorded different ids.

| Case | Query | Dispatched | Run ID |
|---|---|---|---|
| vqa | What type of terrain dominates this scene? | `vqa` | `run_fef26e91e7e6` |
| caption | Describe the main visual characteristics of this scene. | `caption` | `run_96281bdfcc08` |
| grounding | Where are the visible buildings in this image? | `grounding` | `run_e49adc8d319f` |
| change | What changed between the earlier and later image? | `change` | `run_aedc59cbcdc9` |
| change_vqa | Did the coastline advance between the two observations? | `change_vqa` | `run_62ca98d510be` |
| optical_sar | …combining the optical and SAR observations? | `optical_sar` | `run_beacf6aa4e21` |
| **grounding** | **Where are the built-up areas in this image?** | **`grounding`** | **`run_467ffa406f22`** |
| **grounding** | **Where is the new airport?** | **`grounding`** | **`run_46980ba55c62`** |

The last two are the router-defect queries. Both previously collapsed to `vqa` and answered "River".

### Screenshots

Eight captures from the post-fix live run (headed browser, 1384Γ—855, one upload per case). Each
panel shows the run identifier, the frozen config hash `78f1e3700da15aa1`, and the evidence list
returned by the specialist β€” nothing is mocked.

| | |
|---|---|
| ![Grounding β€” built-up areas](screenshots/analyze-grounding.png) | ![Optical-SAR](screenshots/analyze-optical-sar.png) |
| **Grounding** β€” "Where are the built-up areas in this image?" β€” the fixed router defect (`run_467ffa406f22`, dispatched `grounding`, not `vqa`) | **Optical-SAR** fusion on a real optical/SAR GeoTIFF pair (`run_beacf6aa4e21`, fused class 18) |
| ![Grounding β€” new airport](screenshots/analyze-grounding-new-airport.png) | ![Grounding β€” visible buildings](screenshots/analyze-grounding-buildings.png) |
| **Grounding** β€” "Where is the new airport?" β€” second defect query (`run_46980ba55c62`, dispatched `grounding`) | **Grounding** β€” "Where are the visible buildings in this image?" (`run_e49adc8d319f`) |
| ![Change](screenshots/analyze-change.png) | ![Caption](screenshots/analyze-caption.png) |
| **Change** detection on a same-shape temporal pair | **Caption** of a single scene (`run_96281bdfcc08`, calibrated confidence 1.000) |
| ![VQA](screenshots/analyze-vqa.png) | ![Change-VQA](screenshots/analyze-change-vqa.png) |
| **VQA** β€” "What type of terrain dominates this scene?" | **Change-VQA** β€” "Did the coastline advance between the two observations?" |

All eight are reproduced byte-for-byte in the evidence archive (Phase 7) with SHA-256 recorded in
`RELEASE_MANIFEST.md`.

## Installation

Python 3.11+ and a CPU are sufficient. No CUDA requirement β€” device is selected via
`SATQUERY_DEVICE`; all placement is `.to(device)`, never `.cuda()`.

```bash
git clone https://github.com/Anish-lab-blip/SatQuery-AI
cd SatQuery-AI
python -m venv .venv
source .venv/Scripts/activate      # Windows git-bash; use .venv/bin/activate on Linux/macOS
pip install -r requirements.txt
```

Backbones are fetched from the Hugging Face Hub on first use, pinned by revision in
`configs/base.yaml`. The frozen config hash is **`78f1e3700da15aa1`** β€” the loader refuses to run a
config that violates the recorded invariants (for example `fusion.input_dim == 3*encoder_dim + 12 + 2`).

## Local development

```bash
# Inference service, CPU (this is the launcher the Codespace runs)
PORT=8000 python deploy/codespace/serve.py

# Health
curl localhost:8000/v1/health
```

The frontend is fully static and needs no build step to serve locally:

```bash
python -m http.server 5500 --directory frontend
```

Run the frontend regression suite:

```bash
python -m pytest tests/unit/test_frontend_live_wiring.py -q
```

## Deployment

The live topology is **Cloudflare Pages β†’ Render β†’ outbound tunnel β†’ GitHub Codespace**.

| Layer | Role | Source |
|---|---|---|
| Cloudflare Pages | Static frontend at **https://satquery.pages.dev** | `frontend/` |
| Render | Orchestrator / API gateway, `/api/*`, CORS, wake flow | `gateway/` |
| GitHub Codespace | FastAPI inference host, CPU, port 8000 | `app/` |
| Hugging Face | Model cards, released artifacts, checksums | this release |

Deployment sources are **separate repositories** from this release. The wake flow is: Cloudflare β†’
Render β†’ start the Codespace if stopped β†’ poll `/v1/health` β†’ surface *"Waking inference engine…"* β†’
`POST /infer` β†’ result.

### Deployment caveats

- **Cold start.** The inference host may be stopped when idle. The first request after a cold start
  can exceed the client timeout while weights are fetched; a retry a few seconds later normally
  succeeds. Warm the stack before any demonstration and confirm
  `GET /api/health` reports `tunnel.agent_connected: true`.
- **Tunnel gaps.** The tunnel agent can be briefly absent. A request issued during such a gap may hang
  or return HTTP 504. **This is not fixed in production** β€” a prepared patch
  (`forward_unavailable` 503 / `upstream_timeout` 504 plus a `codespace_name` fix) exists and is
  documented, but it was deliberately not deployed. Root cause: in `auto` transport mode a tunnel
  timeout falls through to the forwarded-port path, which then spends the 120 s wake timeout on an
  HTTP 302 β€” the observed ~249 s failure.
- **`codespace_name`** is still reported with a trailing newline by `/api/health` (cosmetic; the wake
  path strips it).

## Reproducibility

1. **Configuration.** `configs/base.yaml` is the single registry; no magic numbers in Python. Its hash
   is recorded in every execution trace. Frozen hash: `78f1e3700da15aa1`.
2. **Backbones.** Pinned by `repo_id` + `revision`, never by floating tag:
   `HuggingFaceTB/SmolVLM-500M-Instruct@a7da5b986cb5`,
   `chendelong/RemoteCLIP@bf1d8a3ccf2d`,
   `sentence-transformers/all-MiniLM-L6-v2@1110a243fdf4`,
   `antofuller/CROMA@0dd28e3d633b`.
3. **Released artifacts.** `models/manifest.json` and `models/checksums.sha256` are **generated from
   the actual files** β€” never hand-typed. Verify with `sha256sum -c models/checksums.sha256`.
4. **Splits.** LEVIR-CD-256: train 7120 / val 1024 / test 2048. Grounding: VRSBench n=16159.
   Fusion: held-out test n=4000. Change-VQA: test n=39686. Leakage isolation is by `scene_id`.
5. **Prompts** are versioned files, frozen before benchmark evaluation.
6. **Negative results are preserved.** Rejected and open rulings are recorded, not removed.

## Known limitations

1. **No system-level end-to-end benchmark exists.** Per-specialist metrics are real; a single
   end-to-end number is **NOT RUN**.
2. **Router accuracy is validation-only** (n=86, corpus-limited). Its **test set was never run**.
3. **Optical-SAR fusion returns a bare class index** (`class_18`), not a human-readable label. The
   modality accounting in the response confirms the right channels reached the fusion head, but the
   presentation is not user-facing.
4. **VQA is weak-but-related.** Asked what terrain dominates a scene, it answers "Grassland".
5. **Fusion macro_F1 is low (0.434)** against 0.931 accuracy β€” rare classes are poorly handled.
6. **Grounding IoU is modest** (0.2838 canonical) β€” useful, not solved.
7. **Calibration makes ECE slightly worse**, and is retained only because it is part of the frozen
   configuration.
8. **Router lexical residuals.** *"What is the new runway?"* reads `change` rather than `vqa` (the
   `new`-as-change heuristic fires outside `where` questions), and *"How much built-up area was
   added?"* reads `vqa` (under-trigger). A lexical router cannot cleanly separate "the new X" from
   "what's new"; a trained intent router exists in `artifacts/router/` but is not attached.
9. **B-07 tunnel gaps are not fixed in production** (see Deployment caveats).
10. **No license has been selected** for this repository. Until one is, the artifacts carry
    `license: unknown` and no reuse rights should be assumed. This is an open owner decision.
11. **The Anatomy of a Run page** renders a recorded run whose plate uses the 720Γ—720 variant of an
    image analysed at 730Γ—730 β€” identical content, scaled by the canvas, but the "actual analysed
    image" wording is slightly loose.

## Links

| | |
|---|---|
| Live demo | https://satquery.pages.dev |
| GitHub | https://github.com/Anish-lab-blip/SatQuery-AI |
| Hugging Face | https://huggingface.co/thundercode/SatQuery |

Third-party models this work builds on (pinned, not redistributed):

| Model | Revision | Role |
|---|---|---|
| [`HuggingFaceTB/SmolVLM-500M-Instruct`](https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct) | `a7da5b986cb5` | VQA + captioning backbone |
| [`chendelong/RemoteCLIP`](https://huggingface.co/chendelong/RemoteCLIP) | `bf1d8a3ccf2d` | remote-sensing grounding encoder |
| [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | `1110a243fdf4` | router embedding |
| [`antofuller/CROMA`](https://huggingface.co/antofuller/CROMA) | `0dd28e3d633b` | optical/SAR fusion encoder |

Datasets referenced by the evaluations: LEVIR-CD-256 (change), VRSBench (grounding),
BigEarthNet (optical-SAR fusion, 19 CLC classes). No dataset is redistributed here.

## Citation

No paper accompanies this release. Until one exists, cite the repository:

```bibtex
@misc{satqueryai2026,
  title  = {SatQuery AI: an interactive vision-language assistant for
            multimodal remote-sensing image analysis},
  author = {SatQuery AI contributors},
  year   = {2026},
  url    = {https://github.com/Anish-lab-blip/SatQuery-AI}
}
```

## License

**Not yet selected.** See limitation 10. Backbone models remain under their own upstream licenses.