Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 10,220 Bytes
724b09e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 | # Architecture
**Status vocabulary used throughout these docs:** `IMPLEMENTED` (code exists and runs) ·
`VERIFIED` (checked against evidence) · `MEASURED` (a number was produced) · `ATTEMPTED`
(tried, outcome recorded) · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` (tried and
explicitly not accepted). Every architectural claim below carries one of these tags.
---
## 1. What the system is
SatQuery AI answers natural-language questions about satellite imagery. It is a **router +
specialists** system: a small intent router reads the question, dispatches it to one of six
task specialists, and assembles the specialist's output into a single structured
`ResultEnvelope` that the browser renders as an answer, an evidence list, and a confidence.
It is not a single end-to-end vision-language model. The only learned components are:
- a **50,822-parameter router adapter** over a frozen `all-MiniLM-L6-v2` sentence encoder;
- four trained **task heads/adapters** (change, change_vqa, optical-SAR fusion, grounding);
- one **LoRA adapter** over a frozen `SmolVLM-500M-Instruct` for caption/VQA.
Everything else is frozen, publicly-pinned backbone weights resolved at run time.
## 2. Deployment topology (IMPLEMENTED, VERIFIED)
The shipped system is a four-tier chain. **Status: VERIFIED live on 2026-09-25** — the health,
capabilities and infer endpoints were probed against the running deployment.
```
Browser
│ HTTPS
▼
Cloudflare Pages — satquery.pages.dev (static frontend, 11 pages)
│ HTTPS/JSON
▼
Render — satquery-backend-m4yv.onrender.com (orchestrator / gateway)
│ outbound long-poll tunnel POST /tunnel/agent
▼
GitHub Codespace — FastAPI inference, port 8000 (CPU)
│ build_space_app()
▼
Specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
```
```mermaid
flowchart LR
U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"]
CF -->|"HTTPS JSON<br/>/api/health · /api/capabilities · /api/infer"| R["Render<br/>orchestrator / gateway"]
R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"]
C --> S[(SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet)]
C -->|ResultEnvelope| R
R -->|"envelope + error translation"| CF
```
### 2.1 Why a gateway and a tunnel
- **The gateway (Render)** is deliberately thin and stateless: schema validation, size limits,
per-IP rate limiting, a CORS allowlist (**never** `*`), request IDs, timeouts, secret custody,
and translation of upstream failures into the documented error envelope. It is **not** a model
host and has no database, auth, or queue.
- **The tunnel** exists because the inference host (a Codespace) is not publicly addressable for
this private repository — a forwarded port returns `302`. Instead the Codespace runs
`tunnel_agent.py`, which **dials out** to `POST /tunnel/agent` and long-polls. Work is executed
against `http://127.0.0.1:8000` locally. This inverts the usual direction: the inference host
needs no inbound firewall hole.
> **Status:** IMPLEMENTED and VERIFIED. `GET /api/health` reported `tunnel.agent_connected:true`
> with a non-zero `completed` counter on 2026-09-25.
### 2.2 Cold start and the `transport_mode: auto` fallthrough (KNOWN ISSUE, OPEN)
The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore
tens of seconds and is **documented, not hidden**. A timeout on the tunnel path is surfaced to the
client as a recoverable error.
`SATQUERY_TRANSPORT=auto` means: try the tunnel; if it does not answer within
`SATQUERY_TUNNEL_TIMEOUT_S` (150 s), **fall through to the forward path**. The forward path to a
private repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S`
(120 s) first — so a worst-case failed request can take ≈ 249 s. This is the root of the observed
"transient tunnel gap" and is recorded as **OPEN** (not fixed) in
[`LIMITATIONS.md`](LIMITATIONS.md). A deployed fix (`codespace_name` newline handling on the
wake path) was authored and pushed separately; the fallthrough itself is by design.
## 3. Request lifecycle (IMPLEMENTED, VERIFIED)
The controller is a nine-state machine, declared in `configs/base.yaml`:
```
RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND
```
1. **RECEIVE / PARSE** — the query and 1–2 image assets are decoded.
2. **VALIDATE** — modality inference from band count, dimension equality for pairs, byte cap
(`4,194,304` bytes/file), GeoTIFF acceptance for the optical-SAR path.
3. **PLAN** — the router embeds the query with the frozen MiniLM encoder and the trained adapter
predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change
task.
4. **PREPROCESS** — percentile normalisation for optical (2/98), dB clip for SAR (−30…+5 dB),
tiling for large scenes.
5. **EXECUTE** — the chosen specialist runs.
6. **AGGREGATE** — the specialist's raw output becomes `Evidence` items
(`Region` / `ChangeRegion`).
7. **VERIFY** — confidence is computed; temperature scaling is applied from
`calibration_v001.json`.
8. **RESPOND** — a `ResultEnvelope` is assembled and returned.
### 3.1 The eight execution events
The frontend renders a live trace from eight named events (source of truth:
`frontend/assets/js/core.js`, `SQ.EVENT_NAMES`):
| # | Event | Emitted when |
|---|---|---|
| 1 | `QUERY_RECEIVED` | the request is accepted |
| 2 | `QUERY_UNDERSTOOD` | the router produces task + modality |
| 3 | `ROUTE_SELECTED` | the specialist is chosen |
| 4 | `SPECIALIST_STARTED` | specialist execution begins |
| 5 | `SPECIALIST_COMPLETED` | specialist returns |
| 6 | `EVIDENCE_GENERATED` | evidence items are built |
| 7 | `CONFIDENCE_COMPUTED` | confidence (and calibration) is computed |
| 8 | `RESULT_ASSEMBLED` | the envelope is finalised |
**Live runs emit all eight; the offline preview path emits none of the specialist events and is
labelled as a preview in the UI.** Measured trace fill on every live run: **94.4444 %**.
## 4. Two reading paths: `interpret()` vs `chooseTask()` (IMPLEMENTED, VERIFIED)
There are two separate functions, and they see different information:
- **`interpret()`** — produces the *human-readable* interpretation shown in the console. It is
**asset-count-blind**: it only sees the query text.
- **`chooseTask()`** — performs the *dispatch*. It is **asset-count-aware**: it knows how many
assets are attached and will not route a single-image query to a two-image task.
This asymmetry is real and explains a documented behaviour: for the query *"What changed between
the earlier and later image?"* with a single asset attached, the console can *read* `change` while
dispatch correctly falls back to `change_vqa`. This is intentional, not a bug, and is covered in
[`RESEARCH_NOTES.md`](RESEARCH_NOTES.md).
## 5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED)
The Codespace exposes exactly four routes; the gateway mirrors each under `/api/*`:
| Codespace | Gateway | Cost | Purpose |
|---|---|---|---|
| `GET /v1/health` | `GET /api/health` | cheap | liveness, tunnel state, counters |
| `GET /v1/capabilities` | `GET /api/capabilities` | cheap | per-task availability |
| `POST /v1/analyze` | `POST /api/infer` | **costly** | run a query |
| `POST /v1/assets` | `POST /api/assets` | **costly** | stage an asset handle |
The gateway must **not** retry `POST /api/infer` on its own — a retry would consume inference a
second time. The client decides on retry. Errors are translated into a stable envelope with a
machine code (`invalid_request`, `tunnel_offline`, …) and a `recoverable` flag.
**Verified:** `POST /api/infer {}` returns `422 invalid_request` with header
`x-satquery-transport: tunnel`.
## 6. The frozen configuration (VERIFIED)
All tunables live in `configs/base.yaml`; **no magic numbers in Python**. The loader validates
invariants and computes a hash. The frozen hash is:
```
78f1e3700da15aa1
```
The loader **refuses to run** a config that violates a recorded invariant — for example
`fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels`
(2318) and `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (2048). A mismatch
in the latter is a *silent* shape error otherwise (torch raises only later, after features are
cached), so it is enforced at load time.
Editing `configs/base.yaml` moves the hash, which invalidates every artifact keyed to it. This is
why the stale `configs/deploy.yaml` (which still describes an HF-Space/ZeroGPU target) is left
**undisturbed as frozen paperwork** — it is never read at run time.
## 7. Component map
| Layer | Path (in the source repo) | Notes |
|---|---|---|
| Frontend | `frontend/` | static, staged by `scripts/stage_pages.mjs` |
| Gateway | `deploy/render/` | `main.py` exposes `app`; `render.yaml` blueprint |
| Inference entry | `deploy/codespace/serve.py` | serves `build_space_app()` on `$PORT` |
| Inference app | `app/space_app.py` | `build_space_app()`; 4-endpoint contract |
| Composition root | `app/serving.py` | reuses the local serving path |
| Specialists | `specialists/` | vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA) |
| Router | `router/` | frozen MiniLM + trained adapter |
| Config | `configs/base.yaml` | authoritative registry; hashed |
| Artifacts | `artifacts/` | weights, metrics, calibration, reports |
> **Caveat.** `deploy/` inside the monorepo is **stale and untracked**; the deployed sources live in
> three separate private repositories. See [`DEPLOYMENT.md`](DEPLOYMENT.md).
## 8. What is deliberately absent
- **No database, no auth, no queue.** The gateway is stateless by design.
- **No GPU requirement.** Device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`,
never `.cuda()`. The shipped deployment runs CPU-only.
- **No system-level end-to-end benchmark.** There is no measured end-to-end accuracy number, and
none is claimed. Per-specialist metrics exist and are reported in [`BENCHMARKS.md`](BENCHMARKS.md).
|