Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 26,767 Bytes
b10b360 bf13032 b10b360 bf13032 b10b360 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 | # SatQuery AI
**An interactive vision-language assistant for multimodal remote-sensing image analysis.**
Ask a natural-language question about a satellite or aerial image β or a pair of images β and SatQuery
routes it to the right specialist models, collects evidence, and returns a single confidence-scored
result envelope. It runs on CPU, is served from a static frontend, and is live at
**https://satquery.pages.dev**.
<p align="center">
<img src="screenshots/analyze-grounding.png" alt="The SatQuery Analyze console answering a grounding question against live inference" width="820">
</p>
> **Status: research prototype, pre-1.0.** The architecture is frozen. This repository is the
> documented public release of the system, its trained artifacts, and its measured results β
> **including the negative ones.**
---
## Table of contents
- [Motivation](#motivation)
- [What the system supports](#what-the-system-supports)
- [Supported inputs](#supported-inputs)
- [Architecture](#architecture)
- [Routing and the execution trace](#routing-and-the-execution-trace)
- [Real inference vs. the preview path](#real-inference-vs-the-preview-path)
- [Models](#models)
- [Measured results](#measured-results)
- [Live validation](#live-validation)
- [Installation](#installation)
- [Local development](#local-development)
- [Deployment](#deployment)
- [Reproducibility](#reproducibility)
- [Known limitations](#known-limitations)
- [Links](#links)
- [Citation](#citation)
---
## Motivation
Remote-sensing analysis is fragmented. Detecting change between two acquisitions, localising an
object, captioning a scene, answering a question about it, and fusing optical with SAR each live in a
different model, a different preprocessing convention, and a different output schema. Assembling them
into one answer means re-solving the same problems β tiling, band handling, coordinate systems,
confidence β every time.
SatQuery AI explores a single hypothesis: **a small deterministic router plus a shared evidence
contract can make a heterogeneous specialist ensemble behave like one system**, without a large
language model in the control path. The router *understands* the query. A deterministic policy
*decides* which specialists run. The specialists *compute*. The evidence engine *proves* the answer.
Two design rules follow from that, and they are non-negotiable in the codebase:
- **No LLM-generated coordinates. No LLM-generated confidence.** Coordinates come from detection and
segmentation heads; confidence comes from a calibrated scoring path.
- **Every result carries an observable execution trace** β never chain-of-thought.
## What the system supports
Six specialist tasks. All six are reported `available: true` by the live capability contract
(`GET /api/capabilities`, probed 2026-09-25).
| Task | What it answers | Assets |
|---|---|---|
| `vqa` | A free-form question about a single scene | 1 |
| `caption` | A description of a single scene | 1 |
| `grounding` | *Where* is a described object or region β returns boxes | 1 |
| `change` | *What changed* between two co-registered acquisitions β returns change regions | 2 |
| `change_vqa` | A yes/no or short question about a detected change | 2 |
| `optical_sar` | Joint scene classification from an optical + SAR pair | 2 |
## Supported inputs
Confirmed by the implementation, not assumed:
| Modality | Task(s) | Format |
|---|---|---|
| Optical, single image | `vqa`, `caption`, `grounding` | JPEG, PNG, TIFF |
| Temporal optical pair | `change`, `change_vqa` | Two images of **identical dimensions** |
| Optical + SAR pair | `optical_sar` | GeoTIFF/TIFF preferred |
**Modality is inferred server-side from band count**, not from the file extension: `{1, 2}` bands β
SAR, `{3, 4, 8, 11, 12, 13}` bands β optical. The browser cannot read band count, so the console warns
when a submitted pair looks like two ordinary photographs rather than an optical/SAR pair.
**Per-file upload limit: 4,194,304 bytes (4 MiB).** Larger files are refused with HTTP 413 β imagery
must be downscaled first.
## Architecture
This is the **actually deployed** topology. An older direct-client-to-inference design is superseded.
```mermaid
flowchart TD
B["Browser<br/>(static console)"] -->|HTTPS| CF["Cloudflare Pages<br/>satquery.pages.dev"]
CF -->|"HTTPS JSON Β· /api/*"| R["Render<br/>satquery-orchestrator"]
R -->|"outbound long-poll<br/>POST /tunnel/agent"| T{{"outbound tunnel"}}
T --> C["GitHub Codespace<br/>FastAPI inference Β· CPU Β· :8000"]
C --> S["Specialists"]
S --> M["SmolVLM Β· RemoteCLIP Β· STANet-change<br/>CROMA-fusion Β· MiniLM router"]
M --> E["Evidence engine<br/>+ temperature scaling"]
E --> RE["ResultEnvelope"]
RE -->|"tunnel β Render"| B
```
**Why a tunnel.** The inference host runs in a GitHub Codespace. The forwarded-port path is not
reachable for a private repo (it returns HTTP 302), so the orchestrator keeps a **long-poll tunnel**:
the Codespace dials out to `POST /tunnel/agent` and holds the connection; Render queues work onto it.
`transport_mode` is `auto`, and the tunnel is the live transport.
### Repository map
| Path | Contents |
|---|---|
| `app/` | FastAPI inference service and its composition root |
| `core/` | Config registry, evidence engine, contracts |
| `specialists/` | One module per specialist (vqa, caption, grounding, change, optical_sar) |
| `router/` | MiniLM intent router |
| `gateway/` | Render orchestration hub (`/api/*`, CORS, wake flow) |
| `frontend/` | The static console (HTML/CSS/JS) |
| `configs/base.yaml` | The frozen configuration registry β single source of truth |
| `evaluation/`, `training/` | Evaluation harnesses and training entry points |
| `artifacts/` | Trained heads, checkpoints, evaluation outputs, provenance |
| `docs/` | Architecture, models, benchmarks, deployment, limitations |
## Routing and the execution trace
Routing is deliberately two-stage, and the split matters:
1. **`interpret()` β the reading.** A lexical pass over the query produces a *reading*: task intent,
modality, temporal requirement, spatial scope, and expected evidence kind. It is
**asset-count-blind**.
2. **`chooseTask()` β the dispatch.** The reading is combined with the number of attached assets to
decide the task actually dispatched. This is why a reading of `change` with **one** asset dispatches
`change_vqa` β the documented quantifier upgrade.
> **A router bug worth recording.** An earlier revision evaluated the temporal rule before the
> location rule, so *"Where are the built-up areas in this image?"* matched `\bbuilt\b` as a *change*
> marker and `area` inside *"areas"* as a quantifier. With one asset it collapsed to `vqa` and
> answered "River". Fixed on 2026-09-25 in `frontend/assets/js/mission.js`; the fix is covered by
> regression tests and verified live. The same defect existed on a second surface
> (`SQ.policy` in `core.js`) and was fixed the same day.
### The eight execution events
The console renders an execution trace built from **eight events**, emitted by the frontend around
real network calls (`SQ.EVENT_NAMES` in `frontend/assets/js/core.js`):
| # | Event | Emitted when |
|---|---|---|
| 1 | `QUERY_RECEIVED` | The query and assets are accepted |
| 2 | `QUERY_UNDERSTOOD` | `interpret()` has produced the reading |
| 3 | `ROUTE_SELECTED` | `chooseTask()` has selected the dispatched task |
| 4 | `SPECIALIST_STARTED` | The inference request has been issued |
| 5 | `SPECIALIST_COMPLETED` | The specialist has returned |
| 6 | `EVIDENCE_GENERATED` | Evidence items are available |
| 7 | `CONFIDENCE_COMPUTED` | The calibrated confidence is available |
| 8 | `RESULT_ASSEMBLED` | The `ResultEnvelope` is complete |
These are a **frontend** vocabulary driven by observable events β not a backend protocol and not a
model's reasoning trace. On live runs the trace bar reaches **94.4444 %** (17/18) and every node is
marked live; the preview path is the only source of mock-marked nodes.
## Real inference vs. the preview path
- **Real path (production).** With assets attached, the console calls
`POST /api/infer` on the Render orchestrator. Every run returns a real `run_*` identifier from the
inference service. **Live validation recorded 0 mock nodes across 24 live runs.**
- **Preview path.** With *no* files selected, the console renders a labelled illustrative preview so
the interface is explorable without the stack awake. Preview nodes are explicitly marked `is-mock`
and never appear in a live run.
The distinction is observable, not asserted: a live run shows `live Β· N evidence Β· transport β¦` and
zero `.trace__node.is-mock` elements.
## Models
Six trained artifacts are released. **Four are task heads and two are adapters** β none is a complete
standalone model, and each documents its backbone dependency. Full detail: [`docs/MODELS.md`](docs/MODELS.md),
[`MODEL_CARD.md`](MODEL_CARD.md), and the generated [`models/manifest.json`](models/manifest.json).
| Task | Backbone (pinned) | Custom component | Artifact | Size | Eval data | Metric | Status |
|---|---|---|---|---|---|---|---|
| `change` | STANet-style, ResNet-18 encoder, PAM | change head | `head.pt` | 63,231,009 B | LEVIR-CD-256, test n=2048 | pooled IoU **0.8122** Β· macro IoU **0.8457** Β· pooled F1 **0.8964** | **VERIFIED** |
| `grounding` | `chendelong/RemoteCLIP` ViT-B/32 @ `bf1d8a3ccf2d` (frozen) | trainable head (2048β512) | `head.pt` | 12,639,041 B | VRSBench, n=16159 | mean_best_IoU **0.2838** Β· recall@0.5 **0.2198** (canonical) | measured β two protocols |
| `optical_sar` | `antofuller/CROMA` base @ `0dd28e3d633b` | fusion head (2318β512β19) | `head.pt` | 14,427,457 B | BigEarthNet, 19 CLC classes, test n=4000 | accuracy **0.931** Β· macro_F1 **0.434161** | measured β **ruling OPEN** |
| `change_vqa` | as `change` | change-VQA head | `head.pt` | 5,822,809 B | test n=39686 | accuracy **0.697626** Β· macro_F1 **0.378373** | measured β **ruling OPEN** |
| `router` | `sentence-transformers/all-MiniLM-L6-v2` @ `1110a243fdf4` | intent adapter | `adapter.pt` | 211,961 B | val n=86 | accuracy **0.965116** | **TEST NOT RUN** |
| `vqa` / `caption` | `HuggingFaceTB/SmolVLM-500M-Instruct` @ `a7da5b986cb5` | **LoRA** (r=16, Ξ±=32, dropout 0.05) | `adapter_model.safetensors` | 34,798,048 B | frozen 1000-Q subset | exact_match **0.963** Β· F1 **0.96432** | **ACCEPTANCE-REJECTED** |
Backbones are third-party and pinned by `repo_id` + `revision` in `configs/base.yaml`; they are
fetched from the Hugging Face Hub, not redistributed here.
## Measured results
Every number below traces to an artifact, a test, or a live run. **Nothing here is a system-level
benchmark β no such benchmark exists** (see [Known limitations](#known-limitations)).
| Metric | Value | Split / protocol | Source key | Status |
|---|---|---|---|---|
| Change pooled IoU | 0.8122 | LEVIR-CD-256 test, n=2048, thr 0.50 | `metrics.pooled.iou` | **VERIFIED** |
| Change macro IoU | 0.8457 | same | `metrics.macro.miou` | **VERIFIED** |
| Change pooled F1 | 0.8964 | same | `metrics.pooled.f1` | **VERIFIED** |
| Grounding mean_best_IoU (canonical, `head_threshold`) | 0.2838 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
| Grounding recall@0.5 (canonical, `head_threshold`) | 0.2198 | same | `results.head_threshold.recall.0.50` | measured |
| Grounding mean_best_IoU (matched6, `head_threshold`) | 0.2566 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
| Grounding recall@0.5 (matched6, `head_threshold`) | 0.1938 | same | `results.head_threshold.recall.0.50` | measured |
| Grounding `head_argmax` decode (both protocols) | 0.1215 | same | `results.head_argmax.mean_best_iou` | measured β **worse** |
| Grounding zero-shot baseline | 0.0972 | same | `results.zero_shot_matched.mean_best_iou` | measured |
| Optical-SAR accuracy | 0.931 | BigEarthNet, held-out test n=4000 | `accuracy` | measured β ruling **OPEN** |
| Optical-SAR macro_F1 | 0.434161 | same | `macro_f1` | measured β ruling **OPEN** |
| Change-VQA accuracy (**test**) | 0.697626 | test n=39686 | `verification.test_accuracy` | measured β ruling **OPEN** |
| Change-VQA macro_F1 (**test**) | 0.378373 | same | `verification.test_macro_f1` | measured β ruling **OPEN** |
| Change-VQA accuracy (**test2**) | 0.651469 | second test set | `verification.test2_accuracy` | measured β **lower** |
| Change-VQA macro_F1 (**test2**) | 0.372309 | second test set | `verification.test2_macro_f1` | measured β **lower** |
| VLM adapter exact_match | 0.963 | frozen 1000-Q subset | `artifacts/vlm/phase6_closure.json` | USABLE_VERIFIED β **ACCEPTANCE-REJECTED** |
| VLM adapter F1 | 0.96432 | same | same | USABLE_VERIFIED β **ACCEPTANCE-REJECTED** |
| Router **overall ungated** accuracy | 0.965116 | val, n=86, corpus-limited | `overall_ungated_accuracy` | **TEST NOT RUN** |
| System-level end-to-end benchmark | β | β | β | **NOT RUN β none exists** |
### Grounding: three decode variants, two protocols
The grounding head is evaluated under **two matching protocols** (canonical, matched6) and **three
decode variants**. Quoting a single number would misrepresent the result, so all of them are listed:
| Decode | canonical mean_best_IoU | matched6 mean_best_IoU |
|---|---|---|
| `head_threshold` (the headline number) | **0.2838** | **0.2566** |
| `head_argmax` | 0.1215 | 0.1215 |
| `zero_shot_matched` (baseline, no head) | 0.0972 | 0.0972 |
The head clears the zero-shot baseline, but only the threshold decode is meaningfully above it β the
argmax decode (0.1215) is barely better than zero-shot. The absolute level is modest either way:
**grounding is useful, not solved.**
### Fusion: measured but the ruling is open
Optical-SAR fusion reaches **0.931 accuracy** on a 19-class held-out set of 4,000 β but only
**0.434 macro_F1**. Those two numbers describe very different things: the model is accurate on
frequent classes and weak on rare ones. The acceptance ruling for this head is **OPEN**, and the
headline accuracy must never be quoted without the macro_F1 beside it.
### Change-VQA: two test sets, and they disagree
`artifacts/change_vqa/run/PROMOTION.json` records **two** test evaluations:
| Split | accuracy | macro_F1 |
|---|---|---|
| `test` | 0.697626 | 0.378373 |
| `test2` | **0.651469** | 0.372309 |
The `test` numbers are the higher pair. Both are reported here; quoting only `test` would overstate
the result. The acceptance ruling is **OPEN**.
### Phase 6 / VLM: deployment success β model acceptance
The Phase-6 SmolVLM LoRA adapter reaches **exact_match 0.963** and **F1 0.96432** on a frozen
1,000-question subset. It is marked **USABLE_VERIFIED** and **ACCEPTANCE-REJECTED**.
Those two verdicts are not in conflict, and the distinction is the point:
- **USABLE_VERIFIED** β the adapter loads, runs, and produces the measured numbers in the deployed
pipeline.
- **ACCEPTANCE-REJECTED** β the change did not clear the project's own pre-registered acceptance bar.
A model can be a working engineering artifact and a rejected research result at the same time.
This release keeps both labels.
### Calibration: it got worse, and we say so
The `change_vqa` confidence path applies temperature scaling (`T = 0.9773`). Measured on the
validation split (n=16441):
| | ECE |
|---|---|
| Before temperature scaling | **0.013755** |
| After temperature scaling | **0.014929** |
**Calibration did not improve β it moved slightly worse.** The scaling is retained because it is part
of the frozen configuration, not because it helped. The reliability curve plotted on the Benchmark page
is explicitly labelled as the **pre-scaling** diagram so a reader cannot mistake it for the calibrated
result.
## Live validation
Validation drove the **production site** in a headed browser, one upload per case, with per-case
screenshots and recorded run identifiers.
| Property | Result |
|---|---|
| Independent full passes | **3** |
| Cases per pass | 8 (6 regression + 2 defect) |
| Passes at 8/8 | **3 of 3** |
| Live runs executed | **24** |
| Correct dispatches | **24** |
| Mock-node contamination | **0** on every live run |
| Trace fill | 94.4444 % on every live run |
| Frontend regression suite | **106 passed** (`tests/unit/test_frontend_live_wiring.py`) |
Each pass produced **fresh run identifiers** β no run id is shared between passes.
The harness asserts the form state **before** dispatch β that the query box really holds the intended
query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. This matters:
an earlier harness revision typed with synthetic key events that Chrome silently drops when the window
lacks OS focus, so it dispatched the page's *default* query and still recorded a "result". The
assertions exist because of that failure.
### Representative real run IDs
One full pass (pass 3 of 3) β the same pass the screenshots below are drawn from. Run identifiers
are fresh on every pass; the other two passes recorded different ids.
| Case | Query | Dispatched | Run ID |
|---|---|---|---|
| vqa | What type of terrain dominates this scene? | `vqa` | `run_fef26e91e7e6` |
| caption | Describe the main visual characteristics of this scene. | `caption` | `run_96281bdfcc08` |
| grounding | Where are the visible buildings in this image? | `grounding` | `run_e49adc8d319f` |
| change | What changed between the earlier and later image? | `change` | `run_aedc59cbcdc9` |
| change_vqa | Did the coastline advance between the two observations? | `change_vqa` | `run_62ca98d510be` |
| optical_sar | β¦combining the optical and SAR observations? | `optical_sar` | `run_beacf6aa4e21` |
| **grounding** | **Where are the built-up areas in this image?** | **`grounding`** | **`run_467ffa406f22`** |
| **grounding** | **Where is the new airport?** | **`grounding`** | **`run_46980ba55c62`** |
The last two are the router-defect queries. Both previously collapsed to `vqa` and answered "River".
### Screenshots
Eight captures from the post-fix live run (headed browser, 1384Γ855, one upload per case). Each
panel shows the run identifier, the frozen config hash `78f1e3700da15aa1`, and the evidence list
returned by the specialist β nothing is mocked.
| | |
|---|---|
|  |  |
| **Grounding** β "Where are the built-up areas in this image?" β the fixed router defect (`run_467ffa406f22`, dispatched `grounding`, not `vqa`) | **Optical-SAR** fusion on a real optical/SAR GeoTIFF pair (`run_beacf6aa4e21`, fused class 18) |
|  |  |
| **Grounding** β "Where is the new airport?" β second defect query (`run_46980ba55c62`, dispatched `grounding`) | **Grounding** β "Where are the visible buildings in this image?" (`run_e49adc8d319f`) |
|  |  |
| **Change** detection on a same-shape temporal pair | **Caption** of a single scene (`run_96281bdfcc08`, calibrated confidence 1.000) |
|  |  |
| **VQA** β "What type of terrain dominates this scene?" | **Change-VQA** β "Did the coastline advance between the two observations?" |
All eight are reproduced byte-for-byte in the evidence archive (Phase 7) with SHA-256 recorded in
`RELEASE_MANIFEST.md`.
## Installation
Python 3.11+ and a CPU are sufficient. No CUDA requirement β device is selected via
`SATQUERY_DEVICE`; all placement is `.to(device)`, never `.cuda()`.
```bash
git clone https://github.com/Anish-lab-blip/SatQuery-AI
cd SatQuery-AI
python -m venv .venv
source .venv/Scripts/activate # Windows git-bash; use .venv/bin/activate on Linux/macOS
pip install -r requirements.txt
```
Backbones are fetched from the Hugging Face Hub on first use, pinned by revision in
`configs/base.yaml`. The frozen config hash is **`78f1e3700da15aa1`** β the loader refuses to run a
config that violates the recorded invariants (for example `fusion.input_dim == 3*encoder_dim + 12 + 2`).
## Local development
```bash
# Inference service, CPU (this is the launcher the Codespace runs)
PORT=8000 python deploy/codespace/serve.py
# Health
curl localhost:8000/v1/health
```
The frontend is fully static and needs no build step to serve locally:
```bash
python -m http.server 5500 --directory frontend
```
Run the frontend regression suite:
```bash
python -m pytest tests/unit/test_frontend_live_wiring.py -q
```
## Deployment
The live topology is **Cloudflare Pages β Render β outbound tunnel β GitHub Codespace**.
| Layer | Role | Source |
|---|---|---|
| Cloudflare Pages | Static frontend at **https://satquery.pages.dev** | `frontend/` |
| Render | Orchestrator / API gateway, `/api/*`, CORS, wake flow | `gateway/` |
| GitHub Codespace | FastAPI inference host, CPU, port 8000 | `app/` |
| Hugging Face | Model cards, released artifacts, checksums | this release |
Deployment sources are **separate repositories** from this release. The wake flow is: Cloudflare β
Render β start the Codespace if stopped β poll `/v1/health` β surface *"Waking inference engineβ¦"* β
`POST /infer` β result.
### Deployment caveats
- **Cold start.** The inference host may be stopped when idle. The first request after a cold start
can exceed the client timeout while weights are fetched; a retry a few seconds later normally
succeeds. Warm the stack before any demonstration and confirm
`GET /api/health` reports `tunnel.agent_connected: true`.
- **Tunnel gaps.** The tunnel agent can be briefly absent. A request issued during such a gap may hang
or return HTTP 504. **This is not fixed in production** β a prepared patch
(`forward_unavailable` 503 / `upstream_timeout` 504 plus a `codespace_name` fix) exists and is
documented, but it was deliberately not deployed. Root cause: in `auto` transport mode a tunnel
timeout falls through to the forwarded-port path, which then spends the 120 s wake timeout on an
HTTP 302 β the observed ~249 s failure.
- **`codespace_name`** is still reported with a trailing newline by `/api/health` (cosmetic; the wake
path strips it).
## Reproducibility
1. **Configuration.** `configs/base.yaml` is the single registry; no magic numbers in Python. Its hash
is recorded in every execution trace. Frozen hash: `78f1e3700da15aa1`.
2. **Backbones.** Pinned by `repo_id` + `revision`, never by floating tag:
`HuggingFaceTB/SmolVLM-500M-Instruct@a7da5b986cb5`,
`chendelong/RemoteCLIP@bf1d8a3ccf2d`,
`sentence-transformers/all-MiniLM-L6-v2@1110a243fdf4`,
`antofuller/CROMA@0dd28e3d633b`.
3. **Released artifacts.** `models/manifest.json` and `models/checksums.sha256` are **generated from
the actual files** β never hand-typed. Verify with `sha256sum -c models/checksums.sha256`.
4. **Splits.** LEVIR-CD-256: train 7120 / val 1024 / test 2048. Grounding: VRSBench n=16159.
Fusion: held-out test n=4000. Change-VQA: test n=39686. Leakage isolation is by `scene_id`.
5. **Prompts** are versioned files, frozen before benchmark evaluation.
6. **Negative results are preserved.** Rejected and open rulings are recorded, not removed.
## Known limitations
1. **No system-level end-to-end benchmark exists.** Per-specialist metrics are real; a single
end-to-end number is **NOT RUN**.
2. **Router accuracy is validation-only** (n=86, corpus-limited). Its **test set was never run**.
3. **Optical-SAR fusion returns a bare class index** (`class_18`), not a human-readable label. The
modality accounting in the response confirms the right channels reached the fusion head, but the
presentation is not user-facing.
4. **VQA is weak-but-related.** Asked what terrain dominates a scene, it answers "Grassland".
5. **Fusion macro_F1 is low (0.434)** against 0.931 accuracy β rare classes are poorly handled.
6. **Grounding IoU is modest** (0.2838 canonical) β useful, not solved.
7. **Calibration makes ECE slightly worse**, and is retained only because it is part of the frozen
configuration.
8. **Router lexical residuals.** *"What is the new runway?"* reads `change` rather than `vqa` (the
`new`-as-change heuristic fires outside `where` questions), and *"How much built-up area was
added?"* reads `vqa` (under-trigger). A lexical router cannot cleanly separate "the new X" from
"what's new"; a trained intent router exists in `artifacts/router/` but is not attached.
9. **B-07 tunnel gaps are not fixed in production** (see Deployment caveats).
10. **No license has been selected** for this repository. Until one is, the artifacts carry
`license: unknown` and no reuse rights should be assumed. This is an open owner decision.
11. **The Anatomy of a Run page** renders a recorded run whose plate uses the 720Γ720 variant of an
image analysed at 730Γ730 β identical content, scaled by the canvas, but the "actual analysed
image" wording is slightly loose.
## Links
| | |
|---|---|
| Live demo | https://satquery.pages.dev |
| GitHub | https://github.com/Anish-lab-blip/SatQuery-AI |
| Hugging Face | https://huggingface.co/thundercode/SatQuery |
Third-party models this work builds on (pinned, not redistributed):
| Model | Revision | Role |
|---|---|---|
| [`HuggingFaceTB/SmolVLM-500M-Instruct`](https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct) | `a7da5b986cb5` | VQA + captioning backbone |
| [`chendelong/RemoteCLIP`](https://huggingface.co/chendelong/RemoteCLIP) | `bf1d8a3ccf2d` | remote-sensing grounding encoder |
| [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | `1110a243fdf4` | router embedding |
| [`antofuller/CROMA`](https://huggingface.co/antofuller/CROMA) | `0dd28e3d633b` | optical/SAR fusion encoder |
Datasets referenced by the evaluations: LEVIR-CD-256 (change), VRSBench (grounding),
BigEarthNet (optical-SAR fusion, 19 CLC classes). No dataset is redistributed here.
## Citation
No paper accompanies this release. Until one exists, cite the repository:
```bibtex
@misc{satqueryai2026,
title = {SatQuery AI: an interactive vision-language assistant for
multimodal remote-sensing image analysis},
author = {SatQuery AI contributors},
year = {2026},
url = {https://github.com/Anish-lab-blip/SatQuery-AI}
}
```
## License
**Not yet selected.** See limitation 10. Backbone models remain under their own upstream licenses.
|