SatQuery / README.md
thundercode's picture
release: add README.md
b10b360 verified
|
Raw History Blame
26.8 kB

SatQuery AI

An interactive vision-language assistant for multimodal remote-sensing image analysis.

Ask a natural-language question about a satellite or aerial image β€” or a pair of images β€” and SatQuery routes it to the right specialist models, collects evidence, and returns a single confidence-scored result envelope. It runs on CPU, is served from a static frontend, and is live at https://satquery.pages.dev.

The SatQuery Analyze console answering a grounding question against live inference

Status: research prototype, pre-1.0. The architecture is frozen. This repository is the documented public release of the system, its trained artifacts, and its measured results β€” including the negative ones.


Table of contents


Motivation

Remote-sensing analysis is fragmented. Detecting change between two acquisitions, localising an object, captioning a scene, answering a question about it, and fusing optical with SAR each live in a different model, a different preprocessing convention, and a different output schema. Assembling them into one answer means re-solving the same problems β€” tiling, band handling, coordinate systems, confidence β€” every time.

SatQuery AI explores a single hypothesis: a small deterministic router plus a shared evidence contract can make a heterogeneous specialist ensemble behave like one system, without a large language model in the control path. The router understands the query. A deterministic policy decides which specialists run. The specialists compute. The evidence engine proves the answer.

Two design rules follow from that, and they are non-negotiable in the codebase:

  • No LLM-generated coordinates. No LLM-generated confidence. Coordinates come from detection and segmentation heads; confidence comes from a calibrated scoring path.
  • Every result carries an observable execution trace β€” never chain-of-thought.

What the system supports

Six specialist tasks. All six are reported available: true by the live capability contract (GET /api/capabilities, probed 2026-09-25).

Task What it answers Assets
vqa A free-form question about a single scene 1
caption A description of a single scene 1
grounding Where is a described object or region β€” returns boxes 1
change What changed between two co-registered acquisitions β€” returns change regions 2
change_vqa A yes/no or short question about a detected change 2
optical_sar Joint scene classification from an optical + SAR pair 2

Supported inputs

Confirmed by the implementation, not assumed:

Modality Task(s) Format
Optical, single image vqa, caption, grounding JPEG, PNG, TIFF
Temporal optical pair change, change_vqa Two images of identical dimensions
Optical + SAR pair optical_sar GeoTIFF/TIFF preferred

Modality is inferred server-side from band count, not from the file extension: {1, 2} bands β‡’ SAR, {3, 4, 8, 11, 12, 13} bands β‡’ optical. The browser cannot read band count, so the console warns when a submitted pair looks like two ordinary photographs rather than an optical/SAR pair.

Per-file upload limit: 4,194,304 bytes (4 MiB). Larger files are refused with HTTP 413 β€” imagery must be downscaled first.

Architecture

This is the actually deployed topology. An older direct-client-to-inference design is superseded.

flowchart TD
  B["Browser<br/>(static console)"] -->|HTTPS| CF["Cloudflare Pages<br/>satquery.pages.dev"]
  CF -->|"HTTPS JSON Β· /api/*"| R["Render<br/>satquery-orchestrator"]
  R -->|"outbound long-poll<br/>POST /tunnel/agent"| T{{"outbound tunnel"}}
  T --> C["GitHub Codespace<br/>FastAPI inference Β· CPU Β· :8000"]
  C --> S["Specialists"]
  S --> M["SmolVLM Β· RemoteCLIP Β· STANet-change<br/>CROMA-fusion Β· MiniLM router"]
  M --> E["Evidence engine<br/>+ temperature scaling"]
  E --> RE["ResultEnvelope"]
  RE -->|"tunnel β†’ Render"| B

Why a tunnel. The inference host runs in a GitHub Codespace. The forwarded-port path is not reachable for a private repo (it returns HTTP 302), so the orchestrator keeps a long-poll tunnel: the Codespace dials out to POST /tunnel/agent and holds the connection; Render queues work onto it. transport_mode is auto, and the tunnel is the live transport.

Repository map

Path Contents
app/ FastAPI inference service and its composition root
core/ Config registry, evidence engine, contracts
specialists/ One module per specialist (vqa, caption, grounding, change, optical_sar)
router/ MiniLM intent router
gateway/ Render orchestration hub (/api/*, CORS, wake flow)
frontend/ The static console (HTML/CSS/JS)
configs/base.yaml The frozen configuration registry β€” single source of truth
evaluation/, training/ Evaluation harnesses and training entry points
artifacts/ Trained heads, checkpoints, evaluation outputs, provenance
docs/ Architecture, models, benchmarks, deployment, limitations

Routing and the execution trace

Routing is deliberately two-stage, and the split matters:

  1. interpret() β€” the reading. A lexical pass over the query produces a reading: task intent, modality, temporal requirement, spatial scope, and expected evidence kind. It is asset-count-blind.
  2. chooseTask() β€” the dispatch. The reading is combined with the number of attached assets to decide the task actually dispatched. This is why a reading of change with one asset dispatches change_vqa β€” the documented quantifier upgrade.

A router bug worth recording. An earlier revision evaluated the temporal rule before the location rule, so "Where are the built-up areas in this image?" matched \bbuilt\b as a change marker and area inside "areas" as a quantifier. With one asset it collapsed to vqa and answered "River". Fixed on 2026-09-25 in frontend/assets/js/mission.js; the fix is covered by regression tests and verified live. The same defect existed on a second surface (SQ.policy in core.js) and was fixed the same day.

The eight execution events

The console renders an execution trace built from eight events, emitted by the frontend around real network calls (SQ.EVENT_NAMES in frontend/assets/js/core.js):

# Event Emitted when
1 QUERY_RECEIVED The query and assets are accepted
2 QUERY_UNDERSTOOD interpret() has produced the reading
3 ROUTE_SELECTED chooseTask() has selected the dispatched task
4 SPECIALIST_STARTED The inference request has been issued
5 SPECIALIST_COMPLETED The specialist has returned
6 EVIDENCE_GENERATED Evidence items are available
7 CONFIDENCE_COMPUTED The calibrated confidence is available
8 RESULT_ASSEMBLED The ResultEnvelope is complete

These are a frontend vocabulary driven by observable events β€” not a backend protocol and not a model's reasoning trace. On live runs the trace bar reaches 94.4444 % (17/18) and every node is marked live; the preview path is the only source of mock-marked nodes.

Real inference vs. the preview path

  • Real path (production). With assets attached, the console calls POST /api/infer on the Render orchestrator. Every run returns a real run_* identifier from the inference service. Live validation recorded 0 mock nodes across 24 live runs.
  • Preview path. With no files selected, the console renders a labelled illustrative preview so the interface is explorable without the stack awake. Preview nodes are explicitly marked is-mock and never appear in a live run.

The distinction is observable, not asserted: a live run shows live Β· N evidence Β· transport … and zero .trace__node.is-mock elements.

Models

Six trained artifacts are released. Four are task heads and two are adapters β€” none is a complete standalone model, and each documents its backbone dependency. Full detail: docs/MODELS.md, MODEL_CARD.md, and the generated models/manifest.json.

Task Backbone (pinned) Custom component Artifact Size Eval data Metric Status
change STANet-style, ResNet-18 encoder, PAM change head head.pt 63,231,009 B LEVIR-CD-256, test n=2048 pooled IoU 0.8122 Β· macro IoU 0.8457 Β· pooled F1 0.8964 VERIFIED
grounding chendelong/RemoteCLIP ViT-B/32 @ bf1d8a3ccf2d (frozen) trainable head (2048β†’512) head.pt 12,639,041 B VRSBench, n=16159 mean_best_IoU 0.2838 Β· recall@0.5 0.2198 (canonical) measured β€” two protocols
optical_sar antofuller/CROMA base @ 0dd28e3d633b fusion head (2318β†’512β†’19) head.pt 14,427,457 B BigEarthNet, 19 CLC classes, test n=4000 accuracy 0.931 Β· macro_F1 0.434161 measured β€” ruling OPEN
change_vqa as change change-VQA head head.pt 5,822,809 B test n=39686 accuracy 0.697626 Β· macro_F1 0.378373 measured β€” ruling OPEN
router sentence-transformers/all-MiniLM-L6-v2 @ 1110a243fdf4 intent adapter adapter.pt 211,961 B val n=86 accuracy 0.965116 TEST NOT RUN
vqa / caption HuggingFaceTB/SmolVLM-500M-Instruct @ a7da5b986cb5 LoRA (r=16, Ξ±=32, dropout 0.05) adapter_model.safetensors 34,798,048 B frozen 1000-Q subset exact_match 0.963 Β· F1 0.96432 ACCEPTANCE-REJECTED

Backbones are third-party and pinned by repo_id + revision in configs/base.yaml; they are fetched from the Hugging Face Hub, not redistributed here.

Measured results

Every number below traces to an artifact, a test, or a live run. Nothing here is a system-level benchmark β€” no such benchmark exists (see Known limitations).

Metric Value Split / protocol Source key Status
Change pooled IoU 0.8122 LEVIR-CD-256 test, n=2048, thr 0.50 metrics.pooled.iou VERIFIED
Change macro IoU 0.8457 same metrics.macro.miou VERIFIED
Change pooled F1 0.8964 same metrics.pooled.f1 VERIFIED
Grounding mean_best_IoU (canonical, head_threshold) 0.2838 VRSBench, n=16159 results.head_threshold.mean_best_iou measured
Grounding recall@0.5 (canonical, head_threshold) 0.2198 same results.head_threshold.recall.0.50 measured
Grounding mean_best_IoU (matched6, head_threshold) 0.2566 VRSBench, n=16159 results.head_threshold.mean_best_iou measured
Grounding recall@0.5 (matched6, head_threshold) 0.1938 same results.head_threshold.recall.0.50 measured
Grounding head_argmax decode (both protocols) 0.1215 same results.head_argmax.mean_best_iou measured β€” worse
Grounding zero-shot baseline 0.0972 same results.zero_shot_matched.mean_best_iou measured
Optical-SAR accuracy 0.931 BigEarthNet, held-out test n=4000 accuracy measured β€” ruling OPEN
Optical-SAR macro_F1 0.434161 same macro_f1 measured β€” ruling OPEN
Change-VQA accuracy (test) 0.697626 test n=39686 verification.test_accuracy measured β€” ruling OPEN
Change-VQA macro_F1 (test) 0.378373 same verification.test_macro_f1 measured β€” ruling OPEN
Change-VQA accuracy (test2) 0.651469 second test set verification.test2_accuracy measured β€” lower
Change-VQA macro_F1 (test2) 0.372309 second test set verification.test2_macro_f1 measured β€” lower
VLM adapter exact_match 0.963 frozen 1000-Q subset artifacts/vlm/phase6_closure.json USABLE_VERIFIED β€” ACCEPTANCE-REJECTED
VLM adapter F1 0.96432 same same USABLE_VERIFIED β€” ACCEPTANCE-REJECTED
Router overall ungated accuracy 0.965116 val, n=86, corpus-limited overall_ungated_accuracy TEST NOT RUN
System-level end-to-end benchmark β€” β€” β€” NOT RUN β€” none exists

Grounding: three decode variants, two protocols

The grounding head is evaluated under two matching protocols (canonical, matched6) and three decode variants. Quoting a single number would misrepresent the result, so all of them are listed:

Decode canonical mean_best_IoU matched6 mean_best_IoU
head_threshold (the headline number) 0.2838 0.2566
head_argmax 0.1215 0.1215
zero_shot_matched (baseline, no head) 0.0972 0.0972

The head clears the zero-shot baseline, but only the threshold decode is meaningfully above it β€” the argmax decode (0.1215) is barely better than zero-shot. The absolute level is modest either way: grounding is useful, not solved.

Fusion: measured but the ruling is open

Optical-SAR fusion reaches 0.931 accuracy on a 19-class held-out set of 4,000 β€” but only 0.434 macro_F1. Those two numbers describe very different things: the model is accurate on frequent classes and weak on rare ones. The acceptance ruling for this head is OPEN, and the headline accuracy must never be quoted without the macro_F1 beside it.

Change-VQA: two test sets, and they disagree

artifacts/change_vqa/run/PROMOTION.json records two test evaluations:

Split accuracy macro_F1
test 0.697626 0.378373
test2 0.651469 0.372309

The test numbers are the higher pair. Both are reported here; quoting only test would overstate the result. The acceptance ruling is OPEN.

Phase 6 / VLM: deployment success β‰  model acceptance

The Phase-6 SmolVLM LoRA adapter reaches exact_match 0.963 and F1 0.96432 on a frozen 1,000-question subset. It is marked USABLE_VERIFIED and ACCEPTANCE-REJECTED.

Those two verdicts are not in conflict, and the distinction is the point:

  • USABLE_VERIFIED β€” the adapter loads, runs, and produces the measured numbers in the deployed pipeline.
  • ACCEPTANCE-REJECTED β€” the change did not clear the project's own pre-registered acceptance bar.

A model can be a working engineering artifact and a rejected research result at the same time. This release keeps both labels.

Calibration: it got worse, and we say so

The change_vqa confidence path applies temperature scaling (T = 0.9773). Measured on the validation split (n=16441):

ECE
Before temperature scaling 0.013755
After temperature scaling 0.014929

Calibration did not improve β€” it moved slightly worse. The scaling is retained because it is part of the frozen configuration, not because it helped. The reliability curve plotted on the Benchmark page is explicitly labelled as the pre-scaling diagram so a reader cannot mistake it for the calibrated result.

Live validation

Validation drove the production site in a headed browser, one upload per case, with per-case screenshots and recorded run identifiers.

Property Result
Independent full passes 3
Cases per pass 8 (6 regression + 2 defect)
Passes at 8/8 3 of 3
Live runs executed 24
Correct dispatches 24
Mock-node contamination 0 on every live run
Trace fill 94.4444 % on every live run
Frontend regression suite 106 passed (tests/unit/test_frontend_live_wiring.py)

Each pass produced fresh run identifiers β€” no run id is shared between passes.

The harness asserts the form state before dispatch β€” that the query box really holds the intended query, that #obsTail reads ready, and that both frames are attached for pair tasks. This matters: an earlier harness revision typed with synthetic key events that Chrome silently drops when the window lacks OS focus, so it dispatched the page's default query and still recorded a "result". The assertions exist because of that failure.

Representative real run IDs

One full pass (pass 3 of 3) β€” the same pass the screenshots below are drawn from. Run identifiers are fresh on every pass; the other two passes recorded different ids.

Case Query Dispatched Run ID
vqa What type of terrain dominates this scene? vqa run_fef26e91e7e6
caption Describe the main visual characteristics of this scene. caption run_96281bdfcc08
grounding Where are the visible buildings in this image? grounding run_e49adc8d319f
change What changed between the earlier and later image? change run_aedc59cbcdc9
change_vqa Did the coastline advance between the two observations? change_vqa run_62ca98d510be
optical_sar …combining the optical and SAR observations? optical_sar run_beacf6aa4e21
grounding Where are the built-up areas in this image? grounding run_467ffa406f22
grounding Where is the new airport? grounding run_46980ba55c62

The last two are the router-defect queries. Both previously collapsed to vqa and answered "River".

Screenshots

Eight captures from the post-fix live run (headed browser, 1384Γ—855, one upload per case). Each panel shows the run identifier, the frozen config hash 78f1e3700da15aa1, and the evidence list returned by the specialist β€” nothing is mocked.

Grounding β€” built-up areas Optical-SAR
Grounding β€” "Where are the built-up areas in this image?" β€” the fixed router defect (run_467ffa406f22, dispatched grounding, not vqa) Optical-SAR fusion on a real optical/SAR GeoTIFF pair (run_beacf6aa4e21, fused class 18)
Grounding β€” new airport Grounding β€” visible buildings
Grounding β€” "Where is the new airport?" β€” second defect query (run_46980ba55c62, dispatched grounding) Grounding β€” "Where are the visible buildings in this image?" (run_e49adc8d319f)
Change Caption
Change detection on a same-shape temporal pair Caption of a single scene (run_96281bdfcc08, calibrated confidence 1.000)
VQA Change-VQA
VQA β€” "What type of terrain dominates this scene?" Change-VQA β€” "Did the coastline advance between the two observations?"

All eight are reproduced byte-for-byte in the evidence archive (Phase 7) with SHA-256 recorded in RELEASE_MANIFEST.md.

Installation

Python 3.11+ and a CPU are sufficient. No CUDA requirement β€” device is selected via SATQUERY_DEVICE; all placement is .to(device), never .cuda().

git clone https://github.com/Anish-lab-blip/SatQuery-AI
cd SatQuery-AI
python -m venv .venv
source .venv/Scripts/activate      # Windows git-bash; use .venv/bin/activate on Linux/macOS
pip install -r requirements.txt

Backbones are fetched from the Hugging Face Hub on first use, pinned by revision in configs/base.yaml. The frozen config hash is 78f1e3700da15aa1 β€” the loader refuses to run a config that violates the recorded invariants (for example fusion.input_dim == 3*encoder_dim + 12 + 2).

Local development

# Inference service, CPU (this is the launcher the Codespace runs)
PORT=8000 python deploy/codespace/serve.py

# Health
curl localhost:8000/v1/health

The frontend is fully static and needs no build step to serve locally:

python -m http.server 5500 --directory frontend

Run the frontend regression suite:

python -m pytest tests/unit/test_frontend_live_wiring.py -q

Deployment

The live topology is Cloudflare Pages β†’ Render β†’ outbound tunnel β†’ GitHub Codespace.

Layer Role Source
Cloudflare Pages Static frontend at https://satquery.pages.dev frontend/
Render Orchestrator / API gateway, /api/*, CORS, wake flow gateway/
GitHub Codespace FastAPI inference host, CPU, port 8000 app/
Hugging Face Model cards, released artifacts, checksums this release

Deployment sources are separate repositories from this release. The wake flow is: Cloudflare β†’ Render β†’ start the Codespace if stopped β†’ poll /v1/health β†’ surface "Waking inference engine…" β†’ POST /infer β†’ result.

Deployment caveats

  • Cold start. The inference host may be stopped when idle. The first request after a cold start can exceed the client timeout while weights are fetched; a retry a few seconds later normally succeeds. Warm the stack before any demonstration and confirm GET /api/health reports tunnel.agent_connected: true.
  • Tunnel gaps. The tunnel agent can be briefly absent. A request issued during such a gap may hang or return HTTP 504. This is not fixed in production β€” a prepared patch (forward_unavailable 503 / upstream_timeout 504 plus a codespace_name fix) exists and is documented, but it was deliberately not deployed. Root cause: in auto transport mode a tunnel timeout falls through to the forwarded-port path, which then spends the 120 s wake timeout on an HTTP 302 β€” the observed ~249 s failure.
  • codespace_name is still reported with a trailing newline by /api/health (cosmetic; the wake path strips it).

Reproducibility

  1. Configuration. configs/base.yaml is the single registry; no magic numbers in Python. Its hash is recorded in every execution trace. Frozen hash: 78f1e3700da15aa1.
  2. Backbones. Pinned by repo_id + revision, never by floating tag: HuggingFaceTB/SmolVLM-500M-Instruct@a7da5b986cb5, chendelong/RemoteCLIP@bf1d8a3ccf2d, sentence-transformers/all-MiniLM-L6-v2@1110a243fdf4, antofuller/CROMA@0dd28e3d633b.
  3. Released artifacts. models/manifest.json and models/checksums.sha256 are generated from the actual files β€” never hand-typed. Verify with sha256sum -c models/checksums.sha256.
  4. Splits. LEVIR-CD-256: train 7120 / val 1024 / test 2048. Grounding: VRSBench n=16159. Fusion: held-out test n=4000. Change-VQA: test n=39686. Leakage isolation is by scene_id.
  5. Prompts are versioned files, frozen before benchmark evaluation.
  6. Negative results are preserved. Rejected and open rulings are recorded, not removed.

Known limitations

  1. No system-level end-to-end benchmark exists. Per-specialist metrics are real; a single end-to-end number is NOT RUN.
  2. Router accuracy is validation-only (n=86, corpus-limited). Its test set was never run.
  3. Optical-SAR fusion returns a bare class index (class_18), not a human-readable label. The modality accounting in the response confirms the right channels reached the fusion head, but the presentation is not user-facing.
  4. VQA is weak-but-related. Asked what terrain dominates a scene, it answers "Grassland".
  5. Fusion macro_F1 is low (0.434) against 0.931 accuracy β€” rare classes are poorly handled.
  6. Grounding IoU is modest (0.2838 canonical) β€” useful, not solved.
  7. Calibration makes ECE slightly worse, and is retained only because it is part of the frozen configuration.
  8. Router lexical residuals. "What is the new runway?" reads change rather than vqa (the new-as-change heuristic fires outside where questions), and "How much built-up area was added?" reads vqa (under-trigger). A lexical router cannot cleanly separate "the new X" from "what's new"; a trained intent router exists in artifacts/router/ but is not attached.
  9. B-07 tunnel gaps are not fixed in production (see Deployment caveats).
  10. No license has been selected for this repository. Until one is, the artifacts carry license: unknown and no reuse rights should be assumed. This is an open owner decision.
  11. The Anatomy of a Run page renders a recorded run whose plate uses the 720Γ—720 variant of an image analysed at 730Γ—730 β€” identical content, scaled by the canvas, but the "actual analysed image" wording is slightly loose.

Links

Third-party models this work builds on (pinned, not redistributed):

Model Revision Role
HuggingFaceTB/SmolVLM-500M-Instruct a7da5b986cb5 VQA + captioning backbone
chendelong/RemoteCLIP bf1d8a3ccf2d remote-sensing grounding encoder
sentence-transformers/all-MiniLM-L6-v2 1110a243fdf4 router embedding
antofuller/CROMA 0dd28e3d633b optical/SAR fusion encoder

Datasets referenced by the evaluations: LEVIR-CD-256 (change), VRSBench (grounding), BigEarthNet (optical-SAR fusion, 19 CLC classes). No dataset is redistributed here.

Citation

No paper accompanies this release. Until one exists, cite the repository:

@misc{satqueryai2026,
  title  = {SatQuery AI: an interactive vision-language assistant for
            multimodal remote-sensing image analysis},
  author = {SatQuery AI contributors},
  year   = {2026},
  url    = {https://github.com/Anish-lab-blip/SatQuery-AI}
}

License

Not yet selected. See limitation 10. Backbone models remain under their own upstream licenses.