Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download README.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 26.8 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/README.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/README.md
-
curl -L -o README.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/README.md
SatQuery AI
An interactive vision-language assistant for multimodal remote-sensing image analysis.
Ask a natural-language question about a satellite or aerial image β or a pair of images β and SatQuery routes it to the right specialist models, collects evidence, and returns a single confidence-scored result envelope. It runs on CPU, is served from a static frontend, and is live at https://satquery.pages.dev.
Status: research prototype, pre-1.0. The architecture is frozen. This repository is the documented public release of the system, its trained artifacts, and its measured results β including the negative ones.
Table of contents
- Motivation
- What the system supports
- Supported inputs
- Architecture
- Routing and the execution trace
- Real inference vs. the preview path
- Models
- Measured results
- Live validation
- Installation
- Local development
- Deployment
- Reproducibility
- Known limitations
- Links
- Citation
Motivation
Remote-sensing analysis is fragmented. Detecting change between two acquisitions, localising an object, captioning a scene, answering a question about it, and fusing optical with SAR each live in a different model, a different preprocessing convention, and a different output schema. Assembling them into one answer means re-solving the same problems β tiling, band handling, coordinate systems, confidence β every time.
SatQuery AI explores a single hypothesis: a small deterministic router plus a shared evidence contract can make a heterogeneous specialist ensemble behave like one system, without a large language model in the control path. The router understands the query. A deterministic policy decides which specialists run. The specialists compute. The evidence engine proves the answer.
Two design rules follow from that, and they are non-negotiable in the codebase:
- No LLM-generated coordinates. No LLM-generated confidence. Coordinates come from detection and segmentation heads; confidence comes from a calibrated scoring path.
- Every result carries an observable execution trace β never chain-of-thought.
What the system supports
Six specialist tasks. All six are reported available: true by the live capability contract
(GET /api/capabilities, probed 2026-09-25).
| Task | What it answers | Assets |
|---|---|---|
vqa |
A free-form question about a single scene | 1 |
caption |
A description of a single scene | 1 |
grounding |
Where is a described object or region β returns boxes | 1 |
change |
What changed between two co-registered acquisitions β returns change regions | 2 |
change_vqa |
A yes/no or short question about a detected change | 2 |
optical_sar |
Joint scene classification from an optical + SAR pair | 2 |
Supported inputs
Confirmed by the implementation, not assumed:
| Modality | Task(s) | Format |
|---|---|---|
| Optical, single image | vqa, caption, grounding |
JPEG, PNG, TIFF |
| Temporal optical pair | change, change_vqa |
Two images of identical dimensions |
| Optical + SAR pair | optical_sar |
GeoTIFF/TIFF preferred |
Modality is inferred server-side from band count, not from the file extension: {1, 2} bands β
SAR, {3, 4, 8, 11, 12, 13} bands β optical. The browser cannot read band count, so the console warns
when a submitted pair looks like two ordinary photographs rather than an optical/SAR pair.
Per-file upload limit: 4,194,304 bytes (4 MiB). Larger files are refused with HTTP 413 β imagery must be downscaled first.
Architecture
This is the actually deployed topology. An older direct-client-to-inference design is superseded.
flowchart TD
B["Browser<br/>(static console)"] -->|HTTPS| CF["Cloudflare Pages<br/>satquery.pages.dev"]
CF -->|"HTTPS JSON Β· /api/*"| R["Render<br/>satquery-orchestrator"]
R -->|"outbound long-poll<br/>POST /tunnel/agent"| T{{"outbound tunnel"}}
T --> C["GitHub Codespace<br/>FastAPI inference Β· CPU Β· :8000"]
C --> S["Specialists"]
S --> M["SmolVLM Β· RemoteCLIP Β· STANet-change<br/>CROMA-fusion Β· MiniLM router"]
M --> E["Evidence engine<br/>+ temperature scaling"]
E --> RE["ResultEnvelope"]
RE -->|"tunnel β Render"| B
Why a tunnel. The inference host runs in a GitHub Codespace. The forwarded-port path is not
reachable for a private repo (it returns HTTP 302), so the orchestrator keeps a long-poll tunnel:
the Codespace dials out to POST /tunnel/agent and holds the connection; Render queues work onto it.
transport_mode is auto, and the tunnel is the live transport.
Repository map
| Path | Contents |
|---|---|
app/ |
FastAPI inference service and its composition root |
core/ |
Config registry, evidence engine, contracts |
specialists/ |
One module per specialist (vqa, caption, grounding, change, optical_sar) |
router/ |
MiniLM intent router |
gateway/ |
Render orchestration hub (/api/*, CORS, wake flow) |
frontend/ |
The static console (HTML/CSS/JS) |
configs/base.yaml |
The frozen configuration registry β single source of truth |
evaluation/, training/ |
Evaluation harnesses and training entry points |
artifacts/ |
Trained heads, checkpoints, evaluation outputs, provenance |
docs/ |
Architecture, models, benchmarks, deployment, limitations |
Routing and the execution trace
Routing is deliberately two-stage, and the split matters:
interpret()β the reading. A lexical pass over the query produces a reading: task intent, modality, temporal requirement, spatial scope, and expected evidence kind. It is asset-count-blind.chooseTask()β the dispatch. The reading is combined with the number of attached assets to decide the task actually dispatched. This is why a reading ofchangewith one asset dispatcheschange_vqaβ the documented quantifier upgrade.
A router bug worth recording. An earlier revision evaluated the temporal rule before the location rule, so "Where are the built-up areas in this image?" matched
\bbuilt\bas a change marker andareainside "areas" as a quantifier. With one asset it collapsed tovqaand answered "River". Fixed on 2026-09-25 infrontend/assets/js/mission.js; the fix is covered by regression tests and verified live. The same defect existed on a second surface (SQ.policyincore.js) and was fixed the same day.
The eight execution events
The console renders an execution trace built from eight events, emitted by the frontend around
real network calls (SQ.EVENT_NAMES in frontend/assets/js/core.js):
| # | Event | Emitted when |
|---|---|---|
| 1 | QUERY_RECEIVED |
The query and assets are accepted |
| 2 | QUERY_UNDERSTOOD |
interpret() has produced the reading |
| 3 | ROUTE_SELECTED |
chooseTask() has selected the dispatched task |
| 4 | SPECIALIST_STARTED |
The inference request has been issued |
| 5 | SPECIALIST_COMPLETED |
The specialist has returned |
| 6 | EVIDENCE_GENERATED |
Evidence items are available |
| 7 | CONFIDENCE_COMPUTED |
The calibrated confidence is available |
| 8 | RESULT_ASSEMBLED |
The ResultEnvelope is complete |
These are a frontend vocabulary driven by observable events β not a backend protocol and not a model's reasoning trace. On live runs the trace bar reaches 94.4444 % (17/18) and every node is marked live; the preview path is the only source of mock-marked nodes.
Real inference vs. the preview path
- Real path (production). With assets attached, the console calls
POST /api/inferon the Render orchestrator. Every run returns a realrun_*identifier from the inference service. Live validation recorded 0 mock nodes across 24 live runs. - Preview path. With no files selected, the console renders a labelled illustrative preview so
the interface is explorable without the stack awake. Preview nodes are explicitly marked
is-mockand never appear in a live run.
The distinction is observable, not asserted: a live run shows live Β· N evidence Β· transport β¦ and
zero .trace__node.is-mock elements.
Models
Six trained artifacts are released. Four are task heads and two are adapters β none is a complete
standalone model, and each documents its backbone dependency. Full detail: docs/MODELS.md,
MODEL_CARD.md, and the generated models/manifest.json.
| Task | Backbone (pinned) | Custom component | Artifact | Size | Eval data | Metric | Status |
|---|---|---|---|---|---|---|---|
change |
STANet-style, ResNet-18 encoder, PAM | change head | head.pt |
63,231,009 B | LEVIR-CD-256, test n=2048 | pooled IoU 0.8122 Β· macro IoU 0.8457 Β· pooled F1 0.8964 | VERIFIED |
grounding |
chendelong/RemoteCLIP ViT-B/32 @ bf1d8a3ccf2d (frozen) |
trainable head (2048β512) | head.pt |
12,639,041 B | VRSBench, n=16159 | mean_best_IoU 0.2838 Β· recall@0.5 0.2198 (canonical) | measured β two protocols |
optical_sar |
antofuller/CROMA base @ 0dd28e3d633b |
fusion head (2318β512β19) | head.pt |
14,427,457 B | BigEarthNet, 19 CLC classes, test n=4000 | accuracy 0.931 Β· macro_F1 0.434161 | measured β ruling OPEN |
change_vqa |
as change |
change-VQA head | head.pt |
5,822,809 B | test n=39686 | accuracy 0.697626 Β· macro_F1 0.378373 | measured β ruling OPEN |
router |
sentence-transformers/all-MiniLM-L6-v2 @ 1110a243fdf4 |
intent adapter | adapter.pt |
211,961 B | val n=86 | accuracy 0.965116 | TEST NOT RUN |
vqa / caption |
HuggingFaceTB/SmolVLM-500M-Instruct @ a7da5b986cb5 |
LoRA (r=16, Ξ±=32, dropout 0.05) | adapter_model.safetensors |
34,798,048 B | frozen 1000-Q subset | exact_match 0.963 Β· F1 0.96432 | ACCEPTANCE-REJECTED |
Backbones are third-party and pinned by repo_id + revision in configs/base.yaml; they are
fetched from the Hugging Face Hub, not redistributed here.
Measured results
Every number below traces to an artifact, a test, or a live run. Nothing here is a system-level benchmark β no such benchmark exists (see Known limitations).
| Metric | Value | Split / protocol | Source key | Status |
|---|---|---|---|---|
| Change pooled IoU | 0.8122 | LEVIR-CD-256 test, n=2048, thr 0.50 | metrics.pooled.iou |
VERIFIED |
| Change macro IoU | 0.8457 | same | metrics.macro.miou |
VERIFIED |
| Change pooled F1 | 0.8964 | same | metrics.pooled.f1 |
VERIFIED |
Grounding mean_best_IoU (canonical, head_threshold) |
0.2838 | VRSBench, n=16159 | results.head_threshold.mean_best_iou |
measured |
Grounding recall@0.5 (canonical, head_threshold) |
0.2198 | same | results.head_threshold.recall.0.50 |
measured |
Grounding mean_best_IoU (matched6, head_threshold) |
0.2566 | VRSBench, n=16159 | results.head_threshold.mean_best_iou |
measured |
Grounding recall@0.5 (matched6, head_threshold) |
0.1938 | same | results.head_threshold.recall.0.50 |
measured |
Grounding head_argmax decode (both protocols) |
0.1215 | same | results.head_argmax.mean_best_iou |
measured β worse |
| Grounding zero-shot baseline | 0.0972 | same | results.zero_shot_matched.mean_best_iou |
measured |
| Optical-SAR accuracy | 0.931 | BigEarthNet, held-out test n=4000 | accuracy |
measured β ruling OPEN |
| Optical-SAR macro_F1 | 0.434161 | same | macro_f1 |
measured β ruling OPEN |
| Change-VQA accuracy (test) | 0.697626 | test n=39686 | verification.test_accuracy |
measured β ruling OPEN |
| Change-VQA macro_F1 (test) | 0.378373 | same | verification.test_macro_f1 |
measured β ruling OPEN |
| Change-VQA accuracy (test2) | 0.651469 | second test set | verification.test2_accuracy |
measured β lower |
| Change-VQA macro_F1 (test2) | 0.372309 | second test set | verification.test2_macro_f1 |
measured β lower |
| VLM adapter exact_match | 0.963 | frozen 1000-Q subset | artifacts/vlm/phase6_closure.json |
USABLE_VERIFIED β ACCEPTANCE-REJECTED |
| VLM adapter F1 | 0.96432 | same | same | USABLE_VERIFIED β ACCEPTANCE-REJECTED |
| Router overall ungated accuracy | 0.965116 | val, n=86, corpus-limited | overall_ungated_accuracy |
TEST NOT RUN |
| System-level end-to-end benchmark | β | β | β | NOT RUN β none exists |
Grounding: three decode variants, two protocols
The grounding head is evaluated under two matching protocols (canonical, matched6) and three decode variants. Quoting a single number would misrepresent the result, so all of them are listed:
| Decode | canonical mean_best_IoU | matched6 mean_best_IoU |
|---|---|---|
head_threshold (the headline number) |
0.2838 | 0.2566 |
head_argmax |
0.1215 | 0.1215 |
zero_shot_matched (baseline, no head) |
0.0972 | 0.0972 |
The head clears the zero-shot baseline, but only the threshold decode is meaningfully above it β the argmax decode (0.1215) is barely better than zero-shot. The absolute level is modest either way: grounding is useful, not solved.
Fusion: measured but the ruling is open
Optical-SAR fusion reaches 0.931 accuracy on a 19-class held-out set of 4,000 β but only 0.434 macro_F1. Those two numbers describe very different things: the model is accurate on frequent classes and weak on rare ones. The acceptance ruling for this head is OPEN, and the headline accuracy must never be quoted without the macro_F1 beside it.
Change-VQA: two test sets, and they disagree
artifacts/change_vqa/run/PROMOTION.json records two test evaluations:
| Split | accuracy | macro_F1 |
|---|---|---|
test |
0.697626 | 0.378373 |
test2 |
0.651469 | 0.372309 |
The test numbers are the higher pair. Both are reported here; quoting only test would overstate
the result. The acceptance ruling is OPEN.
Phase 6 / VLM: deployment success β model acceptance
The Phase-6 SmolVLM LoRA adapter reaches exact_match 0.963 and F1 0.96432 on a frozen 1,000-question subset. It is marked USABLE_VERIFIED and ACCEPTANCE-REJECTED.
Those two verdicts are not in conflict, and the distinction is the point:
- USABLE_VERIFIED β the adapter loads, runs, and produces the measured numbers in the deployed pipeline.
- ACCEPTANCE-REJECTED β the change did not clear the project's own pre-registered acceptance bar.
A model can be a working engineering artifact and a rejected research result at the same time. This release keeps both labels.
Calibration: it got worse, and we say so
The change_vqa confidence path applies temperature scaling (T = 0.9773). Measured on the
validation split (n=16441):
| ECE | |
|---|---|
| Before temperature scaling | 0.013755 |
| After temperature scaling | 0.014929 |
Calibration did not improve β it moved slightly worse. The scaling is retained because it is part of the frozen configuration, not because it helped. The reliability curve plotted on the Benchmark page is explicitly labelled as the pre-scaling diagram so a reader cannot mistake it for the calibrated result.
Live validation
Validation drove the production site in a headed browser, one upload per case, with per-case screenshots and recorded run identifiers.
| Property | Result |
|---|---|
| Independent full passes | 3 |
| Cases per pass | 8 (6 regression + 2 defect) |
| Passes at 8/8 | 3 of 3 |
| Live runs executed | 24 |
| Correct dispatches | 24 |
| Mock-node contamination | 0 on every live run |
| Trace fill | 94.4444 % on every live run |
| Frontend regression suite | 106 passed (tests/unit/test_frontend_live_wiring.py) |
Each pass produced fresh run identifiers β no run id is shared between passes.
The harness asserts the form state before dispatch β that the query box really holds the intended
query, that #obsTail reads ready, and that both frames are attached for pair tasks. This matters:
an earlier harness revision typed with synthetic key events that Chrome silently drops when the window
lacks OS focus, so it dispatched the page's default query and still recorded a "result". The
assertions exist because of that failure.
Representative real run IDs
One full pass (pass 3 of 3) β the same pass the screenshots below are drawn from. Run identifiers are fresh on every pass; the other two passes recorded different ids.
| Case | Query | Dispatched | Run ID |
|---|---|---|---|
| vqa | What type of terrain dominates this scene? | vqa |
run_fef26e91e7e6 |
| caption | Describe the main visual characteristics of this scene. | caption |
run_96281bdfcc08 |
| grounding | Where are the visible buildings in this image? | grounding |
run_e49adc8d319f |
| change | What changed between the earlier and later image? | change |
run_aedc59cbcdc9 |
| change_vqa | Did the coastline advance between the two observations? | change_vqa |
run_62ca98d510be |
| optical_sar | β¦combining the optical and SAR observations? | optical_sar |
run_beacf6aa4e21 |
| grounding | Where are the built-up areas in this image? | grounding |
run_467ffa406f22 |
| grounding | Where is the new airport? | grounding |
run_46980ba55c62 |
The last two are the router-defect queries. Both previously collapsed to vqa and answered "River".
Screenshots
Eight captures from the post-fix live run (headed browser, 1384Γ855, one upload per case). Each
panel shows the run identifier, the frozen config hash 78f1e3700da15aa1, and the evidence list
returned by the specialist β nothing is mocked.
All eight are reproduced byte-for-byte in the evidence archive (Phase 7) with SHA-256 recorded in
RELEASE_MANIFEST.md.
Installation
Python 3.11+ and a CPU are sufficient. No CUDA requirement β device is selected via
SATQUERY_DEVICE; all placement is .to(device), never .cuda().
git clone https://github.com/Anish-lab-blip/SatQuery-AI
cd SatQuery-AI
python -m venv .venv
source .venv/Scripts/activate # Windows git-bash; use .venv/bin/activate on Linux/macOS
pip install -r requirements.txt
Backbones are fetched from the Hugging Face Hub on first use, pinned by revision in
configs/base.yaml. The frozen config hash is 78f1e3700da15aa1 β the loader refuses to run a
config that violates the recorded invariants (for example fusion.input_dim == 3*encoder_dim + 12 + 2).
Local development
# Inference service, CPU (this is the launcher the Codespace runs)
PORT=8000 python deploy/codespace/serve.py
# Health
curl localhost:8000/v1/health
The frontend is fully static and needs no build step to serve locally:
python -m http.server 5500 --directory frontend
Run the frontend regression suite:
python -m pytest tests/unit/test_frontend_live_wiring.py -q
Deployment
The live topology is Cloudflare Pages β Render β outbound tunnel β GitHub Codespace.
| Layer | Role | Source |
|---|---|---|
| Cloudflare Pages | Static frontend at https://satquery.pages.dev | frontend/ |
| Render | Orchestrator / API gateway, /api/*, CORS, wake flow |
gateway/ |
| GitHub Codespace | FastAPI inference host, CPU, port 8000 | app/ |
| Hugging Face | Model cards, released artifacts, checksums | this release |
Deployment sources are separate repositories from this release. The wake flow is: Cloudflare β
Render β start the Codespace if stopped β poll /v1/health β surface "Waking inference engineβ¦" β
POST /infer β result.
Deployment caveats
- Cold start. The inference host may be stopped when idle. The first request after a cold start
can exceed the client timeout while weights are fetched; a retry a few seconds later normally
succeeds. Warm the stack before any demonstration and confirm
GET /api/healthreportstunnel.agent_connected: true. - Tunnel gaps. The tunnel agent can be briefly absent. A request issued during such a gap may hang
or return HTTP 504. This is not fixed in production β a prepared patch
(
forward_unavailable503 /upstream_timeout504 plus acodespace_namefix) exists and is documented, but it was deliberately not deployed. Root cause: inautotransport mode a tunnel timeout falls through to the forwarded-port path, which then spends the 120 s wake timeout on an HTTP 302 β the observed ~249 s failure. codespace_nameis still reported with a trailing newline by/api/health(cosmetic; the wake path strips it).
Reproducibility
- Configuration.
configs/base.yamlis the single registry; no magic numbers in Python. Its hash is recorded in every execution trace. Frozen hash:78f1e3700da15aa1. - Backbones. Pinned by
repo_id+revision, never by floating tag:HuggingFaceTB/SmolVLM-500M-Instruct@a7da5b986cb5,chendelong/RemoteCLIP@bf1d8a3ccf2d,sentence-transformers/all-MiniLM-L6-v2@1110a243fdf4,antofuller/CROMA@0dd28e3d633b. - Released artifacts.
models/manifest.jsonandmodels/checksums.sha256are generated from the actual files β never hand-typed. Verify withsha256sum -c models/checksums.sha256. - Splits. LEVIR-CD-256: train 7120 / val 1024 / test 2048. Grounding: VRSBench n=16159.
Fusion: held-out test n=4000. Change-VQA: test n=39686. Leakage isolation is by
scene_id. - Prompts are versioned files, frozen before benchmark evaluation.
- Negative results are preserved. Rejected and open rulings are recorded, not removed.
Known limitations
- No system-level end-to-end benchmark exists. Per-specialist metrics are real; a single end-to-end number is NOT RUN.
- Router accuracy is validation-only (n=86, corpus-limited). Its test set was never run.
- Optical-SAR fusion returns a bare class index (
class_18), not a human-readable label. The modality accounting in the response confirms the right channels reached the fusion head, but the presentation is not user-facing. - VQA is weak-but-related. Asked what terrain dominates a scene, it answers "Grassland".
- Fusion macro_F1 is low (0.434) against 0.931 accuracy β rare classes are poorly handled.
- Grounding IoU is modest (0.2838 canonical) β useful, not solved.
- Calibration makes ECE slightly worse, and is retained only because it is part of the frozen configuration.
- Router lexical residuals. "What is the new runway?" reads
changerather thanvqa(thenew-as-change heuristic fires outsidewherequestions), and "How much built-up area was added?" readsvqa(under-trigger). A lexical router cannot cleanly separate "the new X" from "what's new"; a trained intent router exists inartifacts/router/but is not attached. - B-07 tunnel gaps are not fixed in production (see Deployment caveats).
- No license has been selected for this repository. Until one is, the artifacts carry
license: unknownand no reuse rights should be assumed. This is an open owner decision. - The Anatomy of a Run page renders a recorded run whose plate uses the 720Γ720 variant of an image analysed at 730Γ730 β identical content, scaled by the canvas, but the "actual analysed image" wording is slightly loose.
Links
| Live demo | https://satquery.pages.dev |
| GitHub | https://github.com/Anish-lab-blip/SatQuery-AI |
| Hugging Face | https://huggingface.co/thundercode/SatQuery |
Third-party models this work builds on (pinned, not redistributed):
| Model | Revision | Role |
|---|---|---|
HuggingFaceTB/SmolVLM-500M-Instruct |
a7da5b986cb5 |
VQA + captioning backbone |
chendelong/RemoteCLIP |
bf1d8a3ccf2d |
remote-sensing grounding encoder |
sentence-transformers/all-MiniLM-L6-v2 |
1110a243fdf4 |
router embedding |
antofuller/CROMA |
0dd28e3d633b |
optical/SAR fusion encoder |
Datasets referenced by the evaluations: LEVIR-CD-256 (change), VRSBench (grounding), BigEarthNet (optical-SAR fusion, 19 CLC classes). No dataset is redistributed here.
Citation
No paper accompanies this release. Until one exists, cite the repository:
@misc{satqueryai2026,
title = {SatQuery AI: an interactive vision-language assistant for
multimodal remote-sensing image analysis},
author = {SatQuery AI contributors},
year = {2026},
url = {https://github.com/Anish-lab-blip/SatQuery-AI}
}
License
Not yet selected. See limitation 10. Backbone models remain under their own upstream licenses.






