Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 10,354 Bytes
df0d288 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 | # Models
**Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` ·
`DEFERRED` · `REJECTED`.
SatQuery AI trains **six** artifacts. Four are task heads, one is a router adapter, one is a LoRA
adapter. Every backbone is **frozen** and publicly pinned by revision in `configs/base.yaml` — the
project trains small modules on top of frozen encoders, not end-to-end networks.
> **This file is the human-readable companion to the machine-generated
> [`../models/manifest.json`](../models/manifest.json) and
> [`../models/checksums.sha256`](../models/checksums.sha256) (Phase 3).** Where the two disagree,
> the generated manifest wins — it is computed from the files, this document is written by hand.
---
## 1. The six trained artifacts
| # | Task | Artifact path | Bytes | Kind | Backbone (frozen) |
|---|---|---|---|---|---|
| 1 | `change` | `artifacts/change/levir_change_v001/head.pt` | 63,231,009 | trained head | STANet-style, ResNet-18 + PAM |
| 2 | `change_vqa` | `artifacts/change_vqa/run/head.pt` | 5,822,809 | trained head | over the change encoder's features |
| 3 | `optical_sar` | `artifacts/optical_sar/fusion_head_production_v001/head.pt` | 14,427,457 | trained head (production) | CROMA-base (frozen), 19-class head |
| 4 | `grounding` | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | 12,639,041 | trained head | RemoteCLIP ViT-B/32 (frozen) |
| 5 | `router` | `artifacts/router/router_adapter_v001/adapter.pt` | 211,961 | trained adapter | `all-MiniLM-L6-v2` (frozen) |
| 6 | `vlm` | `.scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors` | 34,798,048 | **LoRA adapter** (PEFT) | `HuggingFaceTB/SmolVLM-500M-Instruct` (frozen) |
Training checkpoints also exist (`checkpoint_last.pt`, `checkpoint-1500`, `checkpoint-2000`) and are
**not** the released artifacts — they are archived as provenance.
## 2. Backbones — pinned, frozen, never retrained
| Role | Repository | Revision | Notes |
|---|---|---|---|
| Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | 90.9 MB, 22,713,216 params, 384-dim embeddings, tokenizer ceiling 256; truncation set to 128 |
| VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | ~1015 MB safetensors |
| Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` | 605.2 MB; transformer width 768, **projected** dim 512 |
| Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` | 777.6 MB; `encoder_dim` 768; `image_resolution` 120 |
These are resolved from the Hugging Face Hub on first use. **No backbone weights are redistributed**
by this project's release — see §6.
## 3. Per-artifact detail
### 3.1 Change (`change`) — `IMPLEMENTED`, `VERIFIED`
STANet-style Siamese change detector. Encoder ResNet-18, self-attention mode **PAM**, tile size 256,
threshold 0.50, minimum component 32 px. Loss is BCE (0.5) + Dice (0.5). Trained on LEVIR-CD-256
(train 7120 / val 1024 / test 2048).
**Measured** on the LEVIR-CD-256 test split (n = 2048, threshold 0.50):
| Metric | Value |
|---|---|
| pooled IoU | **0.8122** |
| macro IoU | **0.8457** |
| pooled F1 | **0.8964** |
This is the only task whose headline number carries the `VERIFIED` tag, because it is the only one
measured against a single, immutable public test split with a frozen threshold.
### 3.2 Change-VQA (`change_vqa`) — `IMPLEMENTED`, `MEASURED`, ruling **OPEN**
A head that answers natural-language change questions over a temporal pair. It is the dispatch
target for change-style questions when only one asset is attached (see
[`ARCHITECTURE.md`](ARCHITECTURE.md) §4).
**Measured on two test sets — both are reported; quoting only the better one would be selective:**
| Test set | accuracy | macro F1 |
|---|---|---|
| `test` (n = 39,686) | **0.697626** | **0.378373** |
| `test2` | **0.651469** | **0.372309** |
The wide gap between accuracy and macro-F1 means the head is carried by common classes and performs
poorly on rare ones. The ruling is **OPEN** — no promotion/acceptance decision has been recorded.
### 3.3 Optical-SAR fusion (`optical_sar`) — `IMPLEMENTED`, `MEASURED`, ruling **OPEN**
Uses frozen CROMA-base to produce optical (768), SAR (768) and joint (768) embeddings, concatenates
them with the 12 optical and 2 SAR channel descriptors, and feeds a 19-class head:
```
input_dim = 3 * 768 + 12 + 2 = 2318 → hidden 512 → num_classes 19 (BigEarthNet CLC)
```
The **availability mask is consumed by the fusion head, not by CROMA** — CROMA always sees the
canonical channel counts (12 optical, 2 SAR).
**Measured** (production head `fusion_head_production_v001`, pre-registered 115-class protocol,
held-out test n = 4000):
| Metric | Value |
|---|---|
| accuracy | **0.931** |
| macro F1 | **0.434161** |
**Both numbers must travel together.** The high accuracy with a low macro-F1 reflects class
imbalance across 19 classes. The pre-registered metric JSON records `macro_f1_denominator` and
`classes_present`/`classes_absent` so the denominator is auditable. The ruling is **OPEN**.
**Limitation:** the live service returns a bare class index (`class_18`), not a human-readable label.
### 3.4 Grounding (`grounding`) — `IMPLEMENTED`, `MEASURED` — **two protocols**
A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Per-cell feature is
`concat([patch, text, patch·text, global_pool])` = `4 × 512 = 2048` (enforced at config load). Cells
are assigned by ground-truth box centre (`cell_relative` decode). Objectness BCE is weighted **20×**
because only ~1 of 49 cells is positive; unweighted, the optimum collapses to "no object".
Image resolution is frozen at **224** — 448 was evaluated and **rejected** (see §4).
**Measured on VRSBench (n = 16,159), reported under two protocols:**
| Protocol | mean best IoU | recall@0.5 |
|---|---|---|
| canonical (head threshold decode) | **0.2838** | **0.2198** |
| matched6 | **0.2566** | **0.1938** |
Two further decode variants exist and are reported for completeness — a reviewer must be able to see
the whole grid, not one cell of it:
| Variant | mean best IoU |
|---|---|
| head argmax decode (canonical) | **0.1215** |
| zero-shot matched (no trained head) | **0.0972** |
The trained head beats the zero-shot baseline by a wide margin, which is the point of the head; the
absolute IoU is low, which is an honest limitation.
### 3.5 Router (`router`) — `IMPLEMENTED`, `MEASURED`, **TEST NOT RUN**
A **50,822-parameter adapter** over the frozen MiniLM encoder. Because the encoder is frozen,
embeddings are cached and the adapter trains on cached vectors — no GPU required (measured: 20
epochs / 4,096 vectors in 0.28 s on CPU). Splits are by **group** (template / hard-negative family),
never by example; hard-negative families are placed in the test split so their accuracy measures
generalisation, not memorisation.
**Measured:** overall **ungated** task accuracy **0.965116** on the validation split, **n = 86**,
corpus-limited. This number is (a) validation-only, (b) ungated, and (c) small. The router **test**
split was **NOT RUN**. Do not read 0.965116 as a test result.
### 3.6 VLM LoRA (`vlm`) — `IMPLEMENTED`, `MEASURED`, **ACCEPTANCE-REJECTED**
A PEFT LoRA adapter on frozen SmolVLM-500M-Instruct: `peft_type=LORA`, `r=16`, `alpha=32`,
`dropout=0.05`, targeting `model.text_model.*.{q,k,v,o,gate,up,down}_proj`. PEFT 0.19.1.
**Measured** on a frozen 1,000-question subset: exact_match **0.963**, F1 **0.96432** (+49.5 pp over
the unadapted baseline).
**Yet the artifact's status is `CLOSED` with headline `ACCEPTANCE-REJECTED`.** This is not a
contradiction — it is the project's central truthfulness distinction:
- **`USABLE_VERIFIED`** — the adapter demonstrably works (the metrics are real and reproducible).
- **`ACCEPTANCE-REJECTED`** — the adapter is *not accepted* for production promotion, on grounds
recorded in `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`).
The deployed caption/VQA path therefore uses the **unadapted** SmolVLM. Deployment success and model
acceptance are different claims, and this document keeps them apart.
## 4. Rejected and deferred model decisions
| Decision | Outcome | Evidence |
|---|---|---|
| Grounding image resolution 448 vs 224 | **224 chosen; 448 REJECTED** | 448 lost on every axis: mean best IoU −0.0147, recall@0.5 −0.0022, all recall thresholds lower, at 1.59× latency. Paired test: mean diff −0.0147, 95 % CI [−0.0160, −0.0134], t = −22.63; 448 better on 8.5 % of records, worse on 20.9 %. Pre-registered rule and the paired test **agree**. |
| VLM adapter promotion | **REJECTED** | metrics usable, acceptance rejected (§3.6) |
| Calibration | **kept but ineffective** | see §5 |
| optical-SAR / change-VQA rulings | **OPEN** | no decision recorded |
## 5. Calibration — `MEASURED`, **not an improvement**
Temperature scaling is enabled (`confidence.temperature_scaling: true`) with
`calibration_v001.json`. Fitted temperature **T = 0.9772732** on the validation split (n = 16,441).
| Metric | Before | After |
|---|---|---|
| ECE | 0.013755 | **0.014929** |
| NLL | 0.689741 | 0.689631 |
**ECE got worse** (`ece_improvement = −0.001174`). The scaling is retained because it is part of the
frozen configuration, **not** because it helped. The reliability diagram on the Benchmark page is
explicitly labelled **pre-scaling** so a reader cannot mistake it for the calibrated result. This is
recorded as a negative result, not smoothed over.
## 6. Distribution and licensing
- **Backbones are not redistributed.** They are fetched from the Hugging Face Hub at run time, pinned
by revision. Their licences are their own (see each model's HF page).
- **The six trained artifacts are published by this project** on the Hugging Face Hub under
`thundercode/SatQuery`, labelled by kind, each with its backbone dependency documented and each
accompanied by a checksum. See [`../HF_RELEASE_VERIFICATION.md`](../HF_RELEASE_VERIFICATION.md).
- **No licence file exists in the source repository.** This is an **OPEN** item flagged in
[`LIMITATIONS.md`](LIMITATIONS.md); the repository README instructs the owner to select one before
any public release of *code*. Model weights carry the terms of their backbone licences.
|