File size: 10,354 Bytes
df0d288
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
# Models

**Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` ·
`DEFERRED` · `REJECTED`.

SatQuery AI trains **six** artifacts. Four are task heads, one is a router adapter, one is a LoRA
adapter. Every backbone is **frozen** and publicly pinned by revision in `configs/base.yaml` — the
project trains small modules on top of frozen encoders, not end-to-end networks.

> **This file is the human-readable companion to the machine-generated
> [`../models/manifest.json`](../models/manifest.json) and
> [`../models/checksums.sha256`](../models/checksums.sha256) (Phase 3).** Where the two disagree,
> the generated manifest wins — it is computed from the files, this document is written by hand.

---

## 1. The six trained artifacts

| # | Task | Artifact path | Bytes | Kind | Backbone (frozen) |
|---|---|---|---|---|---|
| 1 | `change` | `artifacts/change/levir_change_v001/head.pt` | 63,231,009 | trained head | STANet-style, ResNet-18 + PAM |
| 2 | `change_vqa` | `artifacts/change_vqa/run/head.pt` | 5,822,809 | trained head | over the change encoder's features |
| 3 | `optical_sar` | `artifacts/optical_sar/fusion_head_production_v001/head.pt` | 14,427,457 | trained head (production) | CROMA-base (frozen), 19-class head |
| 4 | `grounding` | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | 12,639,041 | trained head | RemoteCLIP ViT-B/32 (frozen) |
| 5 | `router` | `artifacts/router/router_adapter_v001/adapter.pt` | 211,961 | trained adapter | `all-MiniLM-L6-v2` (frozen) |
| 6 | `vlm` | `.scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors` | 34,798,048 | **LoRA adapter** (PEFT) | `HuggingFaceTB/SmolVLM-500M-Instruct` (frozen) |

Training checkpoints also exist (`checkpoint_last.pt`, `checkpoint-1500`, `checkpoint-2000`) and are
**not** the released artifacts — they are archived as provenance.

## 2. Backbones — pinned, frozen, never retrained

| Role | Repository | Revision | Notes |
|---|---|---|---|
| Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | 90.9 MB, 22,713,216 params, 384-dim embeddings, tokenizer ceiling 256; truncation set to 128 |
| VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | ~1015 MB safetensors |
| Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` | 605.2 MB; transformer width 768, **projected** dim 512 |
| Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` | 777.6 MB; `encoder_dim` 768; `image_resolution` 120 |

These are resolved from the Hugging Face Hub on first use. **No backbone weights are redistributed**
by this project's release — see §6.

## 3. Per-artifact detail

### 3.1 Change (`change`) — `IMPLEMENTED`, `VERIFIED`

STANet-style Siamese change detector. Encoder ResNet-18, self-attention mode **PAM**, tile size 256,
threshold 0.50, minimum component 32 px. Loss is BCE (0.5) + Dice (0.5). Trained on LEVIR-CD-256
(train 7120 / val 1024 / test 2048).

**Measured** on the LEVIR-CD-256 test split (n = 2048, threshold 0.50):

| Metric | Value |
|---|---|
| pooled IoU | **0.8122** |
| macro IoU | **0.8457** |
| pooled F1 | **0.8964** |

This is the only task whose headline number carries the `VERIFIED` tag, because it is the only one
measured against a single, immutable public test split with a frozen threshold.

### 3.2 Change-VQA (`change_vqa`) — `IMPLEMENTED`, `MEASURED`, ruling **OPEN**

A head that answers natural-language change questions over a temporal pair. It is the dispatch
target for change-style questions when only one asset is attached (see
[`ARCHITECTURE.md`](ARCHITECTURE.md) §4).

**Measured on two test sets — both are reported; quoting only the better one would be selective:**

| Test set | accuracy | macro F1 |
|---|---|---|
| `test` (n = 39,686) | **0.697626** | **0.378373** |
| `test2` | **0.651469** | **0.372309** |

The wide gap between accuracy and macro-F1 means the head is carried by common classes and performs
poorly on rare ones. The ruling is **OPEN** — no promotion/acceptance decision has been recorded.

### 3.3 Optical-SAR fusion (`optical_sar`) — `IMPLEMENTED`, `MEASURED`, ruling **OPEN**

Uses frozen CROMA-base to produce optical (768), SAR (768) and joint (768) embeddings, concatenates
them with the 12 optical and 2 SAR channel descriptors, and feeds a 19-class head:

```
input_dim = 3 * 768 + 12 + 2 = 2318   →   hidden 512   →   num_classes 19   (BigEarthNet CLC)
```

The **availability mask is consumed by the fusion head, not by CROMA** — CROMA always sees the
canonical channel counts (12 optical, 2 SAR).

**Measured** (production head `fusion_head_production_v001`, pre-registered 115-class protocol,
held-out test n = 4000):

| Metric | Value |
|---|---|
| accuracy | **0.931** |
| macro F1 | **0.434161** |

**Both numbers must travel together.** The high accuracy with a low macro-F1 reflects class
imbalance across 19 classes. The pre-registered metric JSON records `macro_f1_denominator` and
`classes_present`/`classes_absent` so the denominator is auditable. The ruling is **OPEN**.

**Limitation:** the live service returns a bare class index (`class_18`), not a human-readable label.

### 3.4 Grounding (`grounding`) — `IMPLEMENTED`, `MEASURED` — **two protocols**

A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Per-cell feature is
`concat([patch, text, patch·text, global_pool])` = `4 × 512 = 2048` (enforced at config load). Cells
are assigned by ground-truth box centre (`cell_relative` decode). Objectness BCE is weighted **20×**
because only ~1 of 49 cells is positive; unweighted, the optimum collapses to "no object".

Image resolution is frozen at **224** — 448 was evaluated and **rejected** (see §4).

**Measured on VRSBench (n = 16,159), reported under two protocols:**

| Protocol | mean best IoU | recall@0.5 |
|---|---|---|
| canonical (head threshold decode) | **0.2838** | **0.2198** |
| matched6 | **0.2566** | **0.1938** |

Two further decode variants exist and are reported for completeness — a reviewer must be able to see
the whole grid, not one cell of it:

| Variant | mean best IoU |
|---|---|
| head argmax decode (canonical) | **0.1215** |
| zero-shot matched (no trained head) | **0.0972** |

The trained head beats the zero-shot baseline by a wide margin, which is the point of the head; the
absolute IoU is low, which is an honest limitation.

### 3.5 Router (`router`) — `IMPLEMENTED`, `MEASURED`, **TEST NOT RUN**

A **50,822-parameter adapter** over the frozen MiniLM encoder. Because the encoder is frozen,
embeddings are cached and the adapter trains on cached vectors — no GPU required (measured: 20
epochs / 4,096 vectors in 0.28 s on CPU). Splits are by **group** (template / hard-negative family),
never by example; hard-negative families are placed in the test split so their accuracy measures
generalisation, not memorisation.

**Measured:** overall **ungated** task accuracy **0.965116** on the validation split, **n = 86**,
corpus-limited. This number is (a) validation-only, (b) ungated, and (c) small. The router **test**
split was **NOT RUN**. Do not read 0.965116 as a test result.

### 3.6 VLM LoRA (`vlm`) — `IMPLEMENTED`, `MEASURED`, **ACCEPTANCE-REJECTED**

A PEFT LoRA adapter on frozen SmolVLM-500M-Instruct: `peft_type=LORA`, `r=16`, `alpha=32`,
`dropout=0.05`, targeting `model.text_model.*.{q,k,v,o,gate,up,down}_proj`. PEFT 0.19.1.

**Measured** on a frozen 1,000-question subset: exact_match **0.963**, F1 **0.96432** (+49.5 pp over
the unadapted baseline).

**Yet the artifact's status is `CLOSED` with headline `ACCEPTANCE-REJECTED`.** This is not a
contradiction — it is the project's central truthfulness distinction:

- **`USABLE_VERIFIED`** — the adapter demonstrably works (the metrics are real and reproducible).
- **`ACCEPTANCE-REJECTED`** — the adapter is *not accepted* for production promotion, on grounds
  recorded in `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`).

The deployed caption/VQA path therefore uses the **unadapted** SmolVLM. Deployment success and model
acceptance are different claims, and this document keeps them apart.

## 4. Rejected and deferred model decisions

| Decision | Outcome | Evidence |
|---|---|---|
| Grounding image resolution 448 vs 224 | **224 chosen; 448 REJECTED** | 448 lost on every axis: mean best IoU −0.0147, recall@0.5 −0.0022, all recall thresholds lower, at 1.59× latency. Paired test: mean diff −0.0147, 95 % CI [−0.0160, −0.0134], t = −22.63; 448 better on 8.5 % of records, worse on 20.9 %. Pre-registered rule and the paired test **agree**. |
| VLM adapter promotion | **REJECTED** | metrics usable, acceptance rejected (§3.6) |
| Calibration | **kept but ineffective** | see §5 |
| optical-SAR / change-VQA rulings | **OPEN** | no decision recorded |

## 5. Calibration — `MEASURED`, **not an improvement**

Temperature scaling is enabled (`confidence.temperature_scaling: true`) with
`calibration_v001.json`. Fitted temperature **T = 0.9772732** on the validation split (n = 16,441).

| Metric | Before | After |
|---|---|---|
| ECE | 0.013755 | **0.014929** |
| NLL | 0.689741 | 0.689631 |

**ECE got worse** (`ece_improvement = −0.001174`). The scaling is retained because it is part of the
frozen configuration, **not** because it helped. The reliability diagram on the Benchmark page is
explicitly labelled **pre-scaling** so a reader cannot mistake it for the calibrated result. This is
recorded as a negative result, not smoothed over.

## 6. Distribution and licensing

- **Backbones are not redistributed.** They are fetched from the Hugging Face Hub at run time, pinned
  by revision. Their licences are their own (see each model's HF page).
- **The six trained artifacts are published by this project** on the Hugging Face Hub under
  `thundercode/SatQuery`, labelled by kind, each with its backbone dependency documented and each
  accompanied by a checksum. See [`../HF_RELEASE_VERIFICATION.md`](../HF_RELEASE_VERIFICATION.md).
- **No licence file exists in the source repository.** This is an **OPEN** item flagged in
  [`LIMITATIONS.md`](LIMITATIONS.md); the repository README instructs the owner to select one before
  any public release of *code*. Model weights carry the terms of their backbone licences.