File size: 8,372 Bytes
4e54b93
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
# Training

**Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `REJECTED`.

The project trains **small modules on frozen backbones**. No backbone is fine-tuned end-to-end. Every
hyperparameter lives in `configs/base.yaml` (no magic numbers in Python), and every trained artifact
records the frozen config hash it was trained against.

---

## 1. Overview

| Artifact | Backbone (frozen) | Where trained | Selection signal |
|---|---|---|---|
| `router` adapter | `all-MiniLM-L6-v2` | local CPU | val (group-split) |
| `grounding` head | RemoteCLIP ViT-B/32 | local | val fraction 0.10 |
| `change` head | STANet ResNet-18 + PAM | local | val (LEVIR split) |
| `optical_sar` fusion head | CROMA-base | local, **seed sweep** | held-out test (pre-registered) |
| `change_vqa` head | over cached change features | **external GPU (Kaggle)** | val answer accuracy |
| `vlm` LoRA adapter | SmolVLM-500M-Instruct | **external GPU** | frozen 1000-question subset |

**CPU-first.** The router and every head except the VLM adapter train on CPU. The VLM LoRA adapter
requires a GPU (T4-class).

## 2. Router adapter

A **50,822-parameter** adapter over the frozen MiniLM encoder.

| Hyperparameter | Value |
|---|---|
| epochs | 60 |
| batch size | 64 |
| learning rate | 0.001 |
| weight decay | 0.01 |
| task loss weight | 1.0 |
| modality loss weight | 0.3 |
| binary loss weight | 0.5 |
| val ratio | 0.15 |
| hard negatives to test | true |

**Finding F4-2 — the encoder is frozen, so embeddings are cached** and the adapter trains on cached
vectors. **Measured: 20 epochs / 4,096 vectors in 0.28 s on CPU.** No GPU is required.

**Finding F4-3 — splits are by GROUP** (template / hard-negative family), never by example.
Hard-negative families are placed in the **test** split so their accuracy measures generalisation
rather than memorisation.

> **Status:** the adapter is trained and shipped. Its measured number (0.965116) is **validation-only,
> ungated, n = 86**; the router **test split was NOT RUN**.

## 3. Grounding head

A trainable head over the frozen RemoteCLIP ViT-B/32 encoder.

| Hyperparameter | Value |
|---|---|
| learning rate | 0.0001 |
| batch size | 16 |
| epochs | 20 |
| weight decay | 0.0001 |
| warmup ratio | 0.05 |
| grad clip | 1.0 |
| val fraction | 0.10 |
| save every steps | 500 |
| box loss weight | 0.5 |
| GIoU loss weight | 0.3 |
| confidence loss weight | 0.2 |
| `positive_confidence_weight` | **20.0** |

**Architecture constraint (enforced, not documented).** Per-cell feature is
`concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048`. `core/config.py` **rejects any
value other than** `4 × grounding.encoder_projected_dim` at load time, and the specialist asserts the
same 512 against the real model — because a mismatch is a *silent* shape error that torch only raises
at the similarity step, after patch features are already cached.

**Why `positive_confidence_weight = 20.0`.** Objectness BCE sees ~1 positive cell out of 49.
Unweighted, the optimum is "no object" everywhere; the weight is what stops that collapse.

**Resolution is frozen at 224.** 448 was evaluated and **REJECTED** (paired test: mean diff −0.0147,
95 % CI [−0.0160, −0.0134], t = −22.63, at 1.59× latency).

## 4. Change head

STANet-style Siamese detector.

| Hyperparameter | Value |
|---|---|
| encoder | ResNet-18 |
| self-attention | **PAM** (BAM alternative not used) |
| tile size | 256 |
| tile overlap | 0 |
| threshold | 0.50 |
| min component pixels | 32 |
| learning rate | 0.001 |
| batch size | 8 |
| BCE weight | 0.5 |
| Dice weight | 0.5 |

Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048). **This is the only task with a
`VERIFIED` headline metric** (pooled IoU 0.8122 on the immutable test split).

## 5. Optical-SAR fusion head

| Hyperparameter | Value |
|---|---|
| input dim | 2318 = 3 × 768 + 12 + 2 |
| hidden dim | 512 |
| dropout | 0.2 |
| num classes | 19 (BigEarthNet CLC) |

**Seed sweep.** Training was run as two arms (**armA**, **armB**) × five seeds (**100–104**), with a
per-arm variance report (`armA_seed_variance_report.json`). The **production head** is a distinct,
frozen artifact (`fusion_head_production_v001/head.pt`) with its own
`production_head_record.json` and a `phase12_rerun_verification.json`.

**Pre-registration.** The headline metric is a **pre-registered** 115-class protocol
(`pre_registered_11.5`) computed over the 19-class label space on the held-out test split. The
metric JSON records `is_deciding_statistic: False`, i.e. it is a reported measurement, not a
decision statistic.

**Feature caches.** Training consumes cached CROMA features
(`fusion_features/`, `fusion_features_armB/`, ~231 MB each). These caches are **reproducible** and are
not released as model weights.

## 6. Change-VQA head — trained externally

The change-VQA head was trained **outside this repository**, on an external GPU (Kaggle), following
`docs/R02_KAGGLE_TRAINING_GUIDE.md`. That guide's status on entry was
`IMPLEMENTATION_READY_FOR_EXTERNAL_TRAINING`, and its explicit contract is:

> **Training produces an artifact, not a verified capability, and the run record says
> `TRAINED_UNVERIFIED`.**

The returned checkpoint was **promoted** through a byte-identity gate
(`artifacts/change_vqa/run/PROMOTION.json`):

| Property | Value |
|---|---|
| sha256 | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` |
| bytes | 5,822,809 |
| architecture | `change_vqa_head_v1` |
| parameters | 1,453,912 |
| non-finite tensors | **0** |
| weights modified during promotion | **false** |
| byte-identical to source | **true** |
| hash agrees across | `model_metadata.json`, `run_record.json`, `hashes.json` |

**Selection:** epoch **8**, chosen on **Val answer accuracy = 0.700018**, stopped by early stopping.
Seed 42. It trains on **cached change + text features** (specs `change_feat_v1`,
`change_cache_spec c801326f85a185f8`, `text_cache_spec d2801ea1a314354a`), not on raw imagery.

> **Note.** The raw CDVQA loader (see [`DATASETS.md`](DATASETS.md)) loads examples but has **no
> training loop**; the shipped head is a cached-feature model. These are different paths and are not
> conflated.

## 7. VLM LoRA adapter — trained externally

A PEFT LoRA adapter on **frozen** `HuggingFaceTB/SmolVLM-500M-Instruct`.

| Hyperparameter | Value |
|---|---|
| PEFT version | **0.19.1** |
| `r` (rank) | 16 |
| `alpha` | 32 |
| `dropout` | 0.05 |
| target modules | `model.text_model.*.{q,k,v,o,gate,up,down}_proj` |
| precision | **fp16** (finding C-6: T4 is SM 7.5 → **fp16, NOT bf16**) |
| batch size | 2 |
| gradient accumulation | 8 |
| learning rate | 0.0002 |
| epochs | 1 |
| gradient checkpointing | true |
| save every steps | 500 |

**Finding F5-2 (cost).** The processor's default `longest_edge` is 2048, which upscales 512-px tiles
4× and then splits them into **17 sub-images** (`pixel_values (1, 17, 3, 512, 512)`, 1142 prompt
tokens). Pinning `processor_longest_edge: 512` yields `pixel_values (1, 1, 3, 512, 512)`. The plan
estimated a 4× cost overrun; the **measured** figure is ~17×.

**Finding F5-3.** SmolVLM requires one `<image>` token per image in the prompt; hand-written prompt
strings raise `ValueError`. Prompts are always built through `processor.apply_chat_template()`.

**Outcome:** metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is
**ACCEPTANCE-REJECTED** for promotion. The deployed caption/VQA path uses the **unadapted** model.
See [`MODELS.md`](MODELS.md) §3.6.

## 8. Reproducibility contract for training

- **Seed 42** everywhere (`project.seed`).
- **Precision `fp16`** (T4 constraint), never bf16.
- Every artifact records the **frozen config hash** `78f1e3700da15aa1`; a config edit moves the hash
  and invalidates the artifact.
- `save_every_steps: 500`; checkpoints are archived as provenance, not released.
- Training guides state their own entry status and **never** claim a trained artifact is a verified
  capability.

## 9. What was NOT trained

| Item | State |
|---|---|
| Backbone fine-tuning (any) | **NOT DONE** — all backbones frozen |
| Router on the test split | **NOT RUN** |
| Any end-to-end / joint training | **NOT RUN** |
| Re-training of the change head at a second resolution | **NOT RUN** |
| Benchmark adapters | **NOT RUN** |