File size: 11,513 Bytes
848485b
 
267161f
 
 
848485b
 
 
 
 
 
 
 
 
267161f
848485b
 
267161f
 
 
848485b
 
 
 
267161f
 
 
848485b
 
 
 
 
267161f
848485b
 
 
267161f
 
 
 
848485b
 
 
267161f
 
848485b
 
 
 
 
 
 
267161f
 
848485b
 
 
 
267161f
 
 
 
848485b
 
 
 
267161f
 
 
 
 
848485b
 
 
 
 
 
 
 
267161f
 
848485b
 
 
267161f
 
848485b
 
 
267161f
 
848485b
267161f
848485b
267161f
 
848485b
 
 
 
 
 
 
267161f
 
848485b
267161f
848485b
 
 
 
 
 
267161f
 
848485b
 
 
 
 
267161f
848485b
 
 
267161f
848485b
 
 
 
 
 
 
 
 
 
 
 
 
 
267161f
 
848485b
267161f
848485b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
267161f
848485b
 
 
 
267161f
848485b
 
 
 
 
267161f
848485b
 
 
 
 
 
 
 
 
 
 
 
267161f
 
 
 
 
 
848485b
 
 
 
 
 
 
267161f
848485b
 
 
 
 
 
 
 
267161f
 
 
848485b
267161f
848485b
 
267161f
848485b
267161f
848485b
 
 
 
 
 
267161f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
# Research Notes

Engineering findings, negative results, and design decisions that would otherwise be lost. Every one
was learned by **probe or execution**, not assumption, and every one is recorded so it is not
rediscovered.

**Status tags:** `MEASURED` · `RESOLVED` · `REJECTED` · `OPEN` · `ATTEMPTED`.

---

## 1. Findings that changed the code

### F4-1 — the MiniLM tokenizer ceiling (`MEASURED`)

The MiniLM tokenizer's own ceiling is **256** (verified by probe). This project truncates to **128** —
a deliberate truncation *well inside* the ceiling, not the model limit. Satellite queries are short;
halving the sequence halves attention cost for no measurable accuracy loss. The encoder **asserts**
`max_length ≤ 256`, because truncating above the ceiling is a **silent no-op**.

**Consequence:** `router.max_length: 128` with an assertion, rather than a comment.

### F4-2 — the router needs no GPU (`MEASURED`)

The encoder is frozen, so embeddings are **cached** and the 50,822-parameter adapter trains on cached
vectors. **Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.**

**Consequence:** the router can be retrained during development without a GPU.

### F4-3 — splits must be by group (`MEASURED`)

Splits are by **group** (template / hard-negative family), never by example. Hard-negative families
are placed in the **test** split so their accuracy measures generalisation rather than memorisation.
Splitting by example would leak template variants across the boundary and inflate the score.

### F5-1 — `AutoModelForVision2Seq` does not exist (`MEASURED`)

In transformers 5.17.0, `AutoModelForVision2Seq` **does not exist** (it is not merely deprecated);
`AutoModelForImageTextToText` is present.

**Consequence:** the loader is resolved by **feature detection**, never hardcoded to one class name.

### F5-2 — the processor cost overrun is ~17×, not 4× (`MEASURED`)

The processor's default `longest_edge` is **2048**, which upscales a 512-px tile **4×** and then
splits it (`do_image_splitting=True`) into sub-images:

| Setting | `pixel_values` | prompt tokens |
|---|---|---|
| default | `(1, 17, 3, 512, 512)` | 1142 |
| pinned (`processor_longest_edge: 512`) | `(1, 1, 3, 512, 512)` | — |

The plan estimated a 4× cost overrun; the real figure is **~17×**. The value **must** be set
explicitly on the processor at construction time, and `core/config.py` now **enforces**
`processor_longest_edge <= image.tile_size` so the pin is a control rather than a comment.

### F5-3 — prompts must go through the chat template (`MEASURED`)

SmolVLM requires one `<image>` token per image in the prompt. Hand-written prompt strings raise
`ValueError`.

**Consequence:** prompts are always built through `processor.apply_chat_template()`, enforced by the
config key `vlm.prompt_must_use_chat_template: true`.

### P7-1 — the RemoteCLIP projected dimension is 512 (`MEASURED`)

The RemoteCLIP ViT-B/32 transformer width is **768**, but `visual.proj` maps to a **projected** dim of
**512**. The grounding head's per-cell feature is therefore `4 × 512 = 2048`.

**Consequence:** `grounding.encoder_projected_dim: 512` is declared in config so `core/config.py` can
validate the head **without importing torch**, and `specialists/grounding/remoteclip.py` asserts the
same value against the real model at load time.

### C-1 — the availability mask is consumed by the head, not by CROMA (`MEASURED`)

CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied
by the **fusion head** (`input_dim = 3 × 768 + 12 + 2 = 2318`), not by the encoder.

### C-6 — T4 is SM 7.5, so training uses fp16, not bf16 (`MEASURED`)

`training.precision: fp16` because the target GPU (T4) is compute capability 7.5; bf16 is unavailable
there. The loader validates the value is one of `fp16|bf16|fp32`.

### C-7 — `image_resolution % 8 == 0` (`MEASURED`)

CROMA requires `image_resolution % 8 == 0`. The native value **120** yields 225 patches. Enforced at
config load.

### C-8 — ZeroGPU does not support `torch.compile` (`MEASURED`)

`torch.compile` must never be enabled on the (historical) ZeroGPU target. Enforced: the loader
**fails startup** if `deployment.torch_compile` is true.

### C-9 — STANet hyperparameters are upstream-verified (`MEASURED`)

Change detection uses a STANet-style architecture with upstream-verified hyperparameters: ResNet-18
encoder, **PAM** self-attention mode, tile 256, threshold 0.50, BCE 0.5 + Dice 0.5.

---

## 2. The grounding resolution decision — a pre-registered rejection (`REJECTED`)

**Question:** should grounding decode at 448 or 224?

**Answer: 224. 448 was rejected** — notable because the rejection was *pre-registered* and then
*confirmed* by a paired test over identical samples (n = 16,159):

| Comparison (448 vs 224) | Value |
|---|---|
| mean best IoU | **−0.0147** |
| recall@0.5 | −0.0022 |
| recall@0.10 | −0.0699 |
| recall@0.25 | −0.0243 |
| latency | **1.59×** |
| paired mean difference | −0.0147 |
| paired 95 % CI | **[−0.0160, −0.0134]** |
| paired t | **−22.63** |
| 448 better on | 8.5 % of records |
| 448 worse on | **20.9 %** of records |

448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224.
This is the model for how a resolution decision should be made: declared in advance, then tested.

---

## 3. The router defect — a real bug, found and fixed (`RESOLVED`)

### 3.1 Symptom

The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered
**"River"** — instead of routing to `grounding`. A second query, *"Where is the new airport?"*,
behaved the same way.

### 3.2 Root cause

Two functions with different information:

- **`interpret()`** — produces the console's *reading*; **asset-count-blind** (text only).
- **`chooseTask()`** — performs *dispatch*; **asset-count-aware**.

The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell through
to the generic VQA specialist.

### 3.3 Fix and verification

The fix was deployed to `SatQuery-Frontend` and validated by **three independent live passes**:

| Pass | Deployed HEAD | Result |
|---|---|---|
| 1 | `ff46eba42b18` + `d413d3672311` | 8/8 |
| 2 | `2d7ae53b482d` | 8/8 |
| 3 | `2d7ae53b482d` | 8/8 |

Both defect queries now dispatch to `grounding`:

| Query | Run id | Dispatched |
|---|---|---|
| Where are the built-up areas in this image? | `run_467ffa406f22` | `grounding` |
| Where is the new airport? | `run_46980ba55c62` | `grounding` |

**24 live runs, 24 correct dispatches, 0 mock nodes.** Screenshots are in
[`../screenshots/`](../screenshots/).

---

## 4. The harness false-positive — caught before it could lie (`RESOLVED`)

An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome
**silently drops when the window lacks OS focus**. The harness therefore dispatched the page's
*default* query and still recorded a "result" — a **false pass**.

**Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`) and
uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`).

**Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected — its
answers were query-specific and the query text was embedded in the answers. The failure mode is
recorded because it is exactly the silent false-positive an evaluation harness must never have.

---

## 5. The `transport_mode: auto` fallthrough (`OPEN`)

`SATQUERY_TRANSPORT=auto` tries the tunnel, then falls through to the forward path on timeout. The
forward path to a **private** repo returns `302` quickly, but the wake step still consumes
`SATQUERY_WAKE_TIMEOUT_S` (120 s) first — so a worst-case failed request takes ≈ **249 s**
(150 + 120). This is the root shape of the observed transient tunnel gap.

A patch adding `forward_unavailable` (503) and `upstream_timeout` (504) codes plus the
`codespace_name` `.strip()` fix was authored and verified (`py_compile` clean). **Status: OPEN — the
patch is prepared but NOT deployed.**

---

## 6. The `interpret()` / `chooseTask()` asymmetry — intentional (`RESOLVED`)

For *"What changed between the earlier and later image?"* with **one** asset attached, the console
**reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the
reading describes the question's intent; the dispatch respects what can actually be computed with the
assets present. It is documented so it is not mistaken for a defect.

---

## 7. Environment findings (would otherwise cost hours)

| Finding | Detail |
|---|---|
| **Dead proxy in the authoring sandbox** | outbound calls need `--noproxy '*'` (curl) or `ProxyHandler({})` (Python). |
| **The sandbox proxy is slow for uploads** | `huggingface_hub` uploads stalled at ~51 kB/s through `http_proxy=127.0.0.1:58294`; the fix is to unset `http_proxy`/`https_proxy` and set `no_proxy='*'`. |
| **`hf_hub_download` returned an EMPTY file** | sha256 `e3b0c442…` (the empty-content hash), which produced a **false FAIL** for all six artifacts. The honest check is a direct HTTPS download with `ProxyHandler({})`. |
| **pytest is only in the repo venv** | `.venv/Scripts/python.exe`; a bare `pytest` misses it. |
| **The full test suite trips a bulk-delete guard** | sandbox-specific; affects `test_safe_delete_shim`. |
| **Cloudflare `_headers` concatenate** | two matching rules are merged, not overridden; Chromium takes the **first** `max-age`. |
| **Cloudflare 308-redirects `X.html` → `/X`** | reference the extensionless path. |
| **A forwarded Codespace port returns `302`** | for a private repo — this is *why* the tunnel exists. |
| **Chrome drops synthetic CDP key events without OS focus** | the harness false-positive (§4). |
| **`browser-use` block-buffers stdout** | even when redirected; needs explicit line buffering to stream. |

---

## 8. The BigEarthNet format contradiction (`ATTEMPTED`, reported not resolved)

The BigEarthNet data format **contradicts the original plan**. This was **reported rather than
silently patched**, because quietly changing the preprocessing would move the frozen config hash.

Two facts matter:

1. The BigEarthNet documentation — its uses, mentions, or endorsements — does **not** specify a
   percentile stretch. This project nevertheless applies percentile normalisation (2/98) for optical
   inputs to match the CROMA contract. That is a **deliberate, documented choice**, not an upstream
   fact.
2. The local subset is **100 % single-label**, against the official 1–11 multi-label scheme, so
   metrics computed on it are **not comparable** to published multi-label numbers.

---

## 9. Where the evidence lives

| Topic | Evidence |
|---|---|
| Findings F4-1…F5-3, C-1…C-9, P7-1 | `docs/PHASE*.md`, `docs/ARCHITECTURE_*.md`, `docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md` |
| The 448-vs-224 paired test | `docs/PHASE7_RESOLUTION_DECISION.md` |
| The VLM rejection | `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md`, `artifacts/vlm/phase6_closure.json` |
| Router defect + 3 live passes | `.workbuddy-ai/scratch/live_validation/` (`run_output.txt`, `run_final2.txt`, `run_final3.txt`) |
| The undeployed B-07 patch | session scratch `fix-b07-forward-unavailable.patch` |
| Live validation harness | `.workbuddy-ai/scratch/run_all_postfix2.harness`, `recompute_verdicts.py` |