# Research Notes Engineering findings, negative results, and design decisions that would otherwise be lost. Every one was learned by **probe or execution**, not assumption, and every one is recorded so it is not rediscovered. **Status tags:** `MEASURED` · `RESOLVED` · `REJECTED` · `OPEN` · `ATTEMPTED`. --- ## 1. Findings that changed the code ### F4-1 — the MiniLM tokenizer ceiling (`MEASURED`) The MiniLM tokenizer's own ceiling is **256** (verified by probe). This project truncates to **128** — a deliberate truncation *well inside* the ceiling, not the model limit. Satellite queries are short; halving the sequence halves attention cost for no measurable accuracy loss. The encoder **asserts** `max_length ≤ 256`, because truncating above the ceiling is a **silent no-op**. **Consequence:** `router.max_length: 128` with an assertion, rather than a comment. ### F4-2 — the router needs no GPU (`MEASURED`) The encoder is frozen, so embeddings are **cached** and the 50,822-parameter adapter trains on cached vectors. **Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.** **Consequence:** the router can be retrained during development without a GPU. ### F4-3 — splits must be by group (`MEASURED`) Splits are by **group** (template / hard-negative family), never by example. Hard-negative families are placed in the **test** split so their accuracy measures generalisation rather than memorisation. Splitting by example would leak template variants across the boundary and inflate the score. ### F5-1 — `AutoModelForVision2Seq` does not exist (`MEASURED`) In transformers 5.17.0, `AutoModelForVision2Seq` **does not exist** (it is not merely deprecated); `AutoModelForImageTextToText` is present. **Consequence:** the loader is resolved by **feature detection**, never hardcoded to one class name. ### F5-2 — the processor cost overrun is ~17×, not 4× (`MEASURED`) The processor's default `longest_edge` is **2048**, which upscales a 512-px tile **4×** and then splits it (`do_image_splitting=True`) into sub-images: | Setting | `pixel_values` | prompt tokens | |---|---|---| | default | `(1, 17, 3, 512, 512)` | 1142 | | pinned (`processor_longest_edge: 512`) | `(1, 1, 3, 512, 512)` | — | The plan estimated a 4× cost overrun; the real figure is **~17×**. The value **must** be set explicitly on the processor at construction time, and `core/config.py` now **enforces** `processor_longest_edge <= image.tile_size` so the pin is a control rather than a comment. ### F5-3 — prompts must go through the chat template (`MEASURED`) SmolVLM requires one `` token per image in the prompt. Hand-written prompt strings raise `ValueError`. **Consequence:** prompts are always built through `processor.apply_chat_template()`, enforced by the config key `vlm.prompt_must_use_chat_template: true`. ### P7-1 — the RemoteCLIP projected dimension is 512 (`MEASURED`) The RemoteCLIP ViT-B/32 transformer width is **768**, but `visual.proj` maps to a **projected** dim of **512**. The grounding head's per-cell feature is therefore `4 × 512 = 2048`. **Consequence:** `grounding.encoder_projected_dim: 512` is declared in config so `core/config.py` can validate the head **without importing torch**, and `specialists/grounding/remoteclip.py` asserts the same value against the real model at load time. ### C-1 — the availability mask is consumed by the head, not by CROMA (`MEASURED`) CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied by the **fusion head** (`input_dim = 3 × 768 + 12 + 2 = 2318`), not by the encoder. ### C-6 — T4 is SM 7.5, so training uses fp16, not bf16 (`MEASURED`) `training.precision: fp16` because the target GPU (T4) is compute capability 7.5; bf16 is unavailable there. The loader validates the value is one of `fp16|bf16|fp32`. ### C-7 — `image_resolution % 8 == 0` (`MEASURED`) CROMA requires `image_resolution % 8 == 0`. The native value **120** yields 225 patches. Enforced at config load. ### C-8 — ZeroGPU does not support `torch.compile` (`MEASURED`) `torch.compile` must never be enabled on the (historical) ZeroGPU target. Enforced: the loader **fails startup** if `deployment.torch_compile` is true. ### C-9 — STANet hyperparameters are upstream-verified (`MEASURED`) Change detection uses a STANet-style architecture with upstream-verified hyperparameters: ResNet-18 encoder, **PAM** self-attention mode, tile 256, threshold 0.50, BCE 0.5 + Dice 0.5. --- ## 2. The grounding resolution decision — a pre-registered rejection (`REJECTED`) **Question:** should grounding decode at 448 or 224? **Answer: 224. 448 was rejected** — notable because the rejection was *pre-registered* and then *confirmed* by a paired test over identical samples (n = 16,159): | Comparison (448 vs 224) | Value | |---|---| | mean best IoU | **−0.0147** | | recall@0.5 | −0.0022 | | recall@0.10 | −0.0699 | | recall@0.25 | −0.0243 | | latency | **1.59×** | | paired mean difference | −0.0147 | | paired 95 % CI | **[−0.0160, −0.0134]** | | paired t | **−22.63** | | 448 better on | 8.5 % of records | | 448 worse on | **20.9 %** of records | 448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224. This is the model for how a resolution decision should be made: declared in advance, then tested. --- ## 3. The router defect — a real bug, found and fixed (`RESOLVED`) ### 3.1 Symptom The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered **"River"** — instead of routing to `grounding`. A second query, *"Where is the new airport?"*, behaved the same way. ### 3.2 Root cause Two functions with different information: - **`interpret()`** — produces the console's *reading*; **asset-count-blind** (text only). - **`chooseTask()`** — performs *dispatch*; **asset-count-aware**. The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell through to the generic VQA specialist. ### 3.3 Fix and verification The fix was deployed to `SatQuery-Frontend` and validated by **three independent live passes**: | Pass | Deployed HEAD | Result | |---|---|---| | 1 | `ff46eba42b18` + `d413d3672311` | 8/8 | | 2 | `2d7ae53b482d` | 8/8 | | 3 | `2d7ae53b482d` | 8/8 | Both defect queries now dispatch to `grounding`: | Query | Run id | Dispatched | |---|---|---| | Where are the built-up areas in this image? | `run_467ffa406f22` | `grounding` | | Where is the new airport? | `run_46980ba55c62` | `grounding` | **24 live runs, 24 correct dispatches, 0 mock nodes.** Screenshots are in [`../screenshots/`](../screenshots/). --- ## 4. The harness false-positive — caught before it could lie (`RESOLVED`) An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome **silently drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query and still recorded a "result" — a **false pass**. **Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`) and uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`). **Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected — its answers were query-specific and the query text was embedded in the answers. The failure mode is recorded because it is exactly the silent false-positive an evaluation harness must never have. --- ## 5. The `transport_mode: auto` fallthrough (`OPEN`) `SATQUERY_TRANSPORT=auto` tries the tunnel, then falls through to the forward path on timeout. The forward path to a **private** repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S` (120 s) first — so a worst-case failed request takes ≈ **249 s** (150 + 120). This is the root shape of the observed transient tunnel gap. A patch adding `forward_unavailable` (503) and `upstream_timeout` (504) codes plus the `codespace_name` `.strip()` fix was authored and verified (`py_compile` clean). **Status: OPEN — the patch is prepared but NOT deployed.** --- ## 6. The `interpret()` / `chooseTask()` asymmetry — intentional (`RESOLVED`) For *"What changed between the earlier and later image?"* with **one** asset attached, the console **reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the reading describes the question's intent; the dispatch respects what can actually be computed with the assets present. It is documented so it is not mistaken for a defect. --- ## 7. Environment findings (would otherwise cost hours) | Finding | Detail | |---|---| | **Dead proxy in the authoring sandbox** | outbound calls need `--noproxy '*'` (curl) or `ProxyHandler({})` (Python). | | **The sandbox proxy is slow for uploads** | `huggingface_hub` uploads stalled at ~51 kB/s through `http_proxy=127.0.0.1:58294`; the fix is to unset `http_proxy`/`https_proxy` and set `no_proxy='*'`. | | **`hf_hub_download` returned an EMPTY file** | sha256 `e3b0c442…` (the empty-content hash), which produced a **false FAIL** for all six artifacts. The honest check is a direct HTTPS download with `ProxyHandler({})`. | | **pytest is only in the repo venv** | `.venv/Scripts/python.exe`; a bare `pytest` misses it. | | **The full test suite trips a bulk-delete guard** | sandbox-specific; affects `test_safe_delete_shim`. | | **Cloudflare `_headers` concatenate** | two matching rules are merged, not overridden; Chromium takes the **first** `max-age`. | | **Cloudflare 308-redirects `X.html` → `/X`** | reference the extensionless path. | | **A forwarded Codespace port returns `302`** | for a private repo — this is *why* the tunnel exists. | | **Chrome drops synthetic CDP key events without OS focus** | the harness false-positive (§4). | | **`browser-use` block-buffers stdout** | even when redirected; needs explicit line buffering to stream. | --- ## 8. The BigEarthNet format contradiction (`ATTEMPTED`, reported not resolved) The BigEarthNet data format **contradicts the original plan**. This was **reported rather than silently patched**, because quietly changing the preprocessing would move the frozen config hash. Two facts matter: 1. The BigEarthNet documentation — its uses, mentions, or endorsements — does **not** specify a percentile stretch. This project nevertheless applies percentile normalisation (2/98) for optical inputs to match the CROMA contract. That is a **deliberate, documented choice**, not an upstream fact. 2. The local subset is **100 % single-label**, against the official 1–11 multi-label scheme, so metrics computed on it are **not comparable** to published multi-label numbers. --- ## 9. Where the evidence lives | Topic | Evidence | |---|---| | Findings F4-1…F5-3, C-1…C-9, P7-1 | `docs/PHASE*.md`, `docs/ARCHITECTURE_*.md`, `docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md` | | The 448-vs-224 paired test | `docs/PHASE7_RESOLUTION_DECISION.md` | | The VLM rejection | `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md`, `artifacts/vlm/phase6_closure.json` | | Router defect + 3 live passes | `.workbuddy-ai/scratch/live_validation/` (`run_output.txt`, `run_final2.txt`, `run_final3.txt`) | | The undeployed B-07 patch | session scratch `fix-b07-forward-unavailable.patch` | | Live validation harness | `.workbuddy-ai/scratch/run_all_postfix2.harness`, `recompute_verdicts.py` |