SatQuery / docs /RESEARCH_NOTES.md
thundercode's picture
release: add docs/RESEARCH_NOTES.md
848485b verified
|
Raw History Blame
9.2 kB
# Research Notes
Engineering findings, negative results and design decisions that would otherwise be lost. Each was
learned by **probe or execution**, not by assumption, and each is recorded so it is not rediscovered.
**Status tags:** `MEASURED` Β· `RESOLVED` Β· `REJECTED` Β· `OPEN` Β· `ATTEMPTED`.
---
## 1. Findings that changed the code
### F4-1 β€” the MiniLM tokenizer ceiling (`MEASURED`)
The MiniLM tokenizer's own ceiling is **256** (verified by probe). The project truncates to **128** β€”
a deliberate truncation *well inside* the ceiling, not the model limit. Satellite queries are short;
halving the sequence halves attention cost for no measurable accuracy loss. The encoder **asserts**
`max_length ≀ 256`, because truncating above the ceiling is a silent no-op.
### F4-2 β€” the router needs no GPU (`MEASURED`)
The encoder is frozen, so embeddings are **cached** and the 50,822-parameter adapter trains on cached
vectors. **Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.** No GPU required.
### F4-3 β€” splits must be by group (`MEASURED`)
Splits are by **group** (template / hard-negative family), never by example. Hard-negative families
are placed in the **test** split so their accuracy measures generalisation rather than memorisation.
Splitting by example would leak template variants across the boundary.
### F5-1 β€” `AutoModelForVision2Seq` does not exist (`MEASURED`)
In transformers 5.17.0, `AutoModelForVision2Seq` **does not exist** (not merely deprecated);
`AutoModelForImageTextToText` is present. The loader is resolved by **feature detection**, never
hardcoded.
### F5-2 β€” the processor cost overrun is ~17Γ—, not 4Γ— (`MEASURED`)
The processor's default `longest_edge` is **2048**, which upscales 512-px tiles **4Γ—** and then splits
them (`do_image_splitting=True`) into **17 sub-images**:
| Setting | `pixel_values` | prompt tokens |
|---|---|---|
| default | `(1, 17, 3, 512, 512)` | 1142 |
| pinned (`processor_longest_edge: 512`) | `(1, 1, 3, 512, 512)` | β€” |
The plan estimated a 4Γ— cost overrun; the real figure is **~17Γ—**. The value **must** be set
explicitly on the processor at construction time.
### F5-3 β€” prompts must go through the chat template (`MEASURED`)
SmolVLM requires one `<image>` token per image in the prompt. Hand-written prompt strings raise
`ValueError`. Prompts are always built through `processor.apply_chat_template()`.
### P7-1 β€” the RemoteCLIP projected dimension is 512 (`MEASURED`)
The RemoteCLIP ViT-B/32 transformer width is **768**, but `visual.proj` maps to a **projected** dim of
**512**. The grounding head's per-cell feature is `4 Γ— 512 = 2048`, declared in config so
`core/config.py` can validate the head **without importing torch**.
### C-1 β€” the availability mask is consumed by the head, not by CROMA (`MEASURED`)
CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied
by the **fusion head** (`input_dim = 3 Γ— 768 + 12 + 2 = 2318`), not by the encoder.
### C-6 β€” T4 is SM 7.5, so training uses fp16, not bf16 (`MEASURED`)
The training precision is **fp16** because the target GPU (T4) is compute capability 7.5. bf16 is
not available there.
### C-7 β€” `image_resolution % 8 == 0` (`MEASURED`)
CROMA requires `image_resolution % 8 == 0`. The native value **120** yields 225 patches.
### C-8 β€” ZeroGPU does not support `torch.compile` (`MEASURED`)
`torch.compile` must never be enabled on the (historical) ZeroGPU target.
### C-9 β€” STANet change hyperparameters are upstream-verified (`MEASURED`)
Change detection uses STANet-style architecture with upstream-verified hyperparameters (PAM
self-attention, ResNet-18 encoder).
---
## 2. The grounding resolution decision β€” a pre-registered rejection (`REJECTED`)
**Question:** should grounding decode at 448 or 224?
**Answer: 224. 448 was rejected** β€” and the rejection is notable because it was *pre-registered* and
then *confirmed* by a paired test:
| Comparison (over identical samples, n = 16,159) | 448 vs 224 |
|---|---|
| mean best IoU | **βˆ’0.0147** |
| recall@0.5 | βˆ’0.0022 |
| recall@0.10 | βˆ’0.0699 |
| recall@0.25 | βˆ’0.0243 |
| latency | **1.59Γ—** |
| paired 95 % CI | [βˆ’0.0160, βˆ’0.0134] |
| paired t | **βˆ’22.63** |
| 448 better on | 8.5 % of records |
| 448 worse on | **20.9 %** of records |
448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224.
This is a model of how a resolution decision should be made: declared in advance, then tested.
---
## 3. The router defect β€” a real bug, found and fixed
### 3.1 Symptom
The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered
**"River"** β€” instead of routing to `grounding`. A second query, *"Where is the new airport?"*,
behaved the same way.
### 3.2 Root cause
Two functions with different information:
- **`interpret()`** β€” produces the console's *reading*; **asset-count-blind** (text only).
- **`chooseTask()`** β€” performs *dispatch*; **asset-count-aware**.
The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell
through to the generic VQA specialist. See [`ARCHITECTURE.md`](ARCHITECTURE.md) Β§4.
### 3.3 Fix and verification (`RESOLVED`)
The fix was deployed to `SatQuery-Frontend` and validated by **three independent live passes**:
| Pass | Deployed HEAD | Result |
|---|---|---|
| 1 | `ff46eba42b18` + `d413d3672311` | 8/8 |
| 2 | `2d7ae53b482d` | 8/8 |
| 3 | `2d7ae53b482d` | 8/8 |
Both defect queries now dispatch to `grounding`:
| Query | Run id | Dispatched |
|---|---|---|
| Where are the built-up areas in this image? | `run_467ffa406f22` | `grounding` |
| Where is the new airport? | `run_46980ba55c62` | `grounding` |
24 live runs, 24 correct dispatches, **0 mock nodes**. Screenshots are in
[`../screenshots/`](../screenshots/).
---
## 4. The harness false-positive β€” caught before it could lie
An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome
**silently drops when the window lacks OS focus**. The harness therefore dispatched the page's
*default* query and still recorded a "result" β€” a **false pass**.
**Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`), and
uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`).
**Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected β€” its
answers were query-specific and the query text was embedded in the answers. The failure mode is
recorded because it is exactly the silent false-positive an evaluation harness must never have.
---
## 5. The `transport_mode: auto` fallthrough (`OPEN`)
`SATQUERY_TRANSPORT=auto` tries the tunnel, then falls through to the forward path on timeout. The
forward path to a **private** repo returns `302` quickly, but the wake step still consumes
`SATQUERY_WAKE_TIMEOUT_S` (120 s) first β€” so a worst-case failed request takes β‰ˆ 249 s
(150 + 120). This is the root shape of the observed transient tunnel gap. Recorded as **OPEN**; a
deployed fix for the `codespace_name` newline on the wake path was authored separately.
---
## 6. The `interpret()` / `chooseTask()` asymmetry β€” intentional (`RESOLVED`)
For *"What changed between the earlier and later image?"* with **one** asset attached, the console
**reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the
reading describes the question's intent, the dispatch respects what can actually be computed with the
assets present. It is documented so it is not mistaken for a defect.
---
## 7. Environment findings (would otherwise cost hours)
| Finding | Detail |
|---|---|
| **Dead proxy in the sandbox** | outbound calls need `--noproxy '*'` (curl) or `ProxyHandler({})` (Python). |
| **pytest is only in the repo venv** | `.venv/Scripts/python.exe`; a bare `pytest` misses it. |
| **Full-suite pytest trips a bulk-delete guard** | sandbox-specific; affects `test_safe_delete_shim`. |
| **Cloudflare `_headers` concatenate** | two matching rules are merged, not overridden; Chromium takes the **first** `max-age`. |
| **Cloudflare 308-redirects `X.html` β†’ `/X`** | reference the extensionless path. |
| **A forwarded Codespace port returns `302`** for a private repo | this is *why* the tunnel exists. |
| **Chrome drops synthetic CDP key events without OS focus** | the harness false-positive (Β§4). |
| **`browser-use` block-buffers stdout** even when redirected | needs explicit line buffering to stream. |
---
## 8. The BigEarthNet format contradiction (`ATTEMPTED`, reported not resolved)
The BigEarthNet data format **contradicts the original plan**. This was **reported rather than
silently patched**, because quietly changing the preprocessing would move the frozen config hash. The
BigEarthNet documentation does not specify a percentile stretch; this project nevertheless applies
percentile normalisation (2/98) to match the CROMA contract. That is a **deliberate, documented
choice**, not an upstream fact. See [`DATASETS.md`](DATASETS.md) Β§5.3.