Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 9,196 Bytes
848485b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 | # Research Notes
Engineering findings, negative results and design decisions that would otherwise be lost. Each was
learned by **probe or execution**, not by assumption, and each is recorded so it is not rediscovered.
**Status tags:** `MEASURED` · `RESOLVED` · `REJECTED` · `OPEN` · `ATTEMPTED`.
---
## 1. Findings that changed the code
### F4-1 — the MiniLM tokenizer ceiling (`MEASURED`)
The MiniLM tokenizer's own ceiling is **256** (verified by probe). The project truncates to **128** —
a deliberate truncation *well inside* the ceiling, not the model limit. Satellite queries are short;
halving the sequence halves attention cost for no measurable accuracy loss. The encoder **asserts**
`max_length ≤ 256`, because truncating above the ceiling is a silent no-op.
### F4-2 — the router needs no GPU (`MEASURED`)
The encoder is frozen, so embeddings are **cached** and the 50,822-parameter adapter trains on cached
vectors. **Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.** No GPU required.
### F4-3 — splits must be by group (`MEASURED`)
Splits are by **group** (template / hard-negative family), never by example. Hard-negative families
are placed in the **test** split so their accuracy measures generalisation rather than memorisation.
Splitting by example would leak template variants across the boundary.
### F5-1 — `AutoModelForVision2Seq` does not exist (`MEASURED`)
In transformers 5.17.0, `AutoModelForVision2Seq` **does not exist** (not merely deprecated);
`AutoModelForImageTextToText` is present. The loader is resolved by **feature detection**, never
hardcoded.
### F5-2 — the processor cost overrun is ~17×, not 4× (`MEASURED`)
The processor's default `longest_edge` is **2048**, which upscales 512-px tiles **4×** and then splits
them (`do_image_splitting=True`) into **17 sub-images**:
| Setting | `pixel_values` | prompt tokens |
|---|---|---|
| default | `(1, 17, 3, 512, 512)` | 1142 |
| pinned (`processor_longest_edge: 512`) | `(1, 1, 3, 512, 512)` | — |
The plan estimated a 4× cost overrun; the real figure is **~17×**. The value **must** be set
explicitly on the processor at construction time.
### F5-3 — prompts must go through the chat template (`MEASURED`)
SmolVLM requires one `<image>` token per image in the prompt. Hand-written prompt strings raise
`ValueError`. Prompts are always built through `processor.apply_chat_template()`.
### P7-1 — the RemoteCLIP projected dimension is 512 (`MEASURED`)
The RemoteCLIP ViT-B/32 transformer width is **768**, but `visual.proj` maps to a **projected** dim of
**512**. The grounding head's per-cell feature is `4 × 512 = 2048`, declared in config so
`core/config.py` can validate the head **without importing torch**.
### C-1 — the availability mask is consumed by the head, not by CROMA (`MEASURED`)
CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied
by the **fusion head** (`input_dim = 3 × 768 + 12 + 2 = 2318`), not by the encoder.
### C-6 — T4 is SM 7.5, so training uses fp16, not bf16 (`MEASURED`)
The training precision is **fp16** because the target GPU (T4) is compute capability 7.5. bf16 is
not available there.
### C-7 — `image_resolution % 8 == 0` (`MEASURED`)
CROMA requires `image_resolution % 8 == 0`. The native value **120** yields 225 patches.
### C-8 — ZeroGPU does not support `torch.compile` (`MEASURED`)
`torch.compile` must never be enabled on the (historical) ZeroGPU target.
### C-9 — STANet change hyperparameters are upstream-verified (`MEASURED`)
Change detection uses STANet-style architecture with upstream-verified hyperparameters (PAM
self-attention, ResNet-18 encoder).
---
## 2. The grounding resolution decision — a pre-registered rejection (`REJECTED`)
**Question:** should grounding decode at 448 or 224?
**Answer: 224. 448 was rejected** — and the rejection is notable because it was *pre-registered* and
then *confirmed* by a paired test:
| Comparison (over identical samples, n = 16,159) | 448 vs 224 |
|---|---|
| mean best IoU | **−0.0147** |
| recall@0.5 | −0.0022 |
| recall@0.10 | −0.0699 |
| recall@0.25 | −0.0243 |
| latency | **1.59×** |
| paired 95 % CI | [−0.0160, −0.0134] |
| paired t | **−22.63** |
| 448 better on | 8.5 % of records |
| 448 worse on | **20.9 %** of records |
448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224.
This is a model of how a resolution decision should be made: declared in advance, then tested.
---
## 3. The router defect — a real bug, found and fixed
### 3.1 Symptom
The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered
**"River"** — instead of routing to `grounding`. A second query, *"Where is the new airport?"*,
behaved the same way.
### 3.2 Root cause
Two functions with different information:
- **`interpret()`** — produces the console's *reading*; **asset-count-blind** (text only).
- **`chooseTask()`** — performs *dispatch*; **asset-count-aware**.
The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell
through to the generic VQA specialist. See [`ARCHITECTURE.md`](ARCHITECTURE.md) §4.
### 3.3 Fix and verification (`RESOLVED`)
The fix was deployed to `SatQuery-Frontend` and validated by **three independent live passes**:
| Pass | Deployed HEAD | Result |
|---|---|---|
| 1 | `ff46eba42b18` + `d413d3672311` | 8/8 |
| 2 | `2d7ae53b482d` | 8/8 |
| 3 | `2d7ae53b482d` | 8/8 |
Both defect queries now dispatch to `grounding`:
| Query | Run id | Dispatched |
|---|---|---|
| Where are the built-up areas in this image? | `run_467ffa406f22` | `grounding` |
| Where is the new airport? | `run_46980ba55c62` | `grounding` |
24 live runs, 24 correct dispatches, **0 mock nodes**. Screenshots are in
[`../screenshots/`](../screenshots/).
---
## 4. The harness false-positive — caught before it could lie
An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome
**silently drops when the window lacks OS focus**. The harness therefore dispatched the page's
*default* query and still recorded a "result" — a **false pass**.
**Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`), and
uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`).
**Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected — its
answers were query-specific and the query text was embedded in the answers. The failure mode is
recorded because it is exactly the silent false-positive an evaluation harness must never have.
---
## 5. The `transport_mode: auto` fallthrough (`OPEN`)
`SATQUERY_TRANSPORT=auto` tries the tunnel, then falls through to the forward path on timeout. The
forward path to a **private** repo returns `302` quickly, but the wake step still consumes
`SATQUERY_WAKE_TIMEOUT_S` (120 s) first — so a worst-case failed request takes ≈ 249 s
(150 + 120). This is the root shape of the observed transient tunnel gap. Recorded as **OPEN**; a
deployed fix for the `codespace_name` newline on the wake path was authored separately.
---
## 6. The `interpret()` / `chooseTask()` asymmetry — intentional (`RESOLVED`)
For *"What changed between the earlier and later image?"* with **one** asset attached, the console
**reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the
reading describes the question's intent, the dispatch respects what can actually be computed with the
assets present. It is documented so it is not mistaken for a defect.
---
## 7. Environment findings (would otherwise cost hours)
| Finding | Detail |
|---|---|
| **Dead proxy in the sandbox** | outbound calls need `--noproxy '*'` (curl) or `ProxyHandler({})` (Python). |
| **pytest is only in the repo venv** | `.venv/Scripts/python.exe`; a bare `pytest` misses it. |
| **Full-suite pytest trips a bulk-delete guard** | sandbox-specific; affects `test_safe_delete_shim`. |
| **Cloudflare `_headers` concatenate** | two matching rules are merged, not overridden; Chromium takes the **first** `max-age`. |
| **Cloudflare 308-redirects `X.html` → `/X`** | reference the extensionless path. |
| **A forwarded Codespace port returns `302`** for a private repo | this is *why* the tunnel exists. |
| **Chrome drops synthetic CDP key events without OS focus** | the harness false-positive (§4). |
| **`browser-use` block-buffers stdout** even when redirected | needs explicit line buffering to stream. |
---
## 8. The BigEarthNet format contradiction (`ATTEMPTED`, reported not resolved)
The BigEarthNet data format **contradicts the original plan**. This was **reported rather than
silently patched**, because quietly changing the preprocessing would move the frozen config hash. The
BigEarthNet documentation does not specify a percentile stretch; this project nevertheless applies
percentile normalisation (2/98) to match the CROMA contract. That is a **deliberate, documented
choice**, not an upstream fact. See [`DATASETS.md`](DATASETS.md) §5.3.
|