Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/RESEARCH_NOTES.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 9.2 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/5ce82bef20f9510057a0a4f7bcbb0b5ee643d9cb/docs/RESEARCH_NOTES.md
- Command line
-
hf download hf://thundercode/SatQuery@5ce82bef20f9510057a0a4f7bcbb0b5ee643d9cb/docs/RESEARCH_NOTES.md
-
curl -L -o RESEARCH_NOTES.md https://huggingface.co/thundercode/SatQuery/resolve/5ce82bef20f9510057a0a4f7bcbb0b5ee643d9cb/docs/RESEARCH_NOTES.md
9.2 kB
| # Research Notes | |
| Engineering findings, negative results and design decisions that would otherwise be lost. Each was | |
| learned by **probe or execution**, not by assumption, and each is recorded so it is not rediscovered. | |
| **Status tags:** `MEASURED` Β· `RESOLVED` Β· `REJECTED` Β· `OPEN` Β· `ATTEMPTED`. | |
| --- | |
| ## 1. Findings that changed the code | |
| ### F4-1 β the MiniLM tokenizer ceiling (`MEASURED`) | |
| The MiniLM tokenizer's own ceiling is **256** (verified by probe). The project truncates to **128** β | |
| a deliberate truncation *well inside* the ceiling, not the model limit. Satellite queries are short; | |
| halving the sequence halves attention cost for no measurable accuracy loss. The encoder **asserts** | |
| `max_length β€ 256`, because truncating above the ceiling is a silent no-op. | |
| ### F4-2 β the router needs no GPU (`MEASURED`) | |
| The encoder is frozen, so embeddings are **cached** and the 50,822-parameter adapter trains on cached | |
| vectors. **Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.** No GPU required. | |
| ### F4-3 β splits must be by group (`MEASURED`) | |
| Splits are by **group** (template / hard-negative family), never by example. Hard-negative families | |
| are placed in the **test** split so their accuracy measures generalisation rather than memorisation. | |
| Splitting by example would leak template variants across the boundary. | |
| ### F5-1 β `AutoModelForVision2Seq` does not exist (`MEASURED`) | |
| In transformers 5.17.0, `AutoModelForVision2Seq` **does not exist** (not merely deprecated); | |
| `AutoModelForImageTextToText` is present. The loader is resolved by **feature detection**, never | |
| hardcoded. | |
| ### F5-2 β the processor cost overrun is ~17Γ, not 4Γ (`MEASURED`) | |
| The processor's default `longest_edge` is **2048**, which upscales 512-px tiles **4Γ** and then splits | |
| them (`do_image_splitting=True`) into **17 sub-images**: | |
| | Setting | `pixel_values` | prompt tokens | | |
| |---|---|---| | |
| | default | `(1, 17, 3, 512, 512)` | 1142 | | |
| | pinned (`processor_longest_edge: 512`) | `(1, 1, 3, 512, 512)` | β | | |
| The plan estimated a 4Γ cost overrun; the real figure is **~17Γ**. The value **must** be set | |
| explicitly on the processor at construction time. | |
| ### F5-3 β prompts must go through the chat template (`MEASURED`) | |
| SmolVLM requires one `<image>` token per image in the prompt. Hand-written prompt strings raise | |
| `ValueError`. Prompts are always built through `processor.apply_chat_template()`. | |
| ### P7-1 β the RemoteCLIP projected dimension is 512 (`MEASURED`) | |
| The RemoteCLIP ViT-B/32 transformer width is **768**, but `visual.proj` maps to a **projected** dim of | |
| **512**. The grounding head's per-cell feature is `4 Γ 512 = 2048`, declared in config so | |
| `core/config.py` can validate the head **without importing torch**. | |
| ### C-1 β the availability mask is consumed by the head, not by CROMA (`MEASURED`) | |
| CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied | |
| by the **fusion head** (`input_dim = 3 Γ 768 + 12 + 2 = 2318`), not by the encoder. | |
| ### C-6 β T4 is SM 7.5, so training uses fp16, not bf16 (`MEASURED`) | |
| The training precision is **fp16** because the target GPU (T4) is compute capability 7.5. bf16 is | |
| not available there. | |
| ### C-7 β `image_resolution % 8 == 0` (`MEASURED`) | |
| CROMA requires `image_resolution % 8 == 0`. The native value **120** yields 225 patches. | |
| ### C-8 β ZeroGPU does not support `torch.compile` (`MEASURED`) | |
| `torch.compile` must never be enabled on the (historical) ZeroGPU target. | |
| ### C-9 β STANet change hyperparameters are upstream-verified (`MEASURED`) | |
| Change detection uses STANet-style architecture with upstream-verified hyperparameters (PAM | |
| self-attention, ResNet-18 encoder). | |
| --- | |
| ## 2. The grounding resolution decision β a pre-registered rejection (`REJECTED`) | |
| **Question:** should grounding decode at 448 or 224? | |
| **Answer: 224. 448 was rejected** β and the rejection is notable because it was *pre-registered* and | |
| then *confirmed* by a paired test: | |
| | Comparison (over identical samples, n = 16,159) | 448 vs 224 | | |
| |---|---| | |
| | mean best IoU | **β0.0147** | | |
| | recall@0.5 | β0.0022 | | |
| | recall@0.10 | β0.0699 | | |
| | recall@0.25 | β0.0243 | | |
| | latency | **1.59Γ** | | |
| | paired 95 % CI | [β0.0160, β0.0134] | | |
| | paired t | **β22.63** | | |
| | 448 better on | 8.5 % of records | | |
| | 448 worse on | **20.9 %** of records | | |
| 448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224. | |
| This is a model of how a resolution decision should be made: declared in advance, then tested. | |
| --- | |
| ## 3. The router defect β a real bug, found and fixed | |
| ### 3.1 Symptom | |
| The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered | |
| **"River"** β instead of routing to `grounding`. A second query, *"Where is the new airport?"*, | |
| behaved the same way. | |
| ### 3.2 Root cause | |
| Two functions with different information: | |
| - **`interpret()`** β produces the console's *reading*; **asset-count-blind** (text only). | |
| - **`chooseTask()`** β performs *dispatch*; **asset-count-aware**. | |
| The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell | |
| through to the generic VQA specialist. See [`ARCHITECTURE.md`](ARCHITECTURE.md) Β§4. | |
| ### 3.3 Fix and verification (`RESOLVED`) | |
| The fix was deployed to `SatQuery-Frontend` and validated by **three independent live passes**: | |
| | Pass | Deployed HEAD | Result | | |
| |---|---|---| | |
| | 1 | `ff46eba42b18` + `d413d3672311` | 8/8 | | |
| | 2 | `2d7ae53b482d` | 8/8 | | |
| | 3 | `2d7ae53b482d` | 8/8 | | |
| Both defect queries now dispatch to `grounding`: | |
| | Query | Run id | Dispatched | | |
| |---|---|---| | |
| | Where are the built-up areas in this image? | `run_467ffa406f22` | `grounding` | | |
| | Where is the new airport? | `run_46980ba55c62` | `grounding` | | |
| 24 live runs, 24 correct dispatches, **0 mock nodes**. Screenshots are in | |
| [`../screenshots/`](../screenshots/). | |
| --- | |
| ## 4. The harness false-positive β caught before it could lie | |
| An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome | |
| **silently drops when the window lacks OS focus**. The harness therefore dispatched the page's | |
| *default* query and still recorded a "result" β a **false pass**. | |
| **Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`), and | |
| uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`). | |
| **Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected β its | |
| answers were query-specific and the query text was embedded in the answers. The failure mode is | |
| recorded because it is exactly the silent false-positive an evaluation harness must never have. | |
| --- | |
| ## 5. The `transport_mode: auto` fallthrough (`OPEN`) | |
| `SATQUERY_TRANSPORT=auto` tries the tunnel, then falls through to the forward path on timeout. The | |
| forward path to a **private** repo returns `302` quickly, but the wake step still consumes | |
| `SATQUERY_WAKE_TIMEOUT_S` (120 s) first β so a worst-case failed request takes β 249 s | |
| (150 + 120). This is the root shape of the observed transient tunnel gap. Recorded as **OPEN**; a | |
| deployed fix for the `codespace_name` newline on the wake path was authored separately. | |
| --- | |
| ## 6. The `interpret()` / `chooseTask()` asymmetry β intentional (`RESOLVED`) | |
| For *"What changed between the earlier and later image?"* with **one** asset attached, the console | |
| **reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the | |
| reading describes the question's intent, the dispatch respects what can actually be computed with the | |
| assets present. It is documented so it is not mistaken for a defect. | |
| --- | |
| ## 7. Environment findings (would otherwise cost hours) | |
| | Finding | Detail | | |
| |---|---| | |
| | **Dead proxy in the sandbox** | outbound calls need `--noproxy '*'` (curl) or `ProxyHandler({})` (Python). | | |
| | **pytest is only in the repo venv** | `.venv/Scripts/python.exe`; a bare `pytest` misses it. | | |
| | **Full-suite pytest trips a bulk-delete guard** | sandbox-specific; affects `test_safe_delete_shim`. | | |
| | **Cloudflare `_headers` concatenate** | two matching rules are merged, not overridden; Chromium takes the **first** `max-age`. | | |
| | **Cloudflare 308-redirects `X.html` β `/X`** | reference the extensionless path. | | |
| | **A forwarded Codespace port returns `302`** for a private repo | this is *why* the tunnel exists. | | |
| | **Chrome drops synthetic CDP key events without OS focus** | the harness false-positive (Β§4). | | |
| | **`browser-use` block-buffers stdout** even when redirected | needs explicit line buffering to stream. | | |
| --- | |
| ## 8. The BigEarthNet format contradiction (`ATTEMPTED`, reported not resolved) | |
| The BigEarthNet data format **contradicts the original plan**. This was **reported rather than | |
| silently patched**, because quietly changing the preprocessing would move the frozen config hash. The | |
| BigEarthNet documentation does not specify a percentile stretch; this project nevertheless applies | |
| percentile normalisation (2/98) to match the CROMA contract. That is a **deliberate, documented | |
| choice**, not an upstream fact. See [`DATASETS.md`](DATASETS.md) Β§5.3. | |