Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download docs/RESEARCH_NOTES.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 11.5 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/6df4c6ebaa2cfee7087723ec06fe4dee190b5634/docs/RESEARCH_NOTES.md
- Command line
-
hf download hf://thundercode/SatQuery@6df4c6ebaa2cfee7087723ec06fe4dee190b5634/docs/RESEARCH_NOTES.md
-
curl -L -o RESEARCH_NOTES.md https://huggingface.co/thundercode/SatQuery/resolve/6df4c6ebaa2cfee7087723ec06fe4dee190b5634/docs/RESEARCH_NOTES.md
Research Notes
Engineering findings, negative results, and design decisions that would otherwise be lost. Every one was learned by probe or execution, not assumption, and every one is recorded so it is not rediscovered.
Status tags: MEASURED · RESOLVED · REJECTED · OPEN · ATTEMPTED.
1. Findings that changed the code
F4-1 — the MiniLM tokenizer ceiling (MEASURED)
The MiniLM tokenizer's own ceiling is 256 (verified by probe). This project truncates to 128 —
a deliberate truncation well inside the ceiling, not the model limit. Satellite queries are short;
halving the sequence halves attention cost for no measurable accuracy loss. The encoder asserts
max_length ≤ 256, because truncating above the ceiling is a silent no-op.
Consequence: router.max_length: 128 with an assertion, rather than a comment.
F4-2 — the router needs no GPU (MEASURED)
The encoder is frozen, so embeddings are cached and the 50,822-parameter adapter trains on cached vectors. Measured on CPU: 20 epochs / 4,096 vectors in 0.28 s.
Consequence: the router can be retrained during development without a GPU.
F4-3 — splits must be by group (MEASURED)
Splits are by group (template / hard-negative family), never by example. Hard-negative families are placed in the test split so their accuracy measures generalisation rather than memorisation. Splitting by example would leak template variants across the boundary and inflate the score.
F5-1 — AutoModelForVision2Seq does not exist (MEASURED)
In transformers 5.17.0, AutoModelForVision2Seq does not exist (it is not merely deprecated);
AutoModelForImageTextToText is present.
Consequence: the loader is resolved by feature detection, never hardcoded to one class name.
F5-2 — the processor cost overrun is ~17×, not 4× (MEASURED)
The processor's default longest_edge is 2048, which upscales a 512-px tile 4× and then
splits it (do_image_splitting=True) into sub-images:
| Setting | pixel_values |
prompt tokens |
|---|---|---|
| default | (1, 17, 3, 512, 512) |
1142 |
pinned (processor_longest_edge: 512) |
(1, 1, 3, 512, 512) |
— |
The plan estimated a 4× cost overrun; the real figure is ~17×. The value must be set
explicitly on the processor at construction time, and core/config.py now enforces
processor_longest_edge <= image.tile_size so the pin is a control rather than a comment.
F5-3 — prompts must go through the chat template (MEASURED)
SmolVLM requires one <image> token per image in the prompt. Hand-written prompt strings raise
ValueError.
Consequence: prompts are always built through processor.apply_chat_template(), enforced by the
config key vlm.prompt_must_use_chat_template: true.
P7-1 — the RemoteCLIP projected dimension is 512 (MEASURED)
The RemoteCLIP ViT-B/32 transformer width is 768, but visual.proj maps to a projected dim of
512. The grounding head's per-cell feature is therefore 4 × 512 = 2048.
Consequence: grounding.encoder_projected_dim: 512 is declared in config so core/config.py can
validate the head without importing torch, and specialists/grounding/remoteclip.py asserts the
same value against the real model at load time.
C-1 — the availability mask is consumed by the head, not by CROMA (MEASURED)
CROMA always sees the canonical channel counts (12 optical, 2 SAR). The availability mask is applied
by the fusion head (input_dim = 3 × 768 + 12 + 2 = 2318), not by the encoder.
C-6 — T4 is SM 7.5, so training uses fp16, not bf16 (MEASURED)
training.precision: fp16 because the target GPU (T4) is compute capability 7.5; bf16 is unavailable
there. The loader validates the value is one of fp16|bf16|fp32.
C-7 — image_resolution % 8 == 0 (MEASURED)
CROMA requires image_resolution % 8 == 0. The native value 120 yields 225 patches. Enforced at
config load.
C-8 — ZeroGPU does not support torch.compile (MEASURED)
torch.compile must never be enabled on the (historical) ZeroGPU target. Enforced: the loader
fails startup if deployment.torch_compile is true.
C-9 — STANet hyperparameters are upstream-verified (MEASURED)
Change detection uses a STANet-style architecture with upstream-verified hyperparameters: ResNet-18 encoder, PAM self-attention mode, tile 256, threshold 0.50, BCE 0.5 + Dice 0.5.
2. The grounding resolution decision — a pre-registered rejection (REJECTED)
Question: should grounding decode at 448 or 224?
Answer: 224. 448 was rejected — notable because the rejection was pre-registered and then confirmed by a paired test over identical samples (n = 16,159):
| Comparison (448 vs 224) | Value |
|---|---|
| mean best IoU | −0.0147 |
| recall@0.5 | −0.0022 |
| recall@0.10 | −0.0699 |
| recall@0.25 | −0.0243 |
| latency | 1.59× |
| paired mean difference | −0.0147 |
| paired 95 % CI | [−0.0160, −0.0134] |
| paired t | −22.63 |
| 448 better on | 8.5 % of records |
| 448 worse on | 20.9 % of records |
448 lost on every axis. The pre-registered decision rule and the paired test agree on 224. This is the model for how a resolution decision should be made: declared in advance, then tested.
3. The router defect — a real bug, found and fixed (RESOLVED)
3.1 Symptom
The query "Where are the built-up areas in this image?" collapsed to vqa and answered
"River" — instead of routing to grounding. A second query, "Where is the new airport?",
behaved the same way.
3.2 Root cause
Two functions with different information:
interpret()— produces the console's reading; asset-count-blind (text only).chooseTask()— performs dispatch; asset-count-aware.
The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell through to the generic VQA specialist.
3.3 Fix and verification
The fix was deployed to SatQuery-Frontend and validated by three independent live passes:
| Pass | Deployed HEAD | Result |
|---|---|---|
| 1 | ff46eba42b18 + d413d3672311 |
8/8 |
| 2 | 2d7ae53b482d |
8/8 |
| 3 | 2d7ae53b482d |
8/8 |
Both defect queries now dispatch to grounding:
| Query | Run id | Dispatched |
|---|---|---|
| Where are the built-up areas in this image? | run_467ffa406f22 |
grounding |
| Where is the new airport? | run_46980ba55c62 |
grounding |
24 live runs, 24 correct dispatches, 0 mock nodes. Screenshots are in
../screenshots/.
4. The harness false-positive — caught before it could lie (RESOLVED)
An earlier live-validation harness typed queries with synthetic CDP key events, which Chrome silently drops when the window lacks OS focus. The harness therefore dispatched the page's default query and still recorded a "result" — a false pass.
Fix: the current harness asserts form state before dispatch (q_ok, obs_ok, t0_ok) and
uses deterministic query entry (js() value-set + type_text() via CDP Input.insertText).
Independent check: the earlier 8/8 run was re-examined and confirmed not infected — its answers were query-specific and the query text was embedded in the answers. The failure mode is recorded because it is exactly the silent false-positive an evaluation harness must never have.
5. The transport_mode: auto fallthrough (OPEN)
SATQUERY_TRANSPORT=auto tries the tunnel, then falls through to the forward path on timeout. The
forward path to a private repo returns 302 quickly, but the wake step still consumes
SATQUERY_WAKE_TIMEOUT_S (120 s) first — so a worst-case failed request takes ≈ 249 s
(150 + 120). This is the root shape of the observed transient tunnel gap.
A patch adding forward_unavailable (503) and upstream_timeout (504) codes plus the
codespace_name .strip() fix was authored and verified (py_compile clean). Status: OPEN — the
patch is prepared but NOT deployed.
6. The interpret() / chooseTask() asymmetry — intentional (RESOLVED)
For "What changed between the earlier and later image?" with one asset attached, the console
reads change while dispatch correctly falls back to change_vqa. This is not a bug: the
reading describes the question's intent; the dispatch respects what can actually be computed with the
assets present. It is documented so it is not mistaken for a defect.
7. Environment findings (would otherwise cost hours)
| Finding | Detail |
|---|---|
| Dead proxy in the authoring sandbox | outbound calls need --noproxy '*' (curl) or ProxyHandler({}) (Python). |
| The sandbox proxy is slow for uploads | huggingface_hub uploads stalled at ~51 kB/s through http_proxy=127.0.0.1:58294; the fix is to unset http_proxy/https_proxy and set no_proxy='*'. |
hf_hub_download returned an EMPTY file |
sha256 e3b0c442… (the empty-content hash), which produced a false FAIL for all six artifacts. The honest check is a direct HTTPS download with ProxyHandler({}). |
| pytest is only in the repo venv | .venv/Scripts/python.exe; a bare pytest misses it. |
| The full test suite trips a bulk-delete guard | sandbox-specific; affects test_safe_delete_shim. |
Cloudflare _headers concatenate |
two matching rules are merged, not overridden; Chromium takes the first max-age. |
Cloudflare 308-redirects X.html → /X |
reference the extensionless path. |
A forwarded Codespace port returns 302 |
for a private repo — this is why the tunnel exists. |
| Chrome drops synthetic CDP key events without OS focus | the harness false-positive (§4). |
browser-use block-buffers stdout |
even when redirected; needs explicit line buffering to stream. |
8. The BigEarthNet format contradiction (ATTEMPTED, reported not resolved)
The BigEarthNet data format contradicts the original plan. This was reported rather than silently patched, because quietly changing the preprocessing would move the frozen config hash.
Two facts matter:
- The BigEarthNet documentation — its uses, mentions, or endorsements — does not specify a percentile stretch. This project nevertheless applies percentile normalisation (2/98) for optical inputs to match the CROMA contract. That is a deliberate, documented choice, not an upstream fact.
- The local subset is 100 % single-label, against the official 1–11 multi-label scheme, so metrics computed on it are not comparable to published multi-label numbers.
9. Where the evidence lives
| Topic | Evidence |
|---|---|
| Findings F4-1…F5-3, C-1…C-9, P7-1 | docs/PHASE*.md, docs/ARCHITECTURE_*.md, docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md |
| The 448-vs-224 paired test | docs/PHASE7_RESOLUTION_DECISION.md |
| The VLM rejection | docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md, artifacts/vlm/phase6_closure.json |
| Router defect + 3 live passes | .workbuddy-ai/scratch/live_validation/ (run_output.txt, run_final2.txt, run_final3.txt) |
| The undeployed B-07 patch | session scratch fix-b07-forward-unavailable.patch |
| Live validation harness | .workbuddy-ai/scratch/run_all_postfix2.harness, recompute_verdicts.py |