Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 8,372 Bytes
4e54b93 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 | # Training
**Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `REJECTED`.
The project trains **small modules on frozen backbones**. No backbone is fine-tuned end-to-end. Every
hyperparameter lives in `configs/base.yaml` (no magic numbers in Python), and every trained artifact
records the frozen config hash it was trained against.
---
## 1. Overview
| Artifact | Backbone (frozen) | Where trained | Selection signal |
|---|---|---|---|
| `router` adapter | `all-MiniLM-L6-v2` | local CPU | val (group-split) |
| `grounding` head | RemoteCLIP ViT-B/32 | local | val fraction 0.10 |
| `change` head | STANet ResNet-18 + PAM | local | val (LEVIR split) |
| `optical_sar` fusion head | CROMA-base | local, **seed sweep** | held-out test (pre-registered) |
| `change_vqa` head | over cached change features | **external GPU (Kaggle)** | val answer accuracy |
| `vlm` LoRA adapter | SmolVLM-500M-Instruct | **external GPU** | frozen 1000-question subset |
**CPU-first.** The router and every head except the VLM adapter train on CPU. The VLM LoRA adapter
requires a GPU (T4-class).
## 2. Router adapter
A **50,822-parameter** adapter over the frozen MiniLM encoder.
| Hyperparameter | Value |
|---|---|
| epochs | 60 |
| batch size | 64 |
| learning rate | 0.001 |
| weight decay | 0.01 |
| task loss weight | 1.0 |
| modality loss weight | 0.3 |
| binary loss weight | 0.5 |
| val ratio | 0.15 |
| hard negatives to test | true |
**Finding F4-2 — the encoder is frozen, so embeddings are cached** and the adapter trains on cached
vectors. **Measured: 20 epochs / 4,096 vectors in 0.28 s on CPU.** No GPU is required.
**Finding F4-3 — splits are by GROUP** (template / hard-negative family), never by example.
Hard-negative families are placed in the **test** split so their accuracy measures generalisation
rather than memorisation.
> **Status:** the adapter is trained and shipped. Its measured number (0.965116) is **validation-only,
> ungated, n = 86**; the router **test split was NOT RUN**.
## 3. Grounding head
A trainable head over the frozen RemoteCLIP ViT-B/32 encoder.
| Hyperparameter | Value |
|---|---|
| learning rate | 0.0001 |
| batch size | 16 |
| epochs | 20 |
| weight decay | 0.0001 |
| warmup ratio | 0.05 |
| grad clip | 1.0 |
| val fraction | 0.10 |
| save every steps | 500 |
| box loss weight | 0.5 |
| GIoU loss weight | 0.3 |
| confidence loss weight | 0.2 |
| `positive_confidence_weight` | **20.0** |
**Architecture constraint (enforced, not documented).** Per-cell feature is
`concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048`. `core/config.py` **rejects any
value other than** `4 × grounding.encoder_projected_dim` at load time, and the specialist asserts the
same 512 against the real model — because a mismatch is a *silent* shape error that torch only raises
at the similarity step, after patch features are already cached.
**Why `positive_confidence_weight = 20.0`.** Objectness BCE sees ~1 positive cell out of 49.
Unweighted, the optimum is "no object" everywhere; the weight is what stops that collapse.
**Resolution is frozen at 224.** 448 was evaluated and **REJECTED** (paired test: mean diff −0.0147,
95 % CI [−0.0160, −0.0134], t = −22.63, at 1.59× latency).
## 4. Change head
STANet-style Siamese detector.
| Hyperparameter | Value |
|---|---|
| encoder | ResNet-18 |
| self-attention | **PAM** (BAM alternative not used) |
| tile size | 256 |
| tile overlap | 0 |
| threshold | 0.50 |
| min component pixels | 32 |
| learning rate | 0.001 |
| batch size | 8 |
| BCE weight | 0.5 |
| Dice weight | 0.5 |
Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048). **This is the only task with a
`VERIFIED` headline metric** (pooled IoU 0.8122 on the immutable test split).
## 5. Optical-SAR fusion head
| Hyperparameter | Value |
|---|---|
| input dim | 2318 = 3 × 768 + 12 + 2 |
| hidden dim | 512 |
| dropout | 0.2 |
| num classes | 19 (BigEarthNet CLC) |
**Seed sweep.** Training was run as two arms (**armA**, **armB**) × five seeds (**100–104**), with a
per-arm variance report (`armA_seed_variance_report.json`). The **production head** is a distinct,
frozen artifact (`fusion_head_production_v001/head.pt`) with its own
`production_head_record.json` and a `phase12_rerun_verification.json`.
**Pre-registration.** The headline metric is a **pre-registered** 115-class protocol
(`pre_registered_11.5`) computed over the 19-class label space on the held-out test split. The
metric JSON records `is_deciding_statistic: False`, i.e. it is a reported measurement, not a
decision statistic.
**Feature caches.** Training consumes cached CROMA features
(`fusion_features/`, `fusion_features_armB/`, ~231 MB each). These caches are **reproducible** and are
not released as model weights.
## 6. Change-VQA head — trained externally
The change-VQA head was trained **outside this repository**, on an external GPU (Kaggle), following
`docs/R02_KAGGLE_TRAINING_GUIDE.md`. That guide's status on entry was
`IMPLEMENTATION_READY_FOR_EXTERNAL_TRAINING`, and its explicit contract is:
> **Training produces an artifact, not a verified capability, and the run record says
> `TRAINED_UNVERIFIED`.**
The returned checkpoint was **promoted** through a byte-identity gate
(`artifacts/change_vqa/run/PROMOTION.json`):
| Property | Value |
|---|---|
| sha256 | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` |
| bytes | 5,822,809 |
| architecture | `change_vqa_head_v1` |
| parameters | 1,453,912 |
| non-finite tensors | **0** |
| weights modified during promotion | **false** |
| byte-identical to source | **true** |
| hash agrees across | `model_metadata.json`, `run_record.json`, `hashes.json` |
**Selection:** epoch **8**, chosen on **Val answer accuracy = 0.700018**, stopped by early stopping.
Seed 42. It trains on **cached change + text features** (specs `change_feat_v1`,
`change_cache_spec c801326f85a185f8`, `text_cache_spec d2801ea1a314354a`), not on raw imagery.
> **Note.** The raw CDVQA loader (see [`DATASETS.md`](DATASETS.md)) loads examples but has **no
> training loop**; the shipped head is a cached-feature model. These are different paths and are not
> conflated.
## 7. VLM LoRA adapter — trained externally
A PEFT LoRA adapter on **frozen** `HuggingFaceTB/SmolVLM-500M-Instruct`.
| Hyperparameter | Value |
|---|---|
| PEFT version | **0.19.1** |
| `r` (rank) | 16 |
| `alpha` | 32 |
| `dropout` | 0.05 |
| target modules | `model.text_model.*.{q,k,v,o,gate,up,down}_proj` |
| precision | **fp16** (finding C-6: T4 is SM 7.5 → **fp16, NOT bf16**) |
| batch size | 2 |
| gradient accumulation | 8 |
| learning rate | 0.0002 |
| epochs | 1 |
| gradient checkpointing | true |
| save every steps | 500 |
**Finding F5-2 (cost).** The processor's default `longest_edge` is 2048, which upscales 512-px tiles
4× and then splits them into **17 sub-images** (`pixel_values (1, 17, 3, 512, 512)`, 1142 prompt
tokens). Pinning `processor_longest_edge: 512` yields `pixel_values (1, 1, 3, 512, 512)`. The plan
estimated a 4× cost overrun; the **measured** figure is ~17×.
**Finding F5-3.** SmolVLM requires one `<image>` token per image in the prompt; hand-written prompt
strings raise `ValueError`. Prompts are always built through `processor.apply_chat_template()`.
**Outcome:** metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is
**ACCEPTANCE-REJECTED** for promotion. The deployed caption/VQA path uses the **unadapted** model.
See [`MODELS.md`](MODELS.md) §3.6.
## 8. Reproducibility contract for training
- **Seed 42** everywhere (`project.seed`).
- **Precision `fp16`** (T4 constraint), never bf16.
- Every artifact records the **frozen config hash** `78f1e3700da15aa1`; a config edit moves the hash
and invalidates the artifact.
- `save_every_steps: 500`; checkpoints are archived as provenance, not released.
- Training guides state their own entry status and **never** claim a trained artifact is a verified
capability.
## 9. What was NOT trained
| Item | State |
|---|---|
| Backbone fine-tuning (any) | **NOT DONE** — all backbones frozen |
| Router on the test split | **NOT RUN** |
| Any end-to-end / joint training | **NOT RUN** |
| Re-training of the change head at a second resolution | **NOT RUN** |
| Benchmark adapters | **NOT RUN** |
|