Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/TRAINING.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 8.37 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/TRAINING.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/docs/TRAINING.md
-
curl -L -o TRAINING.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/TRAINING.md
8.37 kB
| # Training | |
| **Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `REJECTED`. | |
| The project trains **small modules on frozen backbones**. No backbone is fine-tuned end-to-end. Every | |
| hyperparameter lives in `configs/base.yaml` (no magic numbers in Python), and every trained artifact | |
| records the frozen config hash it was trained against. | |
| --- | |
| ## 1. Overview | |
| | Artifact | Backbone (frozen) | Where trained | Selection signal | | |
| |---|---|---|---| | |
| | `router` adapter | `all-MiniLM-L6-v2` | local CPU | val (group-split) | | |
| | `grounding` head | RemoteCLIP ViT-B/32 | local | val fraction 0.10 | | |
| | `change` head | STANet ResNet-18 + PAM | local | val (LEVIR split) | | |
| | `optical_sar` fusion head | CROMA-base | local, **seed sweep** | held-out test (pre-registered) | | |
| | `change_vqa` head | over cached change features | **external GPU (Kaggle)** | val answer accuracy | | |
| | `vlm` LoRA adapter | SmolVLM-500M-Instruct | **external GPU** | frozen 1000-question subset | | |
| **CPU-first.** The router and every head except the VLM adapter train on CPU. The VLM LoRA adapter | |
| requires a GPU (T4-class). | |
| ## 2. Router adapter | |
| A **50,822-parameter** adapter over the frozen MiniLM encoder. | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | epochs | 60 | | |
| | batch size | 64 | | |
| | learning rate | 0.001 | | |
| | weight decay | 0.01 | | |
| | task loss weight | 1.0 | | |
| | modality loss weight | 0.3 | | |
| | binary loss weight | 0.5 | | |
| | val ratio | 0.15 | | |
| | hard negatives to test | true | | |
| **Finding F4-2 — the encoder is frozen, so embeddings are cached** and the adapter trains on cached | |
| vectors. **Measured: 20 epochs / 4,096 vectors in 0.28 s on CPU.** No GPU is required. | |
| **Finding F4-3 — splits are by GROUP** (template / hard-negative family), never by example. | |
| Hard-negative families are placed in the **test** split so their accuracy measures generalisation | |
| rather than memorisation. | |
| > **Status:** the adapter is trained and shipped. Its measured number (0.965116) is **validation-only, | |
| > ungated, n = 86**; the router **test split was NOT RUN**. | |
| ## 3. Grounding head | |
| A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | learning rate | 0.0001 | | |
| | batch size | 16 | | |
| | epochs | 20 | | |
| | weight decay | 0.0001 | | |
| | warmup ratio | 0.05 | | |
| | grad clip | 1.0 | | |
| | val fraction | 0.10 | | |
| | save every steps | 500 | | |
| | box loss weight | 0.5 | | |
| | GIoU loss weight | 0.3 | | |
| | confidence loss weight | 0.2 | | |
| | `positive_confidence_weight` | **20.0** | | |
| **Architecture constraint (enforced, not documented).** Per-cell feature is | |
| `concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048`. `core/config.py` **rejects any | |
| value other than** `4 × grounding.encoder_projected_dim` at load time, and the specialist asserts the | |
| same 512 against the real model — because a mismatch is a *silent* shape error that torch only raises | |
| at the similarity step, after patch features are already cached. | |
| **Why `positive_confidence_weight = 20.0`.** Objectness BCE sees ~1 positive cell out of 49. | |
| Unweighted, the optimum is "no object" everywhere; the weight is what stops that collapse. | |
| **Resolution is frozen at 224.** 448 was evaluated and **REJECTED** (paired test: mean diff −0.0147, | |
| 95 % CI [−0.0160, −0.0134], t = −22.63, at 1.59× latency). | |
| ## 4. Change head | |
| STANet-style Siamese detector. | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | encoder | ResNet-18 | | |
| | self-attention | **PAM** (BAM alternative not used) | | |
| | tile size | 256 | | |
| | tile overlap | 0 | | |
| | threshold | 0.50 | | |
| | min component pixels | 32 | | |
| | learning rate | 0.001 | | |
| | batch size | 8 | | |
| | BCE weight | 0.5 | | |
| | Dice weight | 0.5 | | |
| Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048). **This is the only task with a | |
| `VERIFIED` headline metric** (pooled IoU 0.8122 on the immutable test split). | |
| ## 5. Optical-SAR fusion head | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | input dim | 2318 = 3 × 768 + 12 + 2 | | |
| | hidden dim | 512 | | |
| | dropout | 0.2 | | |
| | num classes | 19 (BigEarthNet CLC) | | |
| **Seed sweep.** Training was run as two arms (**armA**, **armB**) × five seeds (**100–104**), with a | |
| per-arm variance report (`armA_seed_variance_report.json`). The **production head** is a distinct, | |
| frozen artifact (`fusion_head_production_v001/head.pt`) with its own | |
| `production_head_record.json` and a `phase12_rerun_verification.json`. | |
| **Pre-registration.** The headline metric is a **pre-registered** 115-class protocol | |
| (`pre_registered_11.5`) computed over the 19-class label space on the held-out test split. The | |
| metric JSON records `is_deciding_statistic: False`, i.e. it is a reported measurement, not a | |
| decision statistic. | |
| **Feature caches.** Training consumes cached CROMA features | |
| (`fusion_features/`, `fusion_features_armB/`, ~231 MB each). These caches are **reproducible** and are | |
| not released as model weights. | |
| ## 6. Change-VQA head — trained externally | |
| The change-VQA head was trained **outside this repository**, on an external GPU (Kaggle), following | |
| `docs/R02_KAGGLE_TRAINING_GUIDE.md`. That guide's status on entry was | |
| `IMPLEMENTATION_READY_FOR_EXTERNAL_TRAINING`, and its explicit contract is: | |
| > **Training produces an artifact, not a verified capability, and the run record says | |
| > `TRAINED_UNVERIFIED`.** | |
| The returned checkpoint was **promoted** through a byte-identity gate | |
| (`artifacts/change_vqa/run/PROMOTION.json`): | |
| | Property | Value | | |
| |---|---| | |
| | sha256 | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` | | |
| | bytes | 5,822,809 | | |
| | architecture | `change_vqa_head_v1` | | |
| | parameters | 1,453,912 | | |
| | non-finite tensors | **0** | | |
| | weights modified during promotion | **false** | | |
| | byte-identical to source | **true** | | |
| | hash agrees across | `model_metadata.json`, `run_record.json`, `hashes.json` | | |
| **Selection:** epoch **8**, chosen on **Val answer accuracy = 0.700018**, stopped by early stopping. | |
| Seed 42. It trains on **cached change + text features** (specs `change_feat_v1`, | |
| `change_cache_spec c801326f85a185f8`, `text_cache_spec d2801ea1a314354a`), not on raw imagery. | |
| > **Note.** The raw CDVQA loader (see [`DATASETS.md`](DATASETS.md)) loads examples but has **no | |
| > training loop**; the shipped head is a cached-feature model. These are different paths and are not | |
| > conflated. | |
| ## 7. VLM LoRA adapter — trained externally | |
| A PEFT LoRA adapter on **frozen** `HuggingFaceTB/SmolVLM-500M-Instruct`. | |
| | Hyperparameter | Value | | |
| |---|---| | |
| | PEFT version | **0.19.1** | | |
| | `r` (rank) | 16 | | |
| | `alpha` | 32 | | |
| | `dropout` | 0.05 | | |
| | target modules | `model.text_model.*.{q,k,v,o,gate,up,down}_proj` | | |
| | precision | **fp16** (finding C-6: T4 is SM 7.5 → **fp16, NOT bf16**) | | |
| | batch size | 2 | | |
| | gradient accumulation | 8 | | |
| | learning rate | 0.0002 | | |
| | epochs | 1 | | |
| | gradient checkpointing | true | | |
| | save every steps | 500 | | |
| **Finding F5-2 (cost).** The processor's default `longest_edge` is 2048, which upscales 512-px tiles | |
| 4× and then splits them into **17 sub-images** (`pixel_values (1, 17, 3, 512, 512)`, 1142 prompt | |
| tokens). Pinning `processor_longest_edge: 512` yields `pixel_values (1, 1, 3, 512, 512)`. The plan | |
| estimated a 4× cost overrun; the **measured** figure is ~17×. | |
| **Finding F5-3.** SmolVLM requires one `<image>` token per image in the prompt; hand-written prompt | |
| strings raise `ValueError`. Prompts are always built through `processor.apply_chat_template()`. | |
| **Outcome:** metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is | |
| **ACCEPTANCE-REJECTED** for promotion. The deployed caption/VQA path uses the **unadapted** model. | |
| See [`MODELS.md`](MODELS.md) §3.6. | |
| ## 8. Reproducibility contract for training | |
| - **Seed 42** everywhere (`project.seed`). | |
| - **Precision `fp16`** (T4 constraint), never bf16. | |
| - Every artifact records the **frozen config hash** `78f1e3700da15aa1`; a config edit moves the hash | |
| and invalidates the artifact. | |
| - `save_every_steps: 500`; checkpoints are archived as provenance, not released. | |
| - Training guides state their own entry status and **never** claim a trained artifact is a verified | |
| capability. | |
| ## 9. What was NOT trained | |
| | Item | State | | |
| |---|---| | |
| | Backbone fine-tuning (any) | **NOT DONE** — all backbones frozen | | |
| | Router on the test split | **NOT RUN** | | |
| | Any end-to-end / joint training | **NOT RUN** | | |
| | Re-training of the change head at a second resolution | **NOT RUN** | | |
| | Benchmark adapters | **NOT RUN** | | |