# Training **Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `REJECTED`. The project trains **small modules on frozen backbones**. No backbone is fine-tuned end-to-end. Every hyperparameter lives in `configs/base.yaml` (no magic numbers in Python), and every trained artifact records the frozen config hash it was trained against. --- ## 1. Overview | Artifact | Backbone (frozen) | Where trained | Selection signal | |---|---|---|---| | `router` adapter | `all-MiniLM-L6-v2` | local CPU | val (group-split) | | `grounding` head | RemoteCLIP ViT-B/32 | local | val fraction 0.10 | | `change` head | STANet ResNet-18 + PAM | local | val (LEVIR split) | | `optical_sar` fusion head | CROMA-base | local, **seed sweep** | held-out test (pre-registered) | | `change_vqa` head | over cached change features | **external GPU (Kaggle)** | val answer accuracy | | `vlm` LoRA adapter | SmolVLM-500M-Instruct | **external GPU** | frozen 1000-question subset | **CPU-first.** The router and every head except the VLM adapter train on CPU. The VLM LoRA adapter requires a GPU (T4-class). ## 2. Router adapter A **50,822-parameter** adapter over the frozen MiniLM encoder. | Hyperparameter | Value | |---|---| | epochs | 60 | | batch size | 64 | | learning rate | 0.001 | | weight decay | 0.01 | | task loss weight | 1.0 | | modality loss weight | 0.3 | | binary loss weight | 0.5 | | val ratio | 0.15 | | hard negatives to test | true | **Finding F4-2 — the encoder is frozen, so embeddings are cached** and the adapter trains on cached vectors. **Measured: 20 epochs / 4,096 vectors in 0.28 s on CPU.** No GPU is required. **Finding F4-3 — splits are by GROUP** (template / hard-negative family), never by example. Hard-negative families are placed in the **test** split so their accuracy measures generalisation rather than memorisation. > **Status:** the adapter is trained and shipped. Its measured number (0.965116) is **validation-only, > ungated, n = 86**; the router **test split was NOT RUN**. ## 3. Grounding head A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. | Hyperparameter | Value | |---|---| | learning rate | 0.0001 | | batch size | 16 | | epochs | 20 | | weight decay | 0.0001 | | warmup ratio | 0.05 | | grad clip | 1.0 | | val fraction | 0.10 | | save every steps | 500 | | box loss weight | 0.5 | | GIoU loss weight | 0.3 | | confidence loss weight | 0.2 | | `positive_confidence_weight` | **20.0** | **Architecture constraint (enforced, not documented).** Per-cell feature is `concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048`. `core/config.py` **rejects any value other than** `4 × grounding.encoder_projected_dim` at load time, and the specialist asserts the same 512 against the real model — because a mismatch is a *silent* shape error that torch only raises at the similarity step, after patch features are already cached. **Why `positive_confidence_weight = 20.0`.** Objectness BCE sees ~1 positive cell out of 49. Unweighted, the optimum is "no object" everywhere; the weight is what stops that collapse. **Resolution is frozen at 224.** 448 was evaluated and **REJECTED** (paired test: mean diff −0.0147, 95 % CI [−0.0160, −0.0134], t = −22.63, at 1.59× latency). ## 4. Change head STANet-style Siamese detector. | Hyperparameter | Value | |---|---| | encoder | ResNet-18 | | self-attention | **PAM** (BAM alternative not used) | | tile size | 256 | | tile overlap | 0 | | threshold | 0.50 | | min component pixels | 32 | | learning rate | 0.001 | | batch size | 8 | | BCE weight | 0.5 | | Dice weight | 0.5 | Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048). **This is the only task with a `VERIFIED` headline metric** (pooled IoU 0.8122 on the immutable test split). ## 5. Optical-SAR fusion head | Hyperparameter | Value | |---|---| | input dim | 2318 = 3 × 768 + 12 + 2 | | hidden dim | 512 | | dropout | 0.2 | | num classes | 19 (BigEarthNet CLC) | **Seed sweep.** Training was run as two arms (**armA**, **armB**) × five seeds (**100–104**), with a per-arm variance report (`armA_seed_variance_report.json`). The **production head** is a distinct, frozen artifact (`fusion_head_production_v001/head.pt`) with its own `production_head_record.json` and a `phase12_rerun_verification.json`. **Pre-registration.** The headline metric is a **pre-registered** 115-class protocol (`pre_registered_11.5`) computed over the 19-class label space on the held-out test split. The metric JSON records `is_deciding_statistic: False`, i.e. it is a reported measurement, not a decision statistic. **Feature caches.** Training consumes cached CROMA features (`fusion_features/`, `fusion_features_armB/`, ~231 MB each). These caches are **reproducible** and are not released as model weights. ## 6. Change-VQA head — trained externally The change-VQA head was trained **outside this repository**, on an external GPU (Kaggle), following `docs/R02_KAGGLE_TRAINING_GUIDE.md`. That guide's status on entry was `IMPLEMENTATION_READY_FOR_EXTERNAL_TRAINING`, and its explicit contract is: > **Training produces an artifact, not a verified capability, and the run record says > `TRAINED_UNVERIFIED`.** The returned checkpoint was **promoted** through a byte-identity gate (`artifacts/change_vqa/run/PROMOTION.json`): | Property | Value | |---|---| | sha256 | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` | | bytes | 5,822,809 | | architecture | `change_vqa_head_v1` | | parameters | 1,453,912 | | non-finite tensors | **0** | | weights modified during promotion | **false** | | byte-identical to source | **true** | | hash agrees across | `model_metadata.json`, `run_record.json`, `hashes.json` | **Selection:** epoch **8**, chosen on **Val answer accuracy = 0.700018**, stopped by early stopping. Seed 42. It trains on **cached change + text features** (specs `change_feat_v1`, `change_cache_spec c801326f85a185f8`, `text_cache_spec d2801ea1a314354a`), not on raw imagery. > **Note.** The raw CDVQA loader (see [`DATASETS.md`](DATASETS.md)) loads examples but has **no > training loop**; the shipped head is a cached-feature model. These are different paths and are not > conflated. ## 7. VLM LoRA adapter — trained externally A PEFT LoRA adapter on **frozen** `HuggingFaceTB/SmolVLM-500M-Instruct`. | Hyperparameter | Value | |---|---| | PEFT version | **0.19.1** | | `r` (rank) | 16 | | `alpha` | 32 | | `dropout` | 0.05 | | target modules | `model.text_model.*.{q,k,v,o,gate,up,down}_proj` | | precision | **fp16** (finding C-6: T4 is SM 7.5 → **fp16, NOT bf16**) | | batch size | 2 | | gradient accumulation | 8 | | learning rate | 0.0002 | | epochs | 1 | | gradient checkpointing | true | | save every steps | 500 | **Finding F5-2 (cost).** The processor's default `longest_edge` is 2048, which upscales 512-px tiles 4× and then splits them into **17 sub-images** (`pixel_values (1, 17, 3, 512, 512)`, 1142 prompt tokens). Pinning `processor_longest_edge: 512` yields `pixel_values (1, 1, 3, 512, 512)`. The plan estimated a 4× cost overrun; the **measured** figure is ~17×. **Finding F5-3.** SmolVLM requires one `` token per image in the prompt; hand-written prompt strings raise `ValueError`. Prompts are always built through `processor.apply_chat_template()`. **Outcome:** metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is **ACCEPTANCE-REJECTED** for promotion. The deployed caption/VQA path uses the **unadapted** model. See [`MODELS.md`](MODELS.md) §3.6. ## 8. Reproducibility contract for training - **Seed 42** everywhere (`project.seed`). - **Precision `fp16`** (T4 constraint), never bf16. - Every artifact records the **frozen config hash** `78f1e3700da15aa1`; a config edit moves the hash and invalidates the artifact. - `save_every_steps: 500`; checkpoints are archived as provenance, not released. - Training guides state their own entry status and **never** claim a trained artifact is a verified capability. ## 9. What was NOT trained | Item | State | |---|---| | Backbone fine-tuning (any) | **NOT DONE** — all backbones frozen | | Router on the test split | **NOT RUN** | | Any end-to-end / joint training | **NOT RUN** | | Re-training of the change head at a second resolution | **NOT RUN** | | Benchmark adapters | **NOT RUN** |