SatQuery / docs /TRAINING.md
thundercode's picture
release: add docs/TRAINING.md
4e54b93 verified
|
Raw History Blame
8.37 kB

Training

Status tags: IMPLEMENTED · VERIFIED · MEASURED · ATTEMPTED · NOT RUN · REJECTED.

The project trains small modules on frozen backbones. No backbone is fine-tuned end-to-end. Every hyperparameter lives in configs/base.yaml (no magic numbers in Python), and every trained artifact records the frozen config hash it was trained against.


1. Overview

Artifact Backbone (frozen) Where trained Selection signal
router adapter all-MiniLM-L6-v2 local CPU val (group-split)
grounding head RemoteCLIP ViT-B/32 local val fraction 0.10
change head STANet ResNet-18 + PAM local val (LEVIR split)
optical_sar fusion head CROMA-base local, seed sweep held-out test (pre-registered)
change_vqa head over cached change features external GPU (Kaggle) val answer accuracy
vlm LoRA adapter SmolVLM-500M-Instruct external GPU frozen 1000-question subset

CPU-first. The router and every head except the VLM adapter train on CPU. The VLM LoRA adapter requires a GPU (T4-class).

2. Router adapter

A 50,822-parameter adapter over the frozen MiniLM encoder.

Hyperparameter Value
epochs 60
batch size 64
learning rate 0.001
weight decay 0.01
task loss weight 1.0
modality loss weight 0.3
binary loss weight 0.5
val ratio 0.15
hard negatives to test true

Finding F4-2 — the encoder is frozen, so embeddings are cached and the adapter trains on cached vectors. Measured: 20 epochs / 4,096 vectors in 0.28 s on CPU. No GPU is required.

Finding F4-3 — splits are by GROUP (template / hard-negative family), never by example. Hard-negative families are placed in the test split so their accuracy measures generalisation rather than memorisation.

Status: the adapter is trained and shipped. Its measured number (0.965116) is validation-only, ungated, n = 86; the router test split was NOT RUN.

3. Grounding head

A trainable head over the frozen RemoteCLIP ViT-B/32 encoder.

Hyperparameter Value
learning rate 0.0001
batch size 16
epochs 20
weight decay 0.0001
warmup ratio 0.05
grad clip 1.0
val fraction 0.10
save every steps 500
box loss weight 0.5
GIoU loss weight 0.3
confidence loss weight 0.2
positive_confidence_weight 20.0

Architecture constraint (enforced, not documented). Per-cell feature is concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048. core/config.py rejects any value other than 4 × grounding.encoder_projected_dim at load time, and the specialist asserts the same 512 against the real model — because a mismatch is a silent shape error that torch only raises at the similarity step, after patch features are already cached.

Why positive_confidence_weight = 20.0. Objectness BCE sees ~1 positive cell out of 49. Unweighted, the optimum is "no object" everywhere; the weight is what stops that collapse.

Resolution is frozen at 224. 448 was evaluated and REJECTED (paired test: mean diff −0.0147, 95 % CI [−0.0160, −0.0134], t = −22.63, at 1.59× latency).

4. Change head

STANet-style Siamese detector.

Hyperparameter Value
encoder ResNet-18
self-attention PAM (BAM alternative not used)
tile size 256
tile overlap 0
threshold 0.50
min component pixels 32
learning rate 0.001
batch size 8
BCE weight 0.5
Dice weight 0.5

Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048). This is the only task with a VERIFIED headline metric (pooled IoU 0.8122 on the immutable test split).

5. Optical-SAR fusion head

Hyperparameter Value
input dim 2318 = 3 × 768 + 12 + 2
hidden dim 512
dropout 0.2
num classes 19 (BigEarthNet CLC)

Seed sweep. Training was run as two arms (armA, armB) × five seeds (100–104), with a per-arm variance report (armA_seed_variance_report.json). The production head is a distinct, frozen artifact (fusion_head_production_v001/head.pt) with its own production_head_record.json and a phase12_rerun_verification.json.

Pre-registration. The headline metric is a pre-registered 115-class protocol (pre_registered_11.5) computed over the 19-class label space on the held-out test split. The metric JSON records is_deciding_statistic: False, i.e. it is a reported measurement, not a decision statistic.

Feature caches. Training consumes cached CROMA features (fusion_features/, fusion_features_armB/, ~231 MB each). These caches are reproducible and are not released as model weights.

6. Change-VQA head — trained externally

The change-VQA head was trained outside this repository, on an external GPU (Kaggle), following docs/R02_KAGGLE_TRAINING_GUIDE.md. That guide's status on entry was IMPLEMENTATION_READY_FOR_EXTERNAL_TRAINING, and its explicit contract is:

Training produces an artifact, not a verified capability, and the run record says TRAINED_UNVERIFIED.

The returned checkpoint was promoted through a byte-identity gate (artifacts/change_vqa/run/PROMOTION.json):

Property Value
sha256 cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a
bytes 5,822,809
architecture change_vqa_head_v1
parameters 1,453,912
non-finite tensors 0
weights modified during promotion false
byte-identical to source true
hash agrees across model_metadata.json, run_record.json, hashes.json

Selection: epoch 8, chosen on Val answer accuracy = 0.700018, stopped by early stopping. Seed 42. It trains on cached change + text features (specs change_feat_v1, change_cache_spec c801326f85a185f8, text_cache_spec d2801ea1a314354a), not on raw imagery.

Note. The raw CDVQA loader (see DATASETS.md) loads examples but has no training loop; the shipped head is a cached-feature model. These are different paths and are not conflated.

7. VLM LoRA adapter — trained externally

A PEFT LoRA adapter on frozen HuggingFaceTB/SmolVLM-500M-Instruct.

Hyperparameter Value
PEFT version 0.19.1
r (rank) 16
alpha 32
dropout 0.05
target modules model.text_model.*.{q,k,v,o,gate,up,down}_proj
precision fp16 (finding C-6: T4 is SM 7.5 → fp16, NOT bf16)
batch size 2
gradient accumulation 8
learning rate 0.0002
epochs 1
gradient checkpointing true
save every steps 500

Finding F5-2 (cost). The processor's default longest_edge is 2048, which upscales 512-px tiles 4× and then splits them into 17 sub-images (pixel_values (1, 17, 3, 512, 512), 1142 prompt tokens). Pinning processor_longest_edge: 512 yields pixel_values (1, 1, 3, 512, 512). The plan estimated a 4× cost overrun; the measured figure is ~17×.

Finding F5-3. SmolVLM requires one <image> token per image in the prompt; hand-written prompt strings raise ValueError. Prompts are always built through processor.apply_chat_template().

Outcome: metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is ACCEPTANCE-REJECTED for promotion. The deployed caption/VQA path uses the unadapted model. See MODELS.md §3.6.

8. Reproducibility contract for training

  • Seed 42 everywhere (project.seed).
  • Precision fp16 (T4 constraint), never bf16.
  • Every artifact records the frozen config hash 78f1e3700da15aa1; a config edit moves the hash and invalidates the artifact.
  • save_every_steps: 500; checkpoints are archived as provenance, not released.
  • Training guides state their own entry status and never claim a trained artifact is a verified capability.

9. What was NOT trained

Item State
Backbone fine-tuning (any) NOT DONE — all backbones frozen
Router on the test split NOT RUN
Any end-to-end / joint training NOT RUN
Re-training of the change head at a second resolution NOT RUN
Benchmark adapters NOT RUN