SatQuery / docs /DATA_PIPELINE.md
thundercode's picture
release: add docs/DATA_PIPELINE.md
5a57728 verified
|
Raw
History Blame Contribute Delete
66.3 kB

Data pipeline β€” deep reference

Status tags used on every substantive claim: IMPLEMENTED Β· VERIFIED Β· MEASURED Β· ATTEMPTED Β· NOT RUN Β· BLOCKED Β· DEFERRED Β· REJECTED Β· OPEN Β· RESOLVED Β· CLOSED.

Every statement in this document was produced by reading a file in C:/Users/anish/satquery-ai/. Where a fact is not established from the available evidence, this document writes UNKNOWN β€” not established from the available evidence. Where a stage is declared in configuration but read by no code, that is stated explicitly rather than presented as a working step. Where a cache spec hash is cited, it was recomputed from the code that produces it, not copied from a report.

This chapter is the reference for the end-to-end data path: from an uploaded asset to the tensor a specialist consumes, including every decode, every normalisation, every cache, and every intermediate artefact. It is written to be sufficient to reconstruct the pipeline from the document alone.

Sibling documents: GEOSPATIAL.md (the raster contract, CRS, coordinate systems, the sensor adapter, the tiling policy and the 224 decision), architecture/03-request-lifecycle.md (the nine controller states), architecture/05-specialists.md (the six specialists), architecture/07-configuration-freeze.md (the frozen config identity), DATASETS.md (the corpora), REPRODUCIBILITY.md (reproduction recipes), TRAINING.md (how each corpus is consumed).


Table of contents

  1. How to read this document
  2. The end-to-end path
  3. Validation and decoding
  4. Modality inference, and where it is consumed
  5. The normalisation stages
  6. Tiling and tile selection
  7. The change-VQA data path
  8. The optical-SAR data path
  9. The router data path
  10. The grounding data path
  11. The VQA / caption data path
  12. Determinism and reproducibility
  13. Where intermediate artefacts live
  14. What is NOT part of the pipeline
  15. What is NOT RUN / OPEN / BLOCKED for this topic
  16. Where the evidence lives

0. How to read this document

The pipeline has two distinct halves, and almost every confusion about it comes from conflating them:

  1. The serving path β€” an uploaded asset becomes a specialist result, request by request. This is what runs in production. It is stateless with respect to data: it reads a raster, prepares it, and calls a model. No cache is consulted except the router's, and no intermediate artefact is required for a result to be produced.
  2. The preparation / training path β€” a corpus on disk becomes a cache of frozen features, and a head is trained over that cache. This runs offline, once per corpus revision. It is where the mean Β± 2Β·std radiometric stretch actually executes, where the change_feat_v1 feature vector is computed, and where the arm (A/B) is baked in.

The caches are the bridge: the training path produces them, and the serving path does not read them (except the router's corpus cache, which is a training-time artefact, and the change-VQA text cache, which is likewise training-time). A reader who assumes the serving path is cache-driven will misread the system; the caches exist so that a frozen encoder is not re-run across an entire corpus on every training experiment.

Three further distinctions are load-bearing and are stated at each point below:

  • Declared versus executed. configs/base.yaml declares the tiling policy, the optical percentile stretch and the SAR dB clip. All three are read by no code on the serving path.
  • Deterministic versus seeded. Some stages are pure functions of their input (the display stretch, the quality gate, the radiometric stretch). Others are stochastic and take a seeded RNG (channel dropout). The distinction determines what "reproducible" means for each.
  • Cache-hit versus cache-miss. Every cache in the project is keyed by a spec hash or a fingerprint, and a mismatch is a miss or a refusal, never a silent stale read.

1. The end-to-end path

1.1 The nine controller states

The controller runs a fixed nine-state machine, declared in configs/base.yaml under agent.states and mirrored by the ControllerState enum in core/schemas.py:79:

RECEIVE β†’ PARSE β†’ VALIDATE β†’ PLAN β†’ PREPROCESS β†’ EXECUTE β†’ AGGREGATE β†’ VERIFY β†’ RESPOND

The division of labour is stated once and holds throughout: the router understands, the policy decides, the specialists compute, the VLM explains, and the evidence proves. The three states that matter for the data pipeline are VALIDATE, PREPROCESS and EXECUTE.

1.2 The path, drawn

POST /v1/assets                          (gateway; five-type allowlist, two-layer size cap)
   β”‚
   └─ asset_id (opaque, 128-bit; never a path)
        β”‚
        β–Ό
POST /v1/analyze { assets: [...], query }
   β”‚
   β”œβ”€ RECEIVE   ── resolve handles to server-side paths
   β”‚
   β”œβ”€ PARSE     ── router: query text β†’ Intent (task, modality, temporal, spatial, language)
   β”‚
   β”œβ”€ VALIDATE  ── core/controller.py::_inspect_asset
   β”‚                 └─ preprocessing.raster.inspect_raster(path, max_pixels=..., ...)
   β”‚                      header-only: dims β†’ bands β†’ dtype β†’ CRS β†’ transform β†’ bounds
   β”‚                      β†’ nodata β†’ modality β†’ AssetMetadata
   β”‚
   β”œβ”€ PLAN      ── policy selects a workflow (which specialists, in what order)
   β”‚
   β”œβ”€ PREPROCESS ── per-specialist preparation (this chapter's core)
   β”‚                 β”œβ”€ decode: read_bands β†’ (C, H, W); load_image_array β†’ (H, W, 3) uint8
   β”‚                 β”œβ”€ sensor adapter: band map + zero-fill + availability mask
   β”‚                 β”œβ”€ normalisation: display stretch / encoder-input stretch
   β”‚                 └─ resize: to the model's canonical resolution
   β”‚
   β”œβ”€ EXECUTE   ── specialist forward pass β†’ SpecialistResult
   β”‚
   β”œβ”€ AGGREGATE ── combine specialist results
   β”œβ”€ VERIFY    ── consistency checks
   └─ RESPOND   ── ResultEnvelope { run_id, result, trace }

architecture/03-request-lifecycle.md records a fact worth repeating here: the PREPROCESS state is never recorded as a distinct TraceStep in the execution trace; the work happens inside each specialist's execute rather than as a controller-level step. So the trace shows VALIDATE and EXECUTE but not the preparation between them, and the preparation's parameters reach the trace through the specialist's evidence payloads (the availability masks, the radiometry report, the decode metadata) rather than through a PREPROCESS step.

1.3 What each state reads and writes

State Reads Writes Cache consulted
RECEIVE the uploaded asset handle a server-side path none
PARSE the query text an Intent router corpus cache (training only)
VALIDATE raster header only AssetMetadata none
PLAN the Intent, the AssetMetadata list a workflow plan none
PREPROCESS raster pixels model-ready arrays none
EXECUTE model-ready arrays a SpecialistResult none (serving)
AGGREGATE / VERIFY / RESPOND specialist results the envelope + trace none

The single most important cell in that table is VALIDATE's "raster header only": the inspection that decides whether an asset is usable does not decode it. inspect_raster's docstring states this β€” "Validate and describe a raster without loading its pixel data" β€” and it is what keeps a 25-megapixel rejection cheap.


2. Validation and decoding

2.1 Header-only validation β€” inspect_raster

Covered in full in GEOSPATIAL.md Β§1.3. The pipeline-relevant points:

  • It reads dimensions, band count, dtypes, CRS, transform, bounds, nodata, resolution, driver and a tiled flag, and constructs a GeoMetadata plus an AssetMetadata.
  • The pixel budget is enforced here: width * height > max_pixels raises OversizedImageError, which is recoverable β€” the caller may downscale. max_pixels comes from image.max_pixels (25,000,000).
  • The modality is assigned here by infer_modality (Β§3).
  • The SHA-256 is computed here only when compute_hash=True (streaming, 1 MiB chunks) β€” it is needed for manifests and dedup detection, and it is the field the change specialist uses to detect a repeated acquisition.

2.2 Band decoding β€” read_bands

preprocessing/raster.py::read_bands(path, *, indexes=None, out_dtype=None) returns (bands, H, W) plus a profile that retains crs, transform and bounds. The profile is the source of georeferencing for every downstream write (change maps, masks) so "downstream code never loses georeferencing." It accepts an optional indexes list and an optional out_dtype cast, and wraps every failure as RasterReadError.

Consumers:

  • preprocessing/imagery.py::load_image_array β€” for the VQA, caption and grounding paths.
  • specialists/optical_sar/specialist.py::execute β€” reads the optical and SAR arrays raw, then hands them to the sensor adapter.
  • specialists/change/specialist.py::_write_artifacts β€” reads T1's profile so the change map is written with T1's georeferencing.

2.3 Display decoding β€” load_image_array

preprocessing/imagery.py::load_image_array turns a raster into a displayable (H, W, 3) uint8 array:

  1. read_bands β†’ (bands, H, W).
  2. Transpose to (H, W, bands); take the first three bands as RGB; repeat a single band three times; or stack a 2-D array three times.
  3. Per-scene 2/98 percentile stretch over finite values, clip to [0, 1], scale to uint8.

Its docstring is explicit about what it does not do: "It does not resample, crop, or reproject. Those change the pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using the ORIGINAL raster's dimensions β€” a silent resize here would put every box in the wrong place." And about determinism: "The percentile stretch is deterministic: the same file always yields the same array, so a grounding box and a VQA answer describe identical pixels."

load_pil_image wraps the same array in a PIL.Image, and notes that georeferencing is not lost because the caller's AssetMetadata already carries it.

2.4 The deterministic quality gate

preprocessing/quality.py is "a deterministic input-quality gate" that sits upstream of the model. It exists because of a measured failure: a loaded SmolVLM-500M-Instruct, given uniform random noise and a prompt that explicitly said "answer only from what is visible", produced "a fluent, specific, entirely fabricated scene description." The gate's rationale is stated directly: "A 500M-parameter VLM will describe something for any input it is given, and asking it to self-assess reliably is asking it to do the thing it just failed at."

The discriminating signal is lag-1 spatial autocorrelation, not variance. The measured separation:

Input Autocorrelation
Real remote-sensing imagery 0.6 – 0.99
Uniform random noise ~0.00
A constant (blank) image undefined; variance ~0

Variance alone cannot separate a flat desert scene from an all-zero tile (both have near-zero variance); autocorrelation is high for both uniform-but-real imagery and textured imagery, and collapses only for noise.

The constants:

Constant Value Meaning
MIN_AUTOCORRELATION 0.10 Below this, the array is not spatially coherent
FLAT_STD_EPSILON 1e-6 Below this normalised std, the array is constant
NOISE_ENTROPY_BITS 7.8 Shannon entropy above which the histogram is noise-like
BLOCKING_VERDICTS {NOISE, INVALID_VALUES} Verdicts that must block a VLM call

assess_image_quality returns an ImageQuality with one of five verdicts:

Verdict Meaning Effect
STRUCTURED Spatial structure consistent with real imagery Safe to analyse
FLAT Effectively constant Usable, but is_degraded β€” lower confidence
NOISE Uncorrelated, near-uniform Blocks the VLM call
TOO_SMALL Fewer than 16 px on a side Usable, flagged
INVALID_VALUES NaN or infinite values present Blocks the VLM call

Two behaviours are worth quoting. First, the non-finite rule is deliberately tolerance-free: "a fixed fraction is size-dependent, so one NaN in a 64x64 array (0.99976) would pass while one NaN in a 10x10 array (0.99) would fail. The same defect must not be tolerated or rejected depending on image dimensions." Any non-finite value is disqualifying. Second, the entropy check is described as a redundant second signal: "MEASURED, and the margin here is thin: structured ramp+texture gives 7.581, uniform noise gives 7.988. That is a 0.22-bit gap against a 7.8 threshold. The autocorrelation check is what actually carries this gate (0.952 vs -0.008, a 0.96 margin against a 0.10 threshold)."

The gate is a pure function β€” "Nothing here is learned, sampled, or probabilistic. Same input, same verdict." A known limitation is recorded: a float GeoTIFF whose nodata sentinel is NaN lands in INVALID_VALUES and is refused, and the comment records that "Filling nodata from the raster profile belongs in the tiling work" β€” i.e. the fix is deferred to a stage that does not exist (Β§5).

2.5 Error types

Every failure in the validation and decoding path is a typed SatQueryError subclass. The relevant ones for this pipeline:

Error Raised by Meaning
RasterReadError inspect_raster, read_bands, write_raster The file is not a readable raster
OversizedImageError inspect_raster Exceeds the pixel budget; recoverable
UnsupportedBandsError inspect_raster, the sensor adapter Band count is zero, or bands cannot be placed
MissingCRSError parse_crs A CRS string is malformed
CoordinateError geospatial/transform.py A conversion lacks the context it needs
SpecialistError specialists A forward pass or preparation failed

The design rule from preprocessing/raster.py applies throughout: "Never raise a bare exception. Every failure is a typed SatQueryError."


3. Modality inference, and where it is consumed

infer_modality(band_count, explicit=None) (preprocessing/raster.py:58) is covered in full in GEOSPATIAL.md Β§6. Its pipeline role is as the first consumer of the band count that VALIDATE read:

  • inspect_raster calls it and stores the result on AssetMetadata.modality.
  • The optical-SAR specialist's _assess_pair reads the declared modality first and falls back to infer_modality(asset.geo.band_count) only when it is UNKNOWN, appending a warning that names the heuristic.
  • The change specialist does not infer modality: it requires exactly two assets and compares their CRS and registration, without asserting a modality for either.

The heuristic's limits are stated in its own comment: "These are heuristics, not ground truth β€” the sensor adapter is authoritative when a sensor descriptor is supplied." A band count is not a sensor declaration, and the adapter is the only component that can map a band to a canonical channel.


4. The normalisation stages

There are three distinct normalisation transforms in this codebase, plus one declared but unimplemented stage. Conflating them is the most common misreading, so they are laid out side by side.

4.1 The display stretch β€” per-scene 2/98

Where: preprocessing/imagery.py::load_image_array (and its grayscale cousin in specialists/change/postprocess.py::to_grayscale_float).

What: one (lo, hi) computed over all finite values across all bands of one scene, then (arr - lo) / (hi - lo), clipped to [0, 1].

Output: uint8 (H, W, 3) for display (or float32 grayscale for registration).

Purpose: presentation β€” so a VQA answer and a grounding box describe identical pixels.

Determinism: pure; same file β†’ same array.

Not the CROMA transform. docs/PHASE14_CROMA_NORMALISATION_CHANGE.md Β§1.1 names three disqualifying differences: per-scene rather than per-channel; uint8 3-band rather than float 12-channel; presentation rather than radiometry. Reusing it for CROMA is explicitly forbidden by that ruling (Β§3.3).

4.2 The encoder-input stretch β€” per-channel mean Β± 2Β·std

Where: specialists/optical_sar/radiometry.py::normalise_for_croma.

What: for each channel c of each sample, lower = mean_c - 2Β·std_c, upper = mean_c + 2Β·std_c, then (x - lower) / (upper - lower), optionally through a uint8 round-trip (Γ—255, clip, quantise, /255). Computed per sample (not per batch) with ddof=1.

Output: float32 of the same shape, in [0, 1].

Purpose: the transform the CROMA authors' own README instructs users to apply before the frozen encoder.

Determinism: "pure and deterministic: no sampling, no learned statistics, no RNG, and no dependence on batch composition. Same input -> same output."

Where it runs: the extraction path (training/fusion/extract.py), not the serving path (Β§7.5). The extraction module applies it "because the arm's use_8_bit has to be applied at exactly one place", and it refuses an encoder that would stretch a second time.

The zero-channel rule: an all-zero channel (the sensor adapter's unavailability convention) has std = 0 and would divide by zero. Such channels are skipped and left at exactly zero, with status unavailable; a present-but-constant channel is skipped with status degenerate. This is the C-1 discipline applied to radiometry β€” "the same rule that forbids inventing a band forbids inventing a dynamic range for a band that does not exist."

4.3 The declared-but-unimplemented conditioning stage

Where declared: configs/base.yaml:

optical:
  normalization: percentile
  lower_percentile: 2
  upper_percentile: 98
sar:
  representation: db
  clip_min_db: -30
  clip_max_db: 5

Status: DECLARED (not read). No code applies a percentile/dB conditioning stage. The radiometry module's docstring states the consequence: "the two optical.* and three sar.* config keys are still read by no code. That is recorded, not fixed." docs/PHASE14_CROMA_NORMALISATION_CHANGE.md Β§6 item 3 adds: "A config key that no code reads is worse than a missing key: it reads as a satisfied requirement."

4.4 The side-by-side

Transform Granularity Output Purpose Implemented Runs on serving path
Display 2/98 per-scene, all bands uint8 (H,W,3) Presentation βœ… βœ…
Encoder-input mean Β± 2Β·std per-channel, per-sample float32 [0,1] Frozen-encoder interface βœ… ❌ (extraction only)
Change grayscale 2/98 per-scene, grayscale float32 [0,1] Registration correlation βœ… βœ…
Percentile / dB conditioning per-scene, per config unbounded rescale base.yaml's intent ❌ ❌

5. Tiling and tile selection

Covered in full in GEOSPATIAL.md Β§10. The pipeline-relevant summary:

  • image.max_pixels: 25000000 is enforced, at inspect_raster time, raising OversizedImageError.
  • image.tile_size: 512, image.tile_overlap: 128, image.max_tiles: 64, image.top_k_tiles: 4 are declared and read only by core/config.py's single relationship check (top_k_tiles <= max_tiles). No serving-path code slices a raster into tiles or selects a top-K.
  • change.tile_size: 256, change.tile_overlap: 0 are likewise declared and not consumed as a tiling step; 256 is the change model's training resolution.

Consequence for the pipeline: a raster within the pixel budget is handed to a specialist whole. The only size-reduction steps that actually run are per-specialist resizes to a model's canonical resolution (e.g. resize_to_canonical to 120Γ—120 for CROMA, and the change feature extractor's bilinear resize to 256Γ—256). Neither is tiling: both resample the whole image rather than selecting a region.

The VLM path's processor_longest_edge: 512 is a control against the processor's default 2048, not a tiling step. core/config.py enforces processor_longest_edge <= image.tile_size, with the measured rationale recorded: the default 2048 upscales a 512 tile 4Γ— and splits it into ~17 sub-images, "the real figure is ~17x" against the plan's estimated 4Γ—.


6. The change-VQA data path

The change-VQA path (Task.CHANGE_VQA, R-02) is the most elaborate in the project: it has a frozen detector, two frozen feature extractors, a trained head, and four cache artefacts. It is the clearest illustration of the serving/preparation split.

6.1 The dataset and the preprocessing version tag

training/change_vqa/dataset.py defines the identity constants:

DATASET_ID = "cdvqa"
PREPROCESSING_VERSION = "change_vqa_preproc_v1"
SPLIT_MAP: dict[str, str] = {
    "Train": "train",
    "Val": "val",
    "Test": "test",
    "Test2": "test",
}

PREPROCESSING_VERSION is "bumped when the record shape or the target definition changes. Recorded in every evaluation so two numbers computed under different definitions can never be silently compared." It is the preprocessing version tag for this path: it names the record and target definition, and is distinct from the feature spec hashes below, which name the tensor definition.

SPLIT_MAP is a leakage control, not a convenience: Test2 maps to test because Test2 is "a SECOND QUESTION SET over the SAME 968 scenes as Test, not an independent sample; it must never be pooled." docs/PHASE10_CDVQA_DATA_STATUS.md Β§4 measured this: Test and Test2 share 100% of their images (968 / 968), while Train/Val/Test are cleanly disjoint.

LABEL_PALETTE decodes the six change classes (NVG surface, trees, low vegetation, water, buildings, playgrounds) from a fixed colour palette; white (255,255,255) is a shared background, not a class.

SceneTargets derives supervision from label1/label2 with four definitions, all fractions of total pixels: class_mag[c] (magnitude), class_delta[c] (signed), class_ratio[c] (per-class ratio) and total_changed. The total_changed definition was validated against gold answers: over 40 real change_ratio rows, count(label1 != label2) / total reproduced the annotated bin 34/40 = 85% of the time, versus 1/40 for the competing definition. The residual 15% is recorded as "annotator/map disagreement, which is the noise floor of the target" rather than smoothed away.

6.2 The change feature extractor

training/change_vqa/features.py defines the frozen change feature. The module explains why the features are frozen: the plan budgets CDVQA at ≀ 3 h on T4Γ—2, and training end-to-end through STANet would re-learn a representation the project already has. Freezing it drops the trainable parameter count from ~15.8 M to ~1.5 M.

CHANGE_FEATURE_DIM = 1045, and it is three things:

Part Dim Composition Answers
level_pooled 1024 4 levels Γ— 128 ch Γ— {mean, max} "what does the change look like locally?"
change_stats 5 mean, std, frac>0.3, frac>0.5, frac>0.7 of the change map "how much of the scene changed?"
change_grid 16 adaptive 4Γ—4 pool of the same map "WHERE did it change?"

The dimensions are derived, never typed β€” "a literal here would drift silently from the pooling code below and produce a shape error three files away":

N_LEVELS = 4
LEVEL_CHANNELS = 128
LEVEL_POOLED_DIM = N_LEVELS * LEVEL_CHANNELS * 2      # 1024
CHANGE_STATS_DIM = 2 + len(CHANGE_STAT_THRESHOLDS)    # 5
CHANGE_GRID_DIM = CHANGE_GRID * CHANGE_GRID           # 16
CHANGE_FEATURE_DIM = LEVEL_POOLED_DIM + CHANGE_STATS_DIM + CHANGE_GRID_DIM  # 1045

The spatial parts are load-bearing: "Without change_grid the head would see a scene-global average and could not distinguish 'buildings changed in one corner' from 'buildings changed everywhere', which is precisely the difference between largest_change and smallest_change."

The default working resolution is DEFAULT_IMAGE_SIZE = 256 β€” "the change model's OWN training resolution (configs/base.yaml: change.tile_size: 256; LEVIR-CD patches are 256x256). Feeding the native 512 would push the frozen encoder outside the distribution it was trained on." 512 is offered as an explicit, recorded alternative.

ChangeFeatureExtractor.extract(t1, t2) re-runs the detector's own submodules to capture the intermediate fused levels that STANetStyleChangeDetector.forward does not return. That is "a second implementation of one forward pass, so it is a drift hazard β€” and it is closed by a test that asserts the probability map produced here is bit-identical to STANetStyleChangeDetector.forward(...).probabilities."

6.3 The text feature extractor

TextFeatureExtractor wraps sentence-transformers/all-MiniLM-L6-v2 β€” the same model the router uses, so "the text side of the reasoning head adds no new model to the deployment." It produces TEXT_FEATURE_DIM = 384 normalised vectors and raises FeatureExtractionError if the encoder's output width disagrees with the declared constant.

6.4 The cache spec hashes

Two independent hash functions key the two caches. Both are sha256(...)[:16].

feature_spec_hash(image_size, *, extractor) hashes the feature definition: {spec, image_size, extractor, levels, level_channels, thresholds, grid, text_encoder}, sorted keys.

text_spec_hash(*, encoder, revision=None) hashes the question-feature definition: {spec: "change_vqa_text_v1", encoder, revision, dim, normalised}, sorted keys. The separate hash is deliberate: "One combined hash would force a full re-extraction of both whenever either changed, and β€” worse β€” would let a stale text cache pass a check it should fail."

Recomputed values (this document computed them from the functions above, with the constants as read from the source):

Input Spec hash
feature_spec_hash(256, extractor="stanet:trained") c801326f85a185f8
feature_spec_hash(256, extractor="stanet:untrained") 714efac5b0e6a6d1
text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2", revision="1110a243fdf4") d2801ea1a314354a
text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2") (unpinned) f87f5226c4c4a69b

So the change-feature cache spec for the trained detector at 256 px is c801326f85a185f8, and the question-feature cache spec with the pinned MiniLM revision (configs/base.yaml: router.revision: 1110a243fdf4) is d2801ea1a314354a. The spec hash encodes trained-vs-untrained, which is how the serving specialist detects a detector-state mismatch (Β§6.6).

6.5 Cache layout, resume, and the artefacts

scripts/prepare_change_vqa.py is the producer. Its default output directory is DEFAULT_OUT = REPO_ROOT / "artifacts" / "change_vqa", and it writes:

Artefact Shape Content
<out>/scene_manifest.jsonl one record per (native split, scene) identity, integrity, split, question counts
<out>/scene_targets.jsonl one record per scene label-derived SceneTargets
<out>/change_features.npz (n_scenes, 1045) + spec hash + extractor config frozen STANet features, keyed by scene_key
<out>/text_features_<split>.npz (n_questions, 384) + spec hash question features, keyed by question_id

The manifest is scene-level, deliberately: "153,130 question rows would produce a ~50 MB JSONL that duplicates text already authoritative in annotations/*.json. The manifest therefore records the 2,968 scenes ... and the question-level records are derived from the annotations on demand."

Resume is spec-guarded. ChangeFeatureCache.read(path, *, expect_spec_hash=...) refuses a cache built under a different spec:

if expect_spec_hash is not None and cache.spec_hash != expect_spec_hash:
    raise FeatureExtractionError(
        f"change feature cache at {p} was built under spec "
        f"{cache.spec_hash!r} but the caller expects {expect_spec_hash!r}; "
        f"re-extract rather than training on a representation that does not match serving"
    )

TextFeatureCache.read applies the same guard. The producer calls ChangeFeatureCache.read(cache_path, expect_spec_hash=extractor.spec_hash) when resuming, so a resolution change or a detector-state change invalidates the cache instead of silently mixing two representations.

A scene whose imagery cannot be read is skipped and reported through on_progress, never silently dropped: "a scene missing from the cache is a training sample the model never sees, and the caller must be able to count how many that was."

6.6 The serving path, and the spec-mismatch refusal

ChangeVQASpecialist (specialists/change/vqa_specialist.py) assembles three pieces: a trained head, a change feature extractor, and a question text encoder. Its execute:

  1. Validates exactly two assets, both present.
  2. Calls unavailable_reason(); if anything is missing or mismatched, returns _unavailable_result with answer="" and degraded=True.
  3. Otherwise loads the two images, extracts (vector, facts) and the text vector, resolves the question type and temporal reference, and calls predict_answers.

Why degraded produces no answer. The module docstring is explicit: "a random change map is visibly noise, whereas an untrained 19-way classifier still emits a fluent, confident-looking yes. A user cannot tell the second from a real answer, so the untrained case returns no answer at all." The SpecialistResult validator independently marks an empty change-VQA answer as degraded.

The spec-mismatch refusal. feature_spec_mismatch() compares the head's trained spec (head_metadata["change_cache_spec"]) with the serving extractor's spec_hash. A mismatch is a refusal, not a warning: "A head is only meaningful for the feature distribution it was fitted on." The measured reason this matters: configs/base.yaml's change: section carries no checkpoint_path key, so the registry builds the change feature extractor with checkpoint_path=None and gets an untrained STANet β€” while scripts/prepare_change_vqa.py defaults to the trained LEVIR checkpoint. Training and serving would therefore consume different representations, and "the head would still emit a fluent yes." The spec hash covers this because it encodes trained-vs-untrained, so the refusal catches the detector-state mismatch as well as a resolution mismatch.

Other serving constants: LOW_CONFIDENCE_THRESHOLD = 0.40 (a reporting threshold, not a calibration) and TEMPORAL_ORDER_NOTE (T1=pre, T2=post; label1=pre/label2=post proven at agreement 1.0000 over 2,968 scenes; im1=pre/im2=post supported statistically, not proven).

A config fact worth recording. The specialist reads change_vqa.image_size (default 256), change_vqa.apply_type_mask (default True) and change_vqa.head_path via config.get(...), but configs/base.yaml has no change_vqa block at all. Those keys therefore always take their code defaults, and the frozen config does not name them. This is the same pattern as the grounding head's DEFAULT_HEAD_PATH (a code constant rather than a config key, to avoid moving Config.hash).


7. The optical-SAR data path

7.1 The serving path, step by step

OpticalSarSpecialist.execute (specialists/optical_sar/specialist.py:320):

  1. Validate. Exactly two assets; _assess_pair establishes one optical and one SAR (Β§9.2 of GEOSPATIAL.md).
  2. Descriptors. _descriptor_for returns the asset's SensorDescriptor if present, else builds one from the band count β€” build_optical_adapter for optical, positional_fallback_descriptor for SAR β€” and appends a warning naming the assumption.
  3. Read pixels. read_bands(optical_asset.path) and read_bands(sar_asset.path).
  4. Pipeline. run_pipeline(...) (below).
  5. Confidence. _confidence_components computes the four plan-section-26 components.
  6. Answer + evidence. _compose_answer builds prose only from computed facts; _build_evidence emits the availability masks, the view items, the fused-representation item and the decision.

7.2 run_pipeline

specialists/optical_sar/inference.py::run_pipeline is the seam between "an image on disk" and "three GAP vectors plus two masks":

optical raster -> sensor adapter -> (12, H, W) + mask[12]
SAR raster     -> sensor adapter -> ( 2, H, W) + mask[2]
resize to CROMA's 120x120
CROMA(SAR_images=..., optical_images=...)     <- no mask, ever
assemble_fusion_input                         <- mask enters HERE
fusion head -> logits -> probabilities

Two properties are load-bearing:

  • Degradation is a first-class outcome. run_pipeline "NEVER raises for a missing model: the degradation is reported, because 'we have no trained head' is a fact about the deployment, not a failure of the request." It raises only for a caller error β€” arrays that cannot be canonicalised at all.
  • The sensor side always runs. "The masks, the canonical channel placement, the availability statistics and the modality validation are all real work that requires no weights, and they are exactly the facts that tell an operator why a result is or is not trustworthy on Cartosat-2S + RISAT."

resize_to_canonical(array, resolution) bilinearly resizes a (C, H, W) array to (C, resolution, resolution). Its docstring notes that nearest-neighbour would be wrong for continuous optical reflectance channels, while for zero-filled channels "interpolating zeros gives zeros." It uses cv2.INTER_LINEAR when OpenCV is available, falling back to torch.nn.functional.interpolate.

7.3 The mask path

build_masks converts the two SensorAdapterOutput masks to (1, C) float32, ready for concatenation. mask_availability_stats computes the fractions and counts that feed both the confidence components and the evidence. The mask never reaches CROMA (finding C-1) β€” the specialist's evidence item records "croma_received_mask": false and "mask_consumed_by": "fusion_head".

7.4 The fusion cache β€” fusion_features

training/fusion/train.py and training/fusion/extract.py implement the Phase 12 frozen-feature trainer and its missing producer. The pipeline:

paired patches -> normalise -> resize -> CROMA (frozen) -> FusionFeature
               -> write_feature_cache -> train_fusion_head

The cache constants:

Constant Value Meaning
CACHE_VERSION "v1" Bumped when the encoder, resolution or stored fields change
DEFAULT_FEATURE_CACHE_DIR artifacts/optical_sar/fusion_features Under artifacts/, scoped to this specialist, not under the frozen artifacts/change/ tree

feature_cache_path(cache_dir=None, tag="") returns base / f"croma_features_{CACHE_VERSION}{suffix}.npz", where the suffix is _<tag> when a tag is given (e.g. a split name). The Phase-12 CLI (scripts/extract_fusion_features.py) instead takes an explicit --out-cache, documented in its own usage as artifacts/optical_sar/fusion_features/train.npz, and writes a .npz plus a .json sidecar. So both naming conventions are in play: the training loop's default is croma_features_v1_<tag>.npz, and the CLI's documented example is <split>.npz. Both are artifacts/optical_sar/fusion_features/ files, and verify_feature_cache checks row_counts_agree and sidecar_n_matches_rows.

The write is atomic-ish. write_feature_cache writes both files to unique temporary names and renames them into place with os.replace, so "no reader ever sees a half-written .npz." The two renames cannot be one operation; a crash in the window leaves a complete-but-stale pair, which verify_feature_cache reports rather than letting a truncated array load cleanly and be silently wrong. A crashed run may leave a .<name>.tmp-<pid>-<uuid>.npz behind, and the function deliberately never deletes it β€” "a delete is not a safe operation to perform on someone else's filesystem."

The multi-label collapse is an explicit policy. BigEarthNet v2.0 is multi-label; the head is a single-label 19-class softmax. The collapse is a parameter, label_policy: Sequence[str] -> int | None, and with no policy: exactly one label β†’ its index; more than one β†’ AmbiguousLabelError; zero β†’ None β†’ the sample is skipped and counted, never mapped to class 0 ("'we do not know' and 'arable land' are different facts and must not be merged").

Resume is provenance-guarded. Re-running with an existing cache loads it, skips the sample_ids already present and appends the rest (existing + new, never a replacement). But "appending rows from a different arm, resolution or config would produce a cache whose metadata describes only half its rows" β€” so check_resume_provenance refuses that and names the field and both values; resume=False rebuilds.

The dry run makes the same decision, not a similar one. plan_extraction calls the same select_samples the run calls, so plan.n_selected is result.n_encoded and plan.n_would_write is result.n_total.

Exit codes. scripts/extract_fusion_features.py returns 0 on success (including a dry run), 2 for a pre-flight or input error (nothing was encoded β€” a missing corpus, a leaked split, an unwritable --out-cache, a missing CROMA checkpoint, an ambiguous label under require_single_label), and 3 for a run that started and then failed with a typed error (e.g. a cached feature that is not 2318 wide). The comment on the FusionTrainingError branch states the distinction: "the run started, so this is 3, not a pre-flight 2."

7.5 The normalisation arm β€” A / B

The Phase 12 trainer selects the normalisation arm through the existing hash-exempt environment channel SATQUERY_CROMA_USE_8_BIT, not by editing configs/base.yaml:

ARM_CONTROL = "A"
ARM_VARIANT = "B"

ARMS: dict[str, Arm] = {
    ARM_CONTROL: Arm(ARM_CONTROL, True,  "... use_8_bit=true ..."),
    ARM_VARIANT: Arm(ARM_VARIANT, False, "... use_8_bit=false ..."),
}

The reason for the env channel is stated: "Adding a fusion_training: block would move the frozen Config.hash off 78f1e3700da15aa1 and detach the Phase 9 benchmark from its config." apply_arm writes the value and returns it so a caller can record it.

A warning that must be repeated, not summarised. The B in this dict is NOT the arm B registered in docs/PHASE14_CROMA_NORMALISATION_CHANGE.md Β§4. Both arms run the same per-channel mean Β± 2Β·std encoder-input stretch; they differ only on whether the result then takes the uint8 round-trip. The axis is the uint8 axis, not the registered "stretch vs no-stretch" axis. The originally preregistered arm B ("percentile/dB conditioning only, no encoder-input stretch") is NOT CONSTRUCTIBLE from this repository: stage 1 is unimplemented, and Arm exposes no field that can express "skip the encoder-input stretch" β€” the stretch in training/fusion/extract.py is applied unconditionally. There is no arm C.

Blocker recorded, not crossed. The arm is baked into the cached features, and artifacts/optical_sar/fusion_features/ holds only the Arm-A cache (train.npz and train.json; val.npz and test.npz were pending as of the Phase-14 correction notes). Arm B therefore needs its own feature cache before it can be trained.

7.6 The extraction module's double-stretch guard

training/fusion/extract.py applies the DEV-2 stretch itself, and refuses an encoder that would apply it again:

CROMAEncoder applies the same stretch inside encode when normalize_input=True (its default). Passing such an encoder here would stretch twice β€” a silent distribution shift that produces a perfectly well-shaped cache full of wrong numbers. _assert_raw_encoder refuses it by name rather than trusting the caller; build_extraction_encoder constructs the raw encoder this module expects.

The CLI prints normalize_input : <value> (must be False), so the guard is visible in the run output as well as enforced in code.


8. The router data path

8.1 The frozen encoder

router/encoder.py::FrozenEncoder wraps SentenceTransformer and holds it frozen. The verified contract (Phase 4 probe, 2026-09-16, sentence-transformers 6.0.1) is recorded at the top of the module:

SentenceTransformer(
    model_name_or_path='sentence-transformers/all-MiniLM-L6-v2',
    revision='1110a243fdf4',      # pinned, verified reachable
    device='cpu',
)
.get_sentence_embedding_dimension()  -> 384
.tokenizer.model_max_length          -> 256     (finding F4-1)
.max_seq_length                      -> 256
params                               -> 22,713,216
encode(64 queries, CPU)              -> 0.118 s

Two constants are declared: VERIFIED_TOKENIZER_MAX_LENGTH = 256 and VERIFIED_EMBEDDING_DIM = 384. The constructor refuses a max_length above 256 β€” "truncation would be a silent no-op" β€” and sets self._model.max_seq_length = max_length so the configured truncation (128, from router.max_length) is "enforced rather than merely documented." The encoder returns detached numpy arrays with no autograd history, which is what makes cached-embedding training possible (finding F4-2).

8.2 The corpus embedding cache

router/train.py is the consumer. CORPUS_CACHE_VERSION = "v1", and _corpus_fingerprint(corpus, encoder_id, max_length) hashes:

payload = {
    "version": CORPUS_CACHE_VERSION,
    "encoder": encoder_id,          # f"{encoder.model_name}@{encoder.revision}"
    "max_length": max_length,
    "texts": sorted(e.text for e in corpus),
}
return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()

embed_corpus_cached writes two files into the cache directory:

File Content
embeddings_<fingerprint[:16]>.npy the (n, 384) float32 embedding matrix
embeddings_<fingerprint[:16]>.json {fingerprint, encoder, max_length, count, dim, seconds}

The cache key includes the encoder id and the corpus text set, so "changing either invalidates it. A stale cache silently training on the wrong vectors would be worse than no cache." A cache hit requires both cached.shape[0] == len(corpus) and meta.get("fingerprint") == fingerprint. A bad cache is a miss, not an error: except Exception: pass falls through to re-encoding.

This cache is a training-time artefact. The serving path does not read it. embed_corpus_cached is defined and used only in router/train.py (defined at :232, used at :518), while the serving path calls self.router.route(request.query) (core/controller.py:476), which encodes the query live β€” a single short string, measured at 0.118 s for 64 queries on CPU, so one query is sub-millisecond.

8.3 The router's training discipline

The module records two findings that shape the data path:

  • F4-2 (frozen encoder β‡’ cached training). "The encoder is FROZEN. Embeddings are therefore a pure function of the query text. So we embed the whole corpus ONCE, cache the vectors to disk, and train the 50,822-parameter adapter on the cached matrix." Router training needs no GPU.
  • F4-3 (splits by group, never by example). "Splitting by example would put 'show me the water body' in train and 'show me the road' in val β€” same template, one token apart β€” and report a fake accuracy." The train function refuses to proceed if the split report is not clean.

The router's reported accuracy is 0.965116 on validation, ungated, n = 86; the test split was NOT RUN. This is a data-pipeline fact: the test split exists and is unused.


9. The grounding data path

9.1 The path

GroundingSpecialist.execute:

  1. Validate exactly one image and a non-empty phrase; check the path exists.
  2. load_pil_image(asset.path) β€” the deterministic display decode (Β§2.3).
  3. encoder.encode_image(image) and encoder.encode_text([query]) β€” RemoteCLIP ViT-B/32 at 224 px, producing (49, 512) patch tokens and a 512-d text embedding.
  4. Decode: _decode_with_head when a trained head is loaded, else _decode_zero_shot.
  5. Confidence from measurable objectness statistics.
  6. Evidence: BOUNDING_BOX items in NORMALIZED_0_1, a GEOLOCATION item in GEO when the CRS allows, a STATISTIC item, and a STATISTIC record of any dropped degenerate candidates.

9.2 There is no feature cache on the grounding serving path

Unlike the change-VQA and optical-SAR paths, grounding has no serving-time feature cache. The encoder runs live on each request. The Phase 7 resolution experiment and the Phase 8 evaluation used cached patch features (the resolution experiment's artifacts are per_sample_224_full.jsonl, per_sample_448_full.jsonl and resolution_experiment_full.json β€” 16,159 lines each), and the single decode implementation exists precisely so a cached-feature path and a live path "cannot disagree":

decode_candidates_from_features (specialists/grounding/inference.py:151) is "THE single implementation of the zero-shot decode. ground_phrase encodes an image then calls this; the evaluation script reads a feature cache then calls this." The module records the defect this closed: the Phase 8 eval originally built its own single-box baseline with argmax_candidate while Phase 7 measured through ground_phrase β€” same 16,159 records, same metric, same cached features, but a different decode:

Phase 7 via ground_phrase  : mean best IoU 0.0972
eval via argmax_candidate  : mean best IoU 0.0092

A 10Γ— gap, from a duplicated decode. "The duplication was the defect, so the duplication is what is removed: there is now exactly one place that turns similarity into boxes."

9.3 The decode

similarity_map L2-normalises both sides before the dot product, because "without that the dot product is dominated by whichever vector happens to have the larger norm, which is a property of the features rather than of the match." decode_candidates_from_features then produces a threshold box (patches within DEFAULT_DELTA = 0.02 of the peak) plus up to top_k local maxima (a patch must beat its 4-neighbours, and adjacent already-accepted peaks are skipped). If nothing qualifies, it falls back to argmax_candidate.

The head path uses decode_cell_relative + sigmoid(objectness) + nms(iou_threshold=0.50), capped at the specialist's max_candidates (6 by default, not the config's 20). The degenerate-box guard (DEGENERATE_AREA_FRACTION = 0.9) applies to the zero-shot path only.


10. The VQA / caption data path

The VQA and caption paths share the display decode and the deterministic quality gate:

  1. load_image_array(asset.path) β†’ (H, W, 3) uint8.
  2. The quality gate (assess_image_quality) β€” a NOISE or INVALID_VALUES verdict blocks the VLM call; a FLAT or TOO_SMALL verdict is usable but degraded. The gate is wired into the serving path, not merely defined: specialists/vqa/inference.py:190 computes it and returns self._refuse(...) when not quality.is_usable, before the model is constructed or called. The comment there states the reason β€” "The VLM will produce a fluent, specific, entirely fabricated scene description for ANY input it is handed, including uniform noise, even when the prompt explicitly instructs it to decline. Measured, not theorised. The guard therefore has to sit here, before the model."
  3. The processor builds pixel_values with processor_longest_edge: 512 pinned, so a 512-px tile is not upscaled and split. The measured contrast is recorded in configs/base.yaml: default β†’ pixel_values (1, 17, 3, 512, 512), 1142 prompt tokens; pinned β†’ pixel_values (1, 1, 3, 512, 512).
  4. Prompts must go through processor.apply_chat_template() β€” hand-written prompt strings raise ValueError because SmolVLM requires one <image> token per image (finding F5-3).
  5. max_images_per_call: 1 β€” the VLM sees one image per call.

The VLM adapter is loaded from SATQUERY_VLM_ADAPTER when set (the resolution order is: explicit argument, then that env var, then none). The adapter's metrics are usable (exact_match 0.963) but its status is ACCEPTANCE-REJECTED β€” USABLE β‰  ACCEPTED.


11. Determinism and reproducibility

11.1 What is deterministic

Stage Determinism Mechanism
load_image_array Deterministic Pure percentile stretch; no sampling
to_grayscale_float Deterministic Pure percentile stretch
assess_image_quality Deterministic "No model, no randomness, no thresholds learned from data"
normalise_for_croma Deterministic "No sampling, no learned statistics, no RNG, and no dependence on batch composition"
infer_modality Deterministic Pure function of band count and explicit label
Sensor adapter Deterministic Pure band placement + zero-fill
assemble_fusion_input Deterministic Concatenation with asserted order
resize_to_canonical Deterministic Bilinear interpolation
Registration measurement Deterministic cv2.phaseCorrelate; no sampling
Post-processing (morphology, components) Deterministic cv2 fixed operations
channel_dropout Seeded Takes an rng; "Seeded deliberately: an unseeded training transform makes a run unreproducible"
Router training Seeded project.seed: 42
Fusion training Seeded Explicit seeding, per train_change_head's discipline

11.2 The global seed

configs/base.yaml declares project.seed: 42. Config.seed returns int(self.get("project.seed", 42)). Training loops take it from the config.

11.3 The config hash

Config.hash is sha256(json.dumps(self._data, sort_keys=True, default=str).encode())[:16]. The frozen value is 78f1e3700da15aa1. It is recorded in every evaluation run, and scripts/eval_change.py refuses to score on drift (exit 3). This is why several pipeline decisions live in code constants or environment channels rather than in configs/base.yaml: adding a key moves the hash and detaches the published benchmark from its config.

11.4 Environment channels β€” two distinct classes

There are two different kinds of environment override, and conflating them would be wrong:

(a) Pre-hash overrides β€” these DO change Config.hash. load_config applies these into the data dictionary before Config(data) is constructed:

Variable Effect
SATQUERY_PRECISION Sets training.precision
SATQUERY_TORCH_COMPILE Sets deployment.torch_compile

Because they are merged before hashing, a run using either has a different Config.hash from the frozen one.

(b) Hash-exempt overrides β€” these do NOT change Config.hash. They are read at use time, outside the config object:

Variable Read by Purpose
SATQUERY_DEVICE core/config.py::Config.device_preference Device selection
SATQUERY_CROMA_USE_8_BIT specialists/optical_sar/radiometry.py::resolve_use_8_bit The uint8 arm (A/B)
SATQUERY_CROMA_CHECKPOINT specialists/optical_sar/croma.py::resolve_checkpoint_path CROMA checkpoint path (first in resolution order)
SATQUERY_VLM_ADAPTER specialists/vqa/model.py VLM adapter path

The distinction matters for reproducibility: a run that sets SATQUERY_DEVICE is comparable to a run that does not (same config identity), while a run that sets SATQUERY_PRECISION is not (different hash).

11.5 Checkpoint resolution order

specialists/optical_sar/croma.py::resolve_checkpoint_path resolves the CROMA checkpoint offline-first: SATQUERY_CROMA_CHECKPOINT (when set and non-empty, and the path must exist β€” otherwise it raises naming the variable), then the Hugging Face Hub cache. The module records that the use_8_bit resolution order is "SATQUERY_CROMA_USE_8_BIT and then croma.use_8_bit."

11.6 What "reproducible" means per artefact

Artefact Reproducible? How
Display array from a raster βœ… byte-identical Deterministic stretch
Availability mask βœ… byte-identical Deterministic placement
normalise_for_croma output βœ… byte-identical Pure function
Change feature cache βœ… given the same spec hash and detector state Spec-guarded
Text feature cache βœ… given the same encoder + revision Spec-guarded
Fusion feature cache βœ… given the same arm, resolution and config Provenance-guarded
Router corpus cache βœ… given the same encoder id + text set Fingerprint-guarded
A trained head ❌ not byte-reproducible across hardware Seeded, but GPU nondeterminism applies

UNKNOWN β€” not established from the available evidence whether bit-exact reproducibility of a trained checkpoint is achieved across GPU runs; the code seeds explicitly but does not claim cross-device determinism.


12. Where intermediate artefacts live

All generated artefacts live under artifacts/, which is "the repository's generated-artefact root (not source, not config, not data/ inputs)."

12.1 Caches (reproducible from a corpus)

Artefact Path Producer Key Reproducible?
Router corpus embeddings <cache_dir>/embeddings_<fingerprint[:16]>.npy + .json router/train.py::embed_corpus_cached corpus fingerprint (v1 + encoder id + max_length + sorted texts) βœ…
Change features artifacts/change_vqa/change_features.npz scripts/prepare_change_vqa.py feature_spec_hash(256, "stanet:trained") = c801326f85a185f8 βœ…
Question features artifacts/change_vqa/text_features_<split>.npz scripts/prepare_change_vqa.py text_spec_hash(...) = d2801ea1a314354a (pinned revision) βœ…
Scene targets artifacts/change_vqa/scene_targets.jsonl scripts/prepare_change_vqa.py change_vqa_preproc_v1 βœ…
Scene manifest artifacts/change_vqa/scene_manifest.jsonl scripts/prepare_change_vqa.py change_vqa_preproc_v1 βœ…
Fusion features (Arm A) artifacts/optical_sar/fusion_features/train.npz + .json scripts/extract_fusion_features.py CACHE_VERSION="v1" + arm + config + checkpoint digest βœ…
Fusion features (Arm B) β€” β€” β€” NOT PRODUCED

A naming caution. There is no fusion_features_armB directory (or any armB/arm_b literal) in the repository β€” a grep for those spellings returns nothing. The arm is a field inside the cache's JSON sidecar and the run record, not a directory name; the Arm-B cache would be a separate cache built under the same artifacts/optical_sar/fusion_features/ root, distinguished by its provenance record. Since that cache has not been built, the root holds only the Arm-A train.npz + train.json.

12.2 Trained artefacts (not reproducible byte-for-byte)

Artefact Path Notes
Grounding head artifacts/grounding/remoteclip_grounding_v001/head.pt The shipped default; DEFAULT_HEAD_PATH in code
Change detector artifacts/change/levir_change_v001/head.pt Test pooled IoU 0.8122; not wired into serving by default
Change-VQA head artifacts/change_vqa/run/head.pt change_vqa_head_v1, 1,453,912 params
Fusion head (production) artifacts/optical_sar/fusion_head_production_v001/ pre_registered_115_metric.json

12.3 Evidence artefacts

Artefact Path
CROMA forward pass artifacts/optical_sar/croma_forward.json
CDVQA imagery verification artifacts/cdvqa/imagery_verification.json
CDVQA/SECOND overlap artifacts/cdvqa/second_overlap.json
CDVQA temporal order artifacts/cdvqa/temporal_order_evidence_v2.json
Resolution experiment per_sample_224_full.jsonl, per_sample_448_full.jsonl, resolution_experiment_full.json
Phase 12 selection manifest artifacts/phase12_selection/selection_manifest_seed10.jsonl

12.4 What the serving path writes

The serving path writes only server-side diagnostic artefacts, and their refs are always null to the client (F-16): the change map (change_map_<stem>.tif) and the optical/SAR views (optical_view_<stem>.png, sar_view_<stem>.png). "The map is still WRITTEN. It is the operator's diagnostic ... What changes is only that its location is a server-side fact, not a client-facing one." No artifact:// URI is fabricated in its place.


13. What is NOT part of the pipeline

Each item below was checked against the code. Where a capability is absent, it is named rather than implied.

Capability Status Evidence
Reprojection NOT part of the pipeline compare_crs reports requires_reprojection; no code reprojects.
Cloud masking NOT part of the pipeline No cloud/QA/SCL handling anywhere.
Atmospheric correction NOT part of the pipeline Rasters are consumed as delivered.
Mosaicking NOT part of the pipeline No composite or seamline code.
Tiling / top-K tile selection DECLARED, not executed Read only by core/config.py's top_k_tiles <= max_tiles check.
Percentile / dB conditioning DECLARED, not executed Read by no code; radiometry.py docstring; PHASE14 Β§1.
Nodata filling NOT part of the pipeline quality.py refuses non-finite input; the fix is "deferred to the tiling work."
Pair-dimension reconciliation NOT part of the pipeline The change path requires equal shapes; no reconciling resize.
A serving-time feature cache NOT part of the pipeline Only the router's training corpus cache and the change-VQA/fusion preparation caches exist; serving encodes live.
Cross-step artefact hand-off NOT part of the pipeline The change-VQA specialist "accepts an optional change_map path in request.params for a future planner to populate. It is unused today" β€” "the controller does not hand artifacts between steps."
A single-pass change+language request NOT part of the pipeline Routing a change+language request plans both a change and a change_vqa step, so "the detector runs twice for one request" β€” the weights load once (the registry caches the instance), but the forward pass runs twice.
Image resolution decoding NOT part of the pipeline CDVQA's res_x/res_y are the opaque strings ".1524m" on every row; "The field is of unknown semantics and is stored as an opaque string. It is not used anywhere."

14. What is NOT RUN / OPEN / BLOCKED for this topic

NOT RUN

  • The router test split. Reported accuracy is 0.965116 on validation, ungated, n = 86; the test split was NOT RUN.
  • The optical-SAR forward pass with a trained head at serving time. CROMA is not loaded locally and no trained head is wired; the specialist runs sensor-only.
  • The CROMA normalisation experiment (option c). Not run; the arm set is frozen at A/B and arm B (the registered one) is non-constructible.
  • An end-to-end system benchmark. No system-level accuracy is claimed.
  • A change+language request end-to-end. The planner would run the detector twice; the artefact hand-off that would make it once is not implemented.

OPEN

  • The percentile/dB conditioning stage. Either implement it or retire the keys.
  • The arm-constructibility question. A human decision, deliberately unresolved.
  • The optical-SAR ruling. Accuracy 0.931 with macro-F1 0.434161; OPEN. Never quote the accuracy without the macro-F1.
  • The change-VQA metric ruling. OPEN and owner-gated (two test sets: test 0.697626/0.378373 and test2 0.651469/0.372309).
  • Calibration. ECE went 0.013755 β†’ 0.014929 β€” worse. Retained only because it is in the frozen config.
  • The VLM adapter. Metrics usable (exact_match 0.963) but status ACCEPTANCE-REJECTED.

BLOCKED

  • The Arm-B fusion feature cache. artifacts/optical_sar/fusion_features/ holds only the Arm-A cache; val.npz and test.npz were pending. Arm B cannot be trained without its own cache.
  • A clean change demo. Blocked on a same-shape pair or a resize step.

DEFERRED

  • codespace_name trailing \n in the /api/health payload (B-02). Cosmetic; the wake path is safe.
  • Nodata filling from the raster profile. Deferred to the tiling work, which is not implemented.

15. Where the evidence lives

Source modules (the pipeline this chapter describes):

Path What it defines
preprocessing/raster.py inspect_raster, read_bands, write_raster, infer_modality, file_sha256, pixel_area_m2
preprocessing/imagery.py load_image_array, load_pil_image β€” the display decode
preprocessing/quality.py assess_image_quality, the deterministic gate
specialists/optical_sar/sensor_adapter.py Band mapping, zero-fill, availability mask
specialists/optical_sar/radiometry.py normalise_for_croma, resolve_use_8_bit, the zero-channel rule
specialists/optical_sar/inference.py run_pipeline, resize_to_canonical, build_masks
specialists/optical_sar/fusion_head.py assemble_fusion_input, channel_dropout, expected_fusion_dim
specialists/optical_sar/specialist.py execute, _descriptor_for, _confidence_components
specialists/change/specialist.py _assess_pair, _predict, _write_artifacts
specialists/change/postprocess.py measure_registration, to_grayscale_float, postprocess_change_map
specialists/change/vqa_specialist.py feature_spec_mismatch, unavailable_reason, execute
specialists/grounding/specialist.py execute, _decode_with_head, _to_geo_boxes
specialists/grounding/inference.py decode_candidates_from_features, ground_phrase
training/change_vqa/features.py feature_spec_hash, text_spec_hash, ChangeFeatureExtractor, the caches
training/change_vqa/dataset.py DATASET_ID, PREPROCESSING_VERSION, SPLIT_MAP, LABEL_PALETTE
training/fusion/train.py CACHE_VERSION, DEFAULT_FEATURE_CACHE_DIR, feature_cache_path, write_feature_cache, ARMS
training/fusion/extract.py The Phase 12 producer, the DEV-2 stretch, _assert_raw_encoder
router/encoder.py FrozenEncoder, the verified MiniLM contract
router/train.py CORPUS_CACHE_VERSION, _corpus_fingerprint, embed_corpus_cached
core/config.py load_config, Config.hash, device_preference, the env overrides
core/controller.py _inspect_asset, the nine-state machine

Scripts (the producers):

  • scripts/prepare_change_vqa.py β€” scene targets, manifest, change features, text features.
  • scripts/extract_fusion_features.py β€” the Phase 12 fusion feature cache (exit codes 0/2/3).
  • scripts/train_fusion.py, scripts/train_change_vqa.py, scripts/train_grounding.py, scripts/train_router.py β€” the training loops.
  • scripts/diagnose_feature_cache.py β€” cache diagnosis.

Configuration:

  • configs/base.yaml β€” project.seed, image.*, optical.*, sar.*, croma.*, fusion.*, grounding.*, change.*, router.*, vlm.*, training.*.

Documents:

  • docs/PHASE14_CROMA_NORMALISATION_CHANGE.md β€” the two ordered stages, the Gate F arm amendment, the Arm-B blocker.
  • docs/PHASE12_ENTRY_GATE.md, docs/PHASE12_R08_TRUTHFUL_BACKFILL.md β€” the Phase 12 run records and the result_status plumbing.
  • docs/PHASE10_CDVQA_DATA_STATUS.md β€” the CDVQA/SECOND layout, the temporal semantics, the four adapter defects found by execution.
  • docs/PHASE7_RESOLUTION_DECISION.md β€” the 224 decision (the cached-feature experiment).
  • docs/API_CONTRACT.md Β§2.5 β€” the asset allowlist and the upload contract.
  • docs/FINAL_DELIVERY_REPORT.md Β§4 (run ids), Β§5 (metrics), Β§6 (the 726Β²/736Β² degraded change demo).

Sibling release chapters: