thundercode commited on
Commit
394471d
Β·
verified Β·
1 Parent(s): 18f4386

release: add docs/MASTER_ARCHITECTURE_PLAN.md

Browse files
Files changed (1) hide show
  1. docs/MASTER_ARCHITECTURE_PLAN.md +4656 -0
docs/MASTER_ARCHITECTURE_PLAN.md ADDED
@@ -0,0 +1,4656 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SatQuery AI β€” Master Architecture & Implementation Plan (ORIGINAL, HISTORICAL)
2
+
3
+ > **STATUS: HISTORICAL / SUPERSEDED IN PARTS.** This is the project's original master plan, preserved
4
+ > verbatim as the design record. It is **not** a description of the shipped system. Where it and the
5
+ > rest of this documentation disagree, the rest of this documentation wins.
6
+ >
7
+ > **Known superseded points:**
8
+ >
9
+ > - It specifies a **Gradio GUI**. The shipped system is a **static frontend** on Cloudflare Pages;
10
+ > `app/space_app.py` serves JSON only and deliberately builds no Gradio interface.
11
+ > - It specifies an **HF Space + ZeroGPU + Railway** deployment. The shipped system is
12
+ > **Cloudflare Pages β†’ Render β†’ outbound tunnel β†’ GitHub Codespace**, CPU-first.
13
+ > - Its **implementation-order phases** (0–19) and stop/go gates describe the plan, not the record of
14
+ > what was executed. For actual status see [`CHANGELOG.md`](CHANGELOG.md),
15
+ > [`LIMITATIONS.md`](LIMITATIONS.md) and the architecture reference in
16
+ > [`ARCHITECTURE.md`](ARCHITECTURE.md).
17
+ > - Several sections describe capabilities that were later **rejected or measured differently** β€” for
18
+ > example the grounding resolution (the plan allows 448; the shipped value is 224, after 448 lost a
19
+ > pre-registered paired test) and the VLM adapter (trained, then **acceptance-rejected**).
20
+ >
21
+ > It is included because the decisions it records β€” the frozen backbones, the typed contracts, the
22
+ > evidence/confidence split, and the "no LLM-generated confidence" rule β€” are the decisions the system
23
+ > still obeys.
24
+
25
+ ---
26
+
27
+
28
+
29
+ # SatQuery AI β€” Revised Master Architecture & Implementation Plan
30
+
31
+ # 0. Executive Decision
32
+
33
+ ## 0.1 Final architecture
34
+
35
+ Build SatQuery AI as a **modular monolith**.
36
+
37
+ One application.
38
+
39
+ One Python backend.
40
+
41
+ One Gradio GUI.
42
+
43
+ One controller.
44
+
45
+ Multiple independently callable specialist modules with stable typed interfaces.
46
+
47
+ ```
48
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
49
+
50
+ ` β”‚ Gradio GUI β”‚`
51
+
52
+ ` β”‚ Upload / Query / Map β”‚`
53
+
54
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
55
+
56
+ ` β”‚`
57
+
58
+ ` β–Ό`
59
+
60
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
61
+
62
+ ` β”‚ Query Processor β”‚`
63
+
64
+ ` β”‚ normalize / sanitize β”‚`
65
+
66
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
67
+
68
+ ` β”‚`
69
+
70
+ ` β–Ό`
71
+
72
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
73
+
74
+ ` β”‚ Tiny NLP Intent β”‚`
75
+
76
+ ` β”‚ Router β”‚`
77
+
78
+ ` β”‚ MiniLM + Adapter β”‚`
79
+
80
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
81
+
82
+ ` β”‚`
83
+
84
+ ` β–Ό`
85
+
86
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
87
+
88
+ ` β”‚ Deterministic Policy β”‚`
89
+
90
+ ` β”‚ / Workflow Controller β”‚`
91
+
92
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
93
+
94
+ ` β”‚`
95
+
96
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
97
+
98
+ ` β”‚ β”‚ β”‚`
99
+
100
+ ` β–Ό β–Ό β–Ό`
101
+
102
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
103
+
104
+ ` β”‚ VLM β”‚ β”‚ Grounding β”‚ β”‚ Change β”‚`
105
+
106
+ ` β”‚ Specialist β”‚ β”‚ Specialist β”‚ β”‚ Specialist β”‚`
107
+
108
+ ` β”‚ SmolVLM β”‚ β”‚ RemoteCLIP β”‚ β”‚ STANet β”‚`
109
+
110
+ ` β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜`
111
+
112
+ ` β”‚ β”‚ β”‚`
113
+
114
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
115
+
116
+ ` β”‚`
117
+
118
+ ` β–Ό`
119
+
120
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
121
+
122
+ ` β”‚ Optical-SAR Specialistβ”‚`
123
+
124
+ ` β”‚ CROMA β”‚`
125
+
126
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
127
+
128
+ ` β”‚`
129
+
130
+ ` β–Ό`
131
+
132
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
133
+
134
+ ` β”‚ Evidence Engine β”‚`
135
+
136
+ ` β”‚ boxes / masks / crops β”‚`
137
+
138
+ ` β”‚ maps / modality proof β”‚`
139
+
140
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
141
+
142
+ ` β”‚`
143
+
144
+ ` β–Ό`
145
+
146
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
147
+
148
+ ` β”‚ Confidence Calibrationβ”‚`
149
+
150
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
151
+
152
+ ` β”‚`
153
+
154
+ ` β–Ό`
155
+
156
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
157
+
158
+ ` β”‚ Result Normalizer β”‚`
159
+
160
+ ` β”‚ JSON / Answer / Trace β”‚`
161
+
162
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
163
+
164
+ ` β”‚`
165
+
166
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
167
+
168
+ ` β–Ό β–Ό`
169
+
170
+ ` GUI visualization PDF/JSON report`
171
+ ```
172
+
173
+
174
+ ## 0.2 Specialist stack
175
+
176
+ | **Role** | **Model/component** | **Decision** |
177
+ | :-: | :-: | :-: |
178
+ | Intent understanding | `sentence-transformers/all-MiniLM-L6-v2` + small classifier adapter | **Primary** |
179
+ | Main VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | **Primary** |
180
+ | Grounding | RemoteCLIP ViT-B/32 + lightweight grounding head | **Primary initial grounding design** |
181
+ | Change detection | STANet-style Siamese network | **Primary** |
182
+ | Optical-SAR representation | CROMA-base | **Primary** |
183
+ | Geospatial processing | Rasterio + pyproj | **Primary** |
184
+ | Report generation | ReportLab | **Primary** |
185
+ | GUI | Gradio | **Primary** |
186
+
187
+ RemoteCLIP provides official pretrained RN50, ViT-B/32 and ViT-L/14 checkpoints and is explicitly built as a remote-sensing vision-language model. The official repository provides the `ViT-B-32` checkpoint and OpenCLIP loading path. ([GitHub](https://github.com/ChenDelong1999/RemoteCLIP?utm_source=chatgpt.com))
188
+
189
+ CROMA's official implementation provides pretrained `base` and `large` weights and explicit optical, SAR and joint representations; its base implementation uses a 768-dimensional representation with separate Sentinel-1 two-channel and Sentinel-2 twelve-channel inputs. ([GitHub](https://github.com/antofuller/CROMA/blob/main/README.md?utm_source=chatgpt.com))
190
+
191
+
192
+ ## 0.3 Mandatory capabilities
193
+
194
+ SatQuery must expose:
195
+
196
+ ```
197
+ `Single-image VQA`
198
+
199
+ `Single-image captioning`
200
+
201
+ `Text-guided grounding`
202
+
203
+ `Bi-temporal change analysis`
204
+
205
+ `Optical-SAR analysis`
206
+
207
+ `Natural-language query routing`
208
+
209
+ `Evidence generation`
210
+
211
+ `Confidence`
212
+
213
+ `Execution trace`
214
+
215
+ `GeoTIFF handling`
216
+
217
+ `Downloadable result/report`
218
+ ```
219
+
220
+ The uploaded project contract already requires VQA, an additional single-image task, temporal change, optical-SAR joint analysis and agentic orchestration. Grounding is now added as a first-class capability.
221
+
222
+
223
+ ## 0.4 The key architectural principle
224
+
225
+ The NLP router **understands**.
226
+
227
+ The policy engine **decides**.
228
+
229
+ The specialists **compute**.
230
+
231
+ The VLM **explains**.
232
+
233
+ The evidence engine **proves**.
234
+
235
+ That separation is the foundation of the entire system.
236
+
237
+
238
+ # 1. Requirement-to-Implementation Traceability Matrix
239
+
240
+ | **Requirement** | Mandatory | **Component** | **Dataset** | **Model** | **Training** | **Runtime** | **Evidence** | **Public Test** | **ISRO/SAC** |
241
+ | :-: | -: | :-: | :-: | :-: | :-: | :-: | :-: | :-: | :-: |
242
+ | Single optical/MS image | Yes | Raster loader + VLM | VRSBench / RSVQA | SmolVLM | PEFT | VQA | selected image/tile | Yes | Yes |
243
+ | Single SAR | Yes | SAR adapter + VLM | BigEarthNet S1 | SmolVLM/CROMA | adapter | VQA | SAR view | Where relevant | Yes |
244
+ | Optical-SAR pair | Yes | Fusion specialist | BigEarthNet S1/S2 | CROMA | fusion head | fusion DAG | optical + SAR + joint | Where prescribed | **Primary** |
245
+ | Bi-temporal pair | Yes | Change specialist | LEVIR-CD/CDVQA | STANet-style | supervised | change DAG | change map | Yes | Yes |
246
+ | VQA | Yes | VLM | RSVQA/VRSBench | SmolVLM | LoRA | VQA | image/tile | Yes | Yes |
247
+ | Captioning | Yes | VLM | VRSBench | SmolVLM | LoRA | caption | image | Yes | Yes |
248
+ | Grounding | Added | Grounding specialist | VRSBench referring expressions | RemoteCLIP + head | head/adapter | grounding | boxes/regions | Yes | Yes where spatial output |
249
+ | Agentic orchestration | Yes | NLP router + policy engine | synthetic routing set | MiniLM | classifier head | controller | trace | Functional | Yes |
250
+ | Evidence | Yes | Evidence engine | all | none | none | postprocessing | boxes/masks/maps/crops | Functional | Yes |
251
+ | Confidence | Yes | Calibration module | validation | specialist scores | calibration only | postprocessing | calibrated score | Yes | Yes |
252
+ | Execution trace | Yes | Trace builder | all | none | none | controller | structured trace | Yes | Yes |
253
+ | Reports | Yes | ReportLab | all | none | none | finalization | PDF/JSON | Functional | Functional |
254
+ | Geospatial preservation | Yes | Rasterio/pyproj | GeoTIFF | none | none | all spatial workflows | CRS/transform | Yes | Yes |
255
+ | Leakage controls | Yes | Dataset registry/evaluator | all | none | none | offline | run manifest | Yes | Yes |
256
+
257
+ The contract requires every mandatory criterion to have an implementation component, test coverage and evaluation path.
258
+
259
+
260
+ # 2. Evaluation-First Strategy
261
+
262
+ ## 2.1 Priority
263
+
264
+ Unless the organisers publish different weights, engineering effort is:
265
+
266
+ ```
267
+ `1. Optical-SAR joint reasoning`
268
+
269
+ `2. Change analysis`
270
+
271
+ `3. Single-image VQA`
272
+
273
+ `4. Grounding`
274
+
275
+ `5. Captioning`
276
+
277
+ `6. Agent/routing/evidence/reliability`
278
+ ```
279
+
280
+ This is an **engineering priority**, not an official score ranking. The supplied specification explicitly makes optical-SAR the primary hidden-set compatibility target.
281
+
282
+
283
+ ## 2.2 Evaluation dimensions
284
+
285
+ ### VQA
286
+
287
+ Measure:
288
+
289
+ ```
290
+ `exact match`
291
+
292
+ `normalized exact match`
293
+
294
+ `F1 where appropriate`
295
+ ```
296
+
297
+ ### Captioning
298
+
299
+ Use the benchmark-prescribed metrics.
300
+
301
+ Locally additionally calculate:
302
+
303
+ ```
304
+ `BLEU`
305
+
306
+ `ROUGE-L`
307
+
308
+ `CIDEr`
309
+
310
+ `BERTScore`
311
+ ```
312
+
313
+ ### Grounding
314
+
315
+ ```
316
+ `IoU`
317
+
318
+ `Recall@IoU`
319
+
320
+ `mAP where applicable`
321
+ ```
322
+
323
+ ### Change
324
+
325
+ ```
326
+ `precision`
327
+
328
+ `recall`
329
+
330
+ `F1`
331
+
332
+ `IoU`
333
+
334
+ `mIoU`
335
+ ```
336
+
337
+ ### Change VQA
338
+
339
+ ```
340
+ `answer accuracy`
341
+
342
+ `semantic match if benchmark specifies it`
343
+ ```
344
+
345
+ ### Optical-SAR
346
+
347
+ This must be **task-dependent**.
348
+
349
+ Possible evaluation:
350
+
351
+ ```
352
+ `classification accuracy/F1`
353
+
354
+ `VQA accuracy`
355
+
356
+ `region IoU`
357
+
358
+ `mask IoU`
359
+
360
+ `change metrics`
361
+ ```
362
+
363
+ Do not invent an official multimodal metric.
364
+
365
+
366
+ # 3. System Architecture
367
+
368
+ ## 3.1 Logical architecture
369
+
370
+ ```
371
+ `flowchart TD`
372
+
373
+ ` A\[User Query + Uploaded Images\]`
374
+
375
+ ` B\[Query Normalizer\]`
376
+
377
+ ` C\[Tiny NLP Intent Router\]`
378
+
379
+ ` D\[Intent Validator\]`
380
+
381
+ ` E\[Deterministic Policy Engine\]`
382
+
383
+ ` F\[Workflow DAG\]`
384
+
385
+
386
+ ` V\[VLM Specialist\]`
387
+
388
+ ` G\[Grounding Specialist\]`
389
+
390
+ ` CD\[Change Specialist\]`
391
+
392
+ ` OS\[Optical-SAR Specialist\]`
393
+
394
+ ` GEO\[Geospatial Engine\]`
395
+
396
+
397
+ ` EV\[Evidence Engine\]`
398
+
399
+ ` CF\[Confidence Calibration\]`
400
+
401
+ ` R\[Result Normalizer\]`
402
+
403
+ ` T\[Execution Trace\]`
404
+
405
+ ` REP\[Report Generator\]`
406
+
407
+
408
+ ` A --\> B`
409
+
410
+ ` B --\> C`
411
+
412
+ ` C --\> D`
413
+
414
+ ` D --\> E`
415
+
416
+ ` E --\> F`
417
+
418
+
419
+ ` F --\> V`
420
+
421
+ ` F --\> G`
422
+
423
+ ` F --\> CD`
424
+
425
+ ` F --\> OS`
426
+
427
+ ` F --\> GEO`
428
+
429
+
430
+ ` V --\> EV`
431
+
432
+ ` G --\> EV`
433
+
434
+ ` CD --\> EV`
435
+
436
+ ` OS --\> EV`
437
+
438
+ ` GEO --\> EV`
439
+
440
+
441
+ ` EV --\> CF`
442
+
443
+ ` CF --\> R`
444
+
445
+ ` R --\> T`
446
+
447
+ ` R --\> REP`
448
+
449
+ ` R --\> A`
450
+ ```
451
+
452
+
453
+ ## 3.2 Deployable competition architecture
454
+
455
+ ```
456
+ `satquery-ai/`
457
+
458
+ ` one Python process`
459
+
460
+ ` β”‚`
461
+
462
+ ` β”œβ”€β”€ controller`
463
+
464
+ ` β”œβ”€β”€ router`
465
+
466
+ ` β”œβ”€β”€ VLM`
467
+
468
+ ` β”œβ”€β”€ grounding`
469
+
470
+ ` β”œβ”€β”€ change`
471
+
472
+ ` β”œβ”€β”€ CROMA`
473
+
474
+ ` β”œβ”€β”€ evidence`
475
+
476
+ ` β”œβ”€β”€ evaluation`
477
+
478
+ ` └── GUI`
479
+ ```
480
+
481
+ No microservices.
482
+
483
+ No Kubernetes.
484
+
485
+ No queue cluster.
486
+
487
+ No vector database.
488
+
489
+ No distributed infrastructure.
490
+
491
+ The **interfaces** are designed so those can be added later if necessary, but they are not implemented for the competition prototype.
492
+
493
+
494
+ # 4. Exact Specialist Model Stack
495
+
496
+ ## 4.1 NLP intent router
497
+
498
+ ### Model
499
+
500
+ `sentence-transformers/all-MiniLM-L6-v2`
501
+
502
+ License: Apache-2.0. ([Hugging Face](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/tree/main?utm_source=chatgpt.com))
503
+
504
+ ### Architecture
505
+
506
+ ```
507
+ `query`
508
+
509
+ ` ↓`
510
+
511
+ `MiniLM encoder`
512
+
513
+ ` ↓`
514
+
515
+ `384-d embedding`
516
+
517
+ ` ↓`
518
+
519
+ `LayerNorm`
520
+
521
+ ` ↓`
522
+
523
+ `Linear(384 β†’ 128)`
524
+
525
+ ` ↓`
526
+
527
+ `GELU`
528
+
529
+ ` ↓`
530
+
531
+ `Dropout(0.1)`
532
+
533
+ ` ↓`
534
+
535
+ `Linear(128 β†’ task classes)`
536
+ ```
537
+
538
+ Also create auxiliary heads:
539
+
540
+ ```
541
+ `task`
542
+
543
+ `modality`
544
+
545
+ `temporal`
546
+
547
+ `spatial\_output`
548
+
549
+ `language\_output`
550
+ ```
551
+
552
+ ### Trainable portion
553
+
554
+ Only the classification adapter initially.
555
+
556
+ Keep MiniLM frozen.
557
+
558
+ This is enough because the router has a tiny, closed ontology.
559
+
560
+
561
+ # 5. Router Training Dataset
562
+
563
+ Create a **manually curated + synthetic query dataset**.
564
+
565
+ Classes:
566
+
567
+ ```
568
+ `vqa`
569
+
570
+ `caption`
571
+
572
+ `grounding`
573
+
574
+ `change`
575
+
576
+ `optical\_sar`
577
+
578
+ `unsupported`
579
+ ```
580
+
581
+ Generate approximately:
582
+
583
+ ```
584
+ `300–500 examples/class`
585
+ ```
586
+
587
+ Target:
588
+
589
+ ```
590
+ `2,000–3,000 queries total`
591
+ ```
592
+
593
+ Example:
594
+
595
+ ```
596
+ `"Describe this image."`
597
+
598
+ `β†’ caption`
599
+
600
+
601
+ `"What buildings can you see?"`
602
+
603
+ `β†’ vqa`
604
+
605
+
606
+ `"Locate the water body."`
607
+
608
+ `β†’ grounding`
609
+
610
+
611
+ `"Has the urban area increased?"`
612
+
613
+ `β†’ change`
614
+
615
+
616
+ `"Compare optical and SAR to identify built-up regions."`
617
+
618
+ `β†’ optical\_sar`
619
+ ```
620
+
621
+ ### Hard negatives
622
+
623
+ This is important.
624
+
625
+ Include confusing examples:
626
+
627
+ ```
628
+ `"Describe the changes."`
629
+ ```
630
+
631
+ vs
632
+
633
+ ```
634
+ `"What changed?"`
635
+ ```
636
+
637
+ vs
638
+
639
+ ```
640
+ `"Where did the change happen?"`
641
+ ```
642
+
643
+ All map to `change`, but the latter also requires spatial output.
644
+
645
+ Similarly:
646
+
647
+ ```
648
+ `"Describe the water body."`
649
+ ```
650
+
651
+ β†’ caption/VQA
652
+
653
+ while:
654
+
655
+ ```
656
+ `"Show me the water body."`
657
+ ```
658
+
659
+ β†’ grounding.
660
+
661
+
662
+ # 6. Tiny NLP Router Decision Process
663
+
664
+ The router returns:
665
+
666
+ ```
667
+ `\{`
668
+
669
+ ` "task": "grounding",`
670
+
671
+ ` "modality": "optical",`
672
+
673
+ ` "temporal": false,`
674
+
675
+ ` "spatial\_output": true,`
676
+
677
+ ` "language\_output": true,`
678
+
679
+ ` "confidence": 0.94`
680
+
681
+ `\}`
682
+ ```
683
+
684
+ Then deterministic validation runs.
685
+
686
+ Example:
687
+
688
+ ```
689
+ `task = grounding`
690
+
691
+ `image\_count = 1`
692
+
693
+ `modality = optical`
694
+ ```
695
+
696
+ Valid.
697
+
698
+ But:
699
+
700
+ ```
701
+ `task = change`
702
+
703
+ `image\_count = 1`
704
+ ```
705
+
706
+ Invalid.
707
+
708
+ The policy engine changes this to:
709
+
710
+ ```
711
+ `\{`
712
+
713
+ ` "status": "invalid\_request",`
714
+
715
+ ` "reason": "change analysis requires two compatible temporal inputs"`
716
+
717
+ `\}`
718
+ ```
719
+
720
+ The router never overrides input reality.
721
+
722
+
723
+ # 7. Agentic Controller
724
+
725
+ The agent is therefore:
726
+
727
+ ```
728
+ `NLP understanding`
729
+
730
+ `+`
731
+
732
+ `deterministic planning`
733
+
734
+ `+`
735
+
736
+ `specialist invocation`
737
+
738
+ `+`
739
+
740
+ `result aggregation`
741
+ ```
742
+
743
+ It is not:
744
+
745
+ ```
746
+ `LLM β†’ arbitrary tool calling`
747
+ ```
748
+
749
+ ### Controller states
750
+
751
+ ```
752
+ `RECEIVE`
753
+
754
+ ` ↓`
755
+
756
+ `PARSE`
757
+
758
+ ` ↓`
759
+
760
+ `VALIDATE`
761
+
762
+ ` ↓`
763
+
764
+ `PLAN`
765
+
766
+ ` ↓`
767
+
768
+ `PREPROCESS`
769
+
770
+ ` ↓`
771
+
772
+ `EXECUTE`
773
+
774
+ ` ↓`
775
+
776
+ `AGGREGATE`
777
+
778
+ ` ↓`
779
+
780
+ `VERIFY`
781
+
782
+ ` ↓`
783
+
784
+ `RESPOND`
785
+ ```
786
+
787
+ Each state has allowed transitions.
788
+
789
+ This satisfies the β€œagentic” requirement while remaining reproducible. The project contract specifically permits and prefers constrained state-machine orchestration over unrestricted autonomous agents.
790
+
791
+
792
+ # 8. Specialist Interface
793
+
794
+ Every specialist must implement the same conceptual interface:
795
+
796
+ ```
797
+ `class Specialist:`
798
+
799
+ ` name: str`
800
+
801
+ ` version: str`
802
+
803
+ ` capabilities: list\[str\]`
804
+
805
+
806
+ ` def validate\_request(self, request):`
807
+
808
+ ` ...`
809
+
810
+
811
+ ` def execute(self, request):`
812
+
813
+ ` ...`
814
+
815
+
816
+ ` def estimate\_confidence(self, output):`
817
+
818
+ ` ...`
819
+
820
+
821
+ ` def produce\_evidence(self, output):`
822
+
823
+ ` ...`
824
+ ```
825
+
826
+ The implementation may be ordinary Python.
827
+
828
+ This is our **future scaling seam**.
829
+
830
+ Today:
831
+
832
+ ```
833
+ `specialist.execute()`
834
+ ```
835
+
836
+ Tomorrow:
837
+
838
+ ```
839
+ `specialist\_worker.execute\_remote()`
840
+ ```
841
+
842
+ The controller doesn't care.
843
+
844
+
845
+ # 9. Workflow Definitions
846
+
847
+ ## 9.1 Workflow A β€” Single-image VQA
848
+
849
+ ```
850
+ `Input`
851
+
852
+ ` ↓`
853
+
854
+ `validate`
855
+
856
+ ` ↓`
857
+
858
+ `identify modality`
859
+
860
+ ` ↓`
861
+
862
+ `normalize`
863
+
864
+ ` ↓`
865
+
866
+ `tile if necessary`
867
+
868
+ ` ↓`
869
+
870
+ `select informative tiles`
871
+
872
+ ` ↓`
873
+
874
+ `SmolVLM`
875
+
876
+ ` ↓`
877
+
878
+ `answer normalization`
879
+
880
+ ` ↓`
881
+
882
+ `evidence generation`
883
+
884
+ ` ↓`
885
+
886
+ `confidence`
887
+
888
+ ` ↓`
889
+
890
+ `result`
891
+ ```
892
+
893
+ ### Tile policy
894
+
895
+ Whole-image view first.
896
+
897
+ If image exceeds configured resolution:
898
+
899
+ ```
900
+ `whole-image thumbnail`
901
+
902
+ ` ↓`
903
+
904
+ `candidate tiles`
905
+
906
+ ` ↓`
907
+
908
+ `top K = 4`
909
+ ```
910
+
911
+ Do not send every tile through the VLM.
912
+
913
+
914
+ # 10. Workflow B β€” Caption
915
+
916
+ ```
917
+ `Input`
918
+
919
+ ` ↓`
920
+
921
+ `validation`
922
+
923
+ ` ↓`
924
+
925
+ `normalization`
926
+
927
+ ` ↓`
928
+
929
+ `whole-image description`
930
+
931
+ ` ↓`
932
+
933
+ `optional selected tiles`
934
+
935
+ ` ↓`
936
+
937
+ `deduplicate repeated objects`
938
+
939
+ ` ↓`
940
+
941
+ `caption composition`
942
+
943
+ ` ↓`
944
+
945
+ `confidence`
946
+
947
+ ` ↓`
948
+
949
+ `result`
950
+ ```
951
+
952
+
953
+ # 11. Workflow C β€” Grounding
954
+
955
+ This is newly promoted to a first-class workflow.
956
+
957
+ ```
958
+ `Query`
959
+
960
+ ` ↓`
961
+
962
+ `NLP router`
963
+
964
+ ` ↓`
965
+
966
+ `grounding intent`
967
+
968
+ ` ↓`
969
+
970
+ `GeoTIFF validation`
971
+
972
+ ` ↓`
973
+
974
+ `image normalization`
975
+
976
+ ` ↓`
977
+
978
+ `RemoteCLIP image encoder`
979
+
980
+ ` ↓`
981
+
982
+ `text encoder`
983
+
984
+ ` ↓`
985
+
986
+ `multiscale candidate generation`
987
+
988
+ ` ↓`
989
+
990
+ `grounding head`
991
+
992
+ ` ↓`
993
+
994
+ `candidate boxes`
995
+
996
+ ` ↓`
997
+
998
+ `NMS`
999
+
1000
+ ` ↓`
1001
+
1002
+ `best region(s)`
1003
+
1004
+ ` ↓`
1005
+
1006
+ `optional mask refinement`
1007
+
1008
+ ` ↓`
1009
+
1010
+ `SmolVLM explanation`
1011
+
1012
+ ` ↓`
1013
+
1014
+ `confidence`
1015
+
1016
+ ` ↓`
1017
+
1018
+ `result`
1019
+ ```
1020
+
1021
+
1022
+ # 12. Grounding Specialist
1023
+
1024
+ RemoteCLIP's official repository provides remote-sensing checkpoints for RN50, ViT-B/32 and ViT-L/14 and explicitly supports image-text semantic alignment. ([GitHub](https://github.com/ChenDelong1999/RemoteCLIP?utm_source=chatgpt.com))
1025
+
1026
+ ### Primary checkpoint
1027
+
1028
+ ```
1029
+ `chendelong/RemoteCLIP`
1030
+
1031
+ `RemoteCLIP-ViT-B-32.pt`
1032
+ ```
1033
+
1034
+ Use ViT-B/32 first.
1035
+
1036
+ Do **not** start with ViT-L/14.
1037
+
1038
+ ### Why?
1039
+
1040
+ ViT-B/32 gives:
1041
+
1042
+ ```
1043
+ `remote-sensing language alignment`
1044
+
1045
+ `+`
1046
+
1047
+ `manageable compute`
1048
+
1049
+ `+`
1050
+
1051
+ `pretrained semantics`
1052
+ ```
1053
+
1054
+ and we only need the model to generate a strong embedding/feature base.
1055
+
1056
+
1057
+ # 13. Grounding Head
1058
+
1059
+ Architecture:
1060
+
1061
+ ```
1062
+ `RemoteCLIP image features`
1063
+
1064
+ ` β”‚`
1065
+
1066
+ ` β”œβ”€β”€ global embedding`
1067
+
1068
+ ` β”‚`
1069
+
1070
+ ` └── patch features`
1071
+
1072
+ ` β”‚`
1073
+
1074
+ ` multi-scale projector`
1075
+
1076
+ ` β”‚`
1077
+
1078
+ ` text/image similarity`
1079
+
1080
+ ` β”‚`
1081
+
1082
+ ` region scoring head`
1083
+
1084
+ ` β”‚`
1085
+
1086
+ ` box regression`
1087
+
1088
+ ` β”‚`
1089
+
1090
+ ` NMS`
1091
+ ```
1092
+
1093
+ ### Initial head
1094
+
1095
+ ```
1096
+ `Linear(D β†’ 512)`
1097
+
1098
+ `GELU`
1099
+
1100
+ `LayerNorm`
1101
+
1102
+ `Linear(512 β†’ 256)`
1103
+
1104
+ `GELU`
1105
+
1106
+ `Linear(256 β†’ 5)`
1107
+ ```
1108
+
1109
+ Output:
1110
+
1111
+ ```
1112
+ `x1`
1113
+
1114
+ `y1`
1115
+
1116
+ `x2`
1117
+
1118
+ `y2`
1119
+
1120
+ `confidence`
1121
+ ```
1122
+
1123
+ Use normalized coordinates internally:
1124
+
1125
+ ```
1126
+ `0–1`
1127
+ ```
1128
+
1129
+ and convert to:
1130
+
1131
+ ```
1132
+ `pixel coordinates`
1133
+ ```
1134
+
1135
+ for output.
1136
+
1137
+ VRSBench's current repository notes that provided evaluation box coordinates are normalized to 0–100, so the evaluator adapter must explicitly convert between its coordinate convention and our internal 0–1 representation rather than quietly treating the numbers as pixels. ([GitHub](https://github.com/lx709/VRSBench?utm_source=chatgpt.com))
1138
+
1139
+
1140
+ # 14. Grounding Training Strategy
1141
+
1142
+ Use VRSBench referring expressions.
1143
+
1144
+ Current VRSBench documentation reports:
1145
+
1146
+ - 29,614 images
1147
+
1148
+ - 52,472 object references
1149
+
1150
+ - 123,221 VQA pairs
1151
+
1152
+ - human-verified captions and referring annotations. ([GitHub](https://github.com/lx709/VRSBench?utm_source=chatgpt.com))
1153
+
1154
+ ### Training split
1155
+
1156
+ Never touch evaluation data while training.
1157
+
1158
+ Create:
1159
+
1160
+ ```
1161
+ `train`
1162
+
1163
+ `validation`
1164
+
1165
+ `immutable test`
1166
+ ```
1167
+
1168
+ at image/scene level.
1169
+
1170
+ ### Training
1171
+
1172
+ Freeze RemoteCLIP.
1173
+
1174
+ Train:
1175
+
1176
+ ```
1177
+ `projection`
1178
+
1179
+ `grounding head`
1180
+ ```
1181
+
1182
+ Loss:
1183
+
1184
+ ```
1185
+ `L = 0.5 \* L1\_box`
1186
+
1187
+ ` + 0.3 \* GIoU`
1188
+
1189
+ ` + 0.2 \* BCE\_confidence`
1190
+ ```
1191
+
1192
+ Initial learning rate:
1193
+
1194
+ ```
1195
+ `1e-4`
1196
+ ```
1197
+
1198
+ Batch:
1199
+
1200
+ ```
1201
+ `16–32`
1202
+ ```
1203
+
1204
+ CPU preprocessing + GPU model.
1205
+
1206
+ This should be cheap compared with VLM fine-tuning.
1207
+
1208
+
1209
+ # 15. Optional Grounding Mask Refinement
1210
+
1211
+ Do not make segmentation a dependency.
1212
+
1213
+ If a lightweight segmentation model becomes available and fits the budget:
1214
+
1215
+ ```
1216
+ `box`
1217
+
1218
+ ` ↓`
1219
+
1220
+ `mask refinement`
1221
+ ```
1222
+
1223
+ Otherwise:
1224
+
1225
+ ```
1226
+ `box`
1227
+
1228
+ ` ↓`
1229
+
1230
+ `polygon rectangle`
1231
+ ```
1232
+
1233
+ The mandatory grounding contract is satisfied by accurate spatial localization.
1234
+
1235
+
1236
+ # 16. Optical-SAR Architecture
1237
+
1238
+ CROMA remains the primary model.
1239
+
1240
+ The official implementation uses:
1241
+
1242
+ ```
1243
+ `SAR encoder`
1244
+
1245
+ `optical encoder`
1246
+
1247
+ `joint cross-modal representation`
1248
+ ```
1249
+
1250
+ and the base model uses:
1251
+
1252
+ ```
1253
+ `768 dimensional feature representation`
1254
+
1255
+ `12-channel optical input`
1256
+
1257
+ `2-channel SAR input`
1258
+ ```
1259
+
1260
+ with a 120Γ—120 default pretraining image size. ([GitHub](https://github.com/antofuller/CROMA/blob/main/README.md?utm_source=chatgpt.com))
1261
+
1262
+
1263
+ # 17. Optical-SAR Fusion
1264
+
1265
+ Use:
1266
+
1267
+ ```
1268
+ `CROMA optical feature`
1269
+
1270
+ `CROMA SAR feature`
1271
+
1272
+ `CROMA joint feature`
1273
+ ```
1274
+
1275
+ then:
1276
+
1277
+ ```
1278
+ `concat`
1279
+
1280
+ ` ↓`
1281
+
1282
+ `LayerNorm`
1283
+
1284
+ ` ↓`
1285
+
1286
+ `Linear`
1287
+
1288
+ ` ↓`
1289
+
1290
+ `GELU`
1291
+
1292
+ ` ↓`
1293
+
1294
+ `Dropout`
1295
+
1296
+ ` ↓`
1297
+
1298
+ `task head`
1299
+ ```
1300
+
1301
+ ### Do not use
1302
+
1303
+ ```
1304
+ `RGB image + colorized SAR`
1305
+ ```
1306
+
1307
+ as the final fusion architecture.
1308
+
1309
+ That is visually convincing and scientifically rather shallow.
1310
+
1311
+
1312
+ # 18. CROMA Sensor Adapter
1313
+
1314
+ CROMA's pretrained architecture is explicitly Sentinel-1/Sentinel-2 oriented. ([GitHub](https://github.com/antofuller/CROMA/blob/main/README.md?utm_source=chatgpt.com))
1315
+
1316
+ The hidden evaluation target described in your specification is Cartosat-2S + RISAT.
1317
+
1318
+ Therefore create:
1319
+
1320
+ ```
1321
+ `sensor\_adapter.py`
1322
+ ```
1323
+
1324
+ with:
1325
+
1326
+ ```
1327
+ `sensor`
1328
+
1329
+ `band\_map`
1330
+
1331
+ `normalization`
1332
+
1333
+ `availability\_mask`
1334
+
1335
+ `resolution`
1336
+ ```
1337
+
1338
+ Example:
1339
+
1340
+ ```
1341
+ `\{`
1342
+
1343
+ ` "sensor": "cartosat\_2s",`
1344
+
1345
+ ` "available\_bands": \["B1", "B2", "B3", "B4"\],`
1346
+
1347
+ ` "mapping": \{`
1348
+
1349
+ ` "B1": "c1",`
1350
+
1351
+ ` "B2": "c2",`
1352
+
1353
+ ` "B3": "c3",`
1354
+
1355
+ ` "B4": "c4"`
1356
+
1357
+ ` \}`
1358
+
1359
+ `\}`
1360
+ ```
1361
+
1362
+ No invented missing bands.
1363
+
1364
+ Missing channels become masked/zero-filled according to the validated adapter policy.
1365
+
1366
+
1367
+ # 19. RISAT Handling
1368
+
1369
+ RISAT imagery may vary in acquisition/polarization characteristics.
1370
+
1371
+ ISRO documentation describes RISAT-1 as a C-band SAR mission with multiple polarization configurations. Therefore the SAR adapter must inspect actual available channels rather than assuming one fixed polarization pair.
1372
+
1373
+ ### Internal representation
1374
+
1375
+ ```
1376
+ `canonical\_sar\[2, H, W\]`
1377
+
1378
+ `+`
1379
+
1380
+ `sar\_channel\_mask\[2\]`
1381
+ ```
1382
+
1383
+ Then CROMA receives both:
1384
+
1385
+ ```
1386
+ `SAR tensor`
1387
+
1388
+ `channel availability metadata`
1389
+ ```
1390
+
1391
+
1392
+ # 20. Sensor Robustness Training
1393
+
1394
+ During fusion-head training randomly mask inputs:
1395
+
1396
+ ```
1397
+ `optical:`
1398
+
1399
+ `100%`
1400
+
1401
+ `80%`
1402
+
1403
+ `60%`
1404
+
1405
+ `40%`
1406
+
1407
+
1408
+ `SAR:`
1409
+
1410
+ `100%`
1411
+
1412
+ `50%`
1413
+ ```
1414
+
1415
+ This creates robustness to modality differences and missing channels.
1416
+
1417
+ Do not synthesize fake SAR data.
1418
+
1419
+ Do not fabricate spectral bands.
1420
+
1421
+
1422
+ # 21. Multitemporal Change Architecture
1423
+
1424
+ ```
1425
+ `T1`
1426
+
1427
+ `T2`
1428
+
1429
+ ` ↓`
1430
+
1431
+ `validation`
1432
+
1433
+ ` ↓`
1434
+
1435
+ `CRS comparison`
1436
+
1437
+ ` ↓`
1438
+
1439
+ `resolution comparison`
1440
+
1441
+ ` ↓`
1442
+
1443
+ `co-registration check`
1444
+
1445
+ ` ↓`
1446
+
1447
+ `shared tiling`
1448
+
1449
+ ` ↓`
1450
+
1451
+ `shared encoder`
1452
+
1453
+ ` ↓`
1454
+
1455
+ `feature difference`
1456
+
1457
+ ` ↓`
1458
+
1459
+ `attention/change representation`
1460
+
1461
+ ` ↓`
1462
+
1463
+ `decoder`
1464
+
1465
+ ` ↓`
1466
+
1467
+ `probability map`
1468
+
1469
+ ` ↓`
1470
+
1471
+ `morphological cleanup`
1472
+
1473
+ ` ↓`
1474
+
1475
+ `connected components`
1476
+
1477
+ ` ↓`
1478
+
1479
+ `change regions`
1480
+
1481
+ ` ↓`
1482
+
1483
+ `SmolVLM interpretation`
1484
+ ```
1485
+
1486
+ STANet's official implementation provides a practical Siamese spatial-temporal attention approach for bitemporal remote-sensing change detection and the LEVIR-CD dataset workflow. ([GitHub](https://github.com/ChenDelong1999/RemoteCLIP?utm_source=chatgpt.com))
1487
+
1488
+
1489
+ # 22. Change Model
1490
+
1491
+ Use a compact ResNet18-class Siamese encoder.
1492
+
1493
+ Architecture:
1494
+
1495
+ ```
1496
+ `T1 ──► Encoder ──► F1`
1497
+
1498
+ ` β”‚`
1499
+
1500
+ `T2 ──► Encoder ──► F2`
1501
+
1502
+
1503
+ `F1 + F2`
1504
+
1505
+ ` ↓`
1506
+
1507
+ `difference/attention`
1508
+
1509
+ ` ↓`
1510
+
1511
+ `decoder`
1512
+
1513
+ ` ↓`
1514
+
1515
+ `change probability`
1516
+ ```
1517
+
1518
+ Shared encoder weights.
1519
+
1520
+ ### Loss
1521
+
1522
+ ```
1523
+ `0.5 BCE`
1524
+
1525
+ `+`
1526
+
1527
+ `0.5 Dice`
1528
+ ```
1529
+
1530
+ Tune around:
1531
+
1532
+ ```
1533
+ `0.25–0.75 BCE weighting`
1534
+ ```
1535
+
1536
+ on validation.
1537
+
1538
+
1539
+ # 23. False Change Handling
1540
+
1541
+ ### Illumination
1542
+
1543
+ Use learned features rather than raw pixel subtraction.
1544
+
1545
+ ### Registration
1546
+
1547
+ Calculate registration quality.
1548
+
1549
+ ### Season
1550
+
1551
+ Augment brightness/contrast/radiometric scaling during training.
1552
+
1553
+ ### Cloud/nodata
1554
+
1555
+ Maintain an invalid-data mask.
1556
+
1557
+ ### Sensor differences
1558
+
1559
+ Normalize each acquisition separately.
1560
+
1561
+ ### Low-quality alignment
1562
+
1563
+ Reduce confidence and optionally refuse spatial conclusions.
1564
+
1565
+ Never convert poor registration into fake certainty.
1566
+
1567
+
1568
+ # 24. Evidence System
1569
+
1570
+ The evidence engine is shared across all specialists.
1571
+
1572
+ ## Evidence types
1573
+
1574
+ ```
1575
+ `IMAGE\_CROP`
1576
+
1577
+ `TILE`
1578
+
1579
+ `BOUNDING\_BOX`
1580
+
1581
+ `MASK`
1582
+
1583
+ `CHANGE\_MAP`
1584
+
1585
+ `OPTICAL\_VIEW`
1586
+
1587
+ `SAR\_VIEW`
1588
+
1589
+ `JOINT\_FEATURE\_REGION`
1590
+
1591
+ `STATISTIC`
1592
+
1593
+ `GEOLOCATION`
1594
+ ```
1595
+
1596
+ Each evidence item has:
1597
+
1598
+ ```
1599
+ `\{`
1600
+
1601
+ ` "id": "evidence\_003",`
1602
+
1603
+ ` "type": "bounding\_box",`
1604
+
1605
+ ` "source": "grounding",`
1606
+
1607
+ ` "coordinates": \[0.21, 0.33, 0.47, 0.61\],`
1608
+
1609
+ ` "score": 0.87`
1610
+
1611
+ `\}`
1612
+ ```
1613
+
1614
+
1615
+ # 25. VLM and Evidence Relationship
1616
+
1617
+ Do not ask SmolVLM to hallucinate coordinates.
1618
+
1619
+ Instead:
1620
+
1621
+ ```
1622
+ `specialist`
1623
+
1624
+ ` ↓`
1625
+
1626
+ `real evidence`
1627
+
1628
+ ` ↓`
1629
+
1630
+ `VLM`
1631
+
1632
+ ` ↓`
1633
+
1634
+ `language explanation`
1635
+ ```
1636
+
1637
+ The VLM gets evidence references.
1638
+
1639
+ Example:
1640
+
1641
+ ```
1642
+ `Grounding model:`
1643
+
1644
+ `water\_region = \[x1,y1,x2,y2\]`
1645
+
1646
+ `confidence=0.91`
1647
+ ```
1648
+
1649
+ Then VLM says:
1650
+
1651
+ > The highlighted region corresponds to the water body.
1652
+
1653
+ The region came from a specialist.
1654
+
1655
+ The sentence came from the language model.
1656
+
1657
+ That distinction matters.
1658
+
1659
+
1660
+ # 26. Confidence System
1661
+
1662
+ No LLM-generated confidence.
1663
+
1664
+ ### Router confidence
1665
+
1666
+ Softmax probability.
1667
+
1668
+ ### Grounding
1669
+
1670
+ ```
1671
+ `box confidence`
1672
+
1673
+ `+`
1674
+
1675
+ `text/image similarity`
1676
+
1677
+ `+`
1678
+
1679
+ `augmentation consistency`
1680
+ ```
1681
+
1682
+ ### Change
1683
+
1684
+ ```
1685
+ `mean pixel probability`
1686
+
1687
+ `+`
1688
+
1689
+ `component stability`
1690
+
1691
+ `+`
1692
+
1693
+ `registration quality`
1694
+ ```
1695
+
1696
+ ### Optical-SAR
1697
+
1698
+ ```
1699
+ `fusion classifier margin`
1700
+
1701
+ `+`
1702
+
1703
+ `optical confidence`
1704
+
1705
+ `+`
1706
+
1707
+ `SAR confidence`
1708
+
1709
+ `+`
1710
+
1711
+ `cross-modal agreement`
1712
+ ```
1713
+
1714
+ Then calibration:
1715
+
1716
+ ```
1717
+ `raw score`
1718
+
1719
+ ` ↓`
1720
+
1721
+ `validation-calibrated mapping`
1722
+
1723
+ ` ↓`
1724
+
1725
+ `final confidence`
1726
+ ```
1727
+
1728
+ Use temperature scaling where appropriate.
1729
+
1730
+
1731
+ # 27. Execution Trace
1732
+
1733
+ ```
1734
+ `\{`
1735
+
1736
+ ` "run\_id": "uuid",`
1737
+
1738
+ ` "query": "...",`
1739
+
1740
+ ` "router": \{`
1741
+
1742
+ ` "task": "grounding",`
1743
+
1744
+ ` "confidence": 0.94`
1745
+
1746
+ ` \},`
1747
+
1748
+ ` "validation": \{`
1749
+
1750
+ ` "input\_count": 1,`
1751
+
1752
+ ` "format": "passed",`
1753
+
1754
+ ` "modality": "optical"`
1755
+
1756
+ ` \},`
1757
+
1758
+ ` "workflow": \[`
1759
+
1760
+ ` "validate",`
1761
+
1762
+ ` "preprocess",`
1763
+
1764
+ ` "ground",`
1765
+
1766
+ ` "evidence",`
1767
+
1768
+ ` "confidence",`
1769
+
1770
+ ` "answer"`
1771
+
1772
+ ` \],`
1773
+
1774
+ ` "models": \[`
1775
+
1776
+ ` \{`
1777
+
1778
+ ` "name": "RemoteCLIP",`
1779
+
1780
+ ` "revision": "..."`
1781
+
1782
+ ` \},`
1783
+
1784
+ ` \{`
1785
+
1786
+ ` "name": "SmolVLM",`
1787
+
1788
+ ` "revision": "..."`
1789
+
1790
+ ` \}`
1791
+
1792
+ ` \],`
1793
+
1794
+ ` "outputs": \[`
1795
+
1796
+ ` "bbox",`
1797
+
1798
+ ` "answer"`
1799
+
1800
+ ` \],`
1801
+
1802
+ ` "confidence": 0.87,`
1803
+
1804
+ ` "timings": \{\},`
1805
+
1806
+ ` "fallbacks": \[\]`
1807
+
1808
+ `\}`
1809
+ ```
1810
+
1811
+ No chain-of-thought.
1812
+
1813
+ Only observable execution facts.
1814
+
1815
+
1816
+ # 28. Natural-Language Router + Deterministic Orchestration
1817
+
1818
+ This is the answer to your original idea.
1819
+
1820
+ ### The router is allowed to answer:
1821
+
1822
+ ```
1823
+ `What does the user want?`
1824
+ ```
1825
+
1826
+ ### The controller answers:
1827
+
1828
+ ```
1829
+ `What inputs are valid?`
1830
+
1831
+ `Which workflow is legal?`
1832
+
1833
+ `Which specialists run?`
1834
+
1835
+ `In what order?`
1836
+
1837
+ `What parameters are permitted?`
1838
+
1839
+ `How are outputs combined?`
1840
+ ```
1841
+
1842
+ That means the system genuinely understands arbitrary phrasing without surrendering control.
1843
+
1844
+
1845
+ # 29. Example Query Routing
1846
+
1847
+ ### Query
1848
+
1849
+ > β€œCan you show me where the water body is?”
1850
+
1851
+ Router:
1852
+
1853
+ ```
1854
+ `\{`
1855
+
1856
+ ` "task": "grounding",`
1857
+
1858
+ ` "spatial\_output": true`
1859
+
1860
+ `\}`
1861
+ ```
1862
+
1863
+ Workflow:
1864
+
1865
+ ```
1866
+ `Grounding`
1867
+
1868
+ `β†’ box`
1869
+
1870
+ `β†’ evidence`
1871
+
1872
+ `β†’ VLM explanation`
1873
+ ```
1874
+
1875
+ ### Query
1876
+
1877
+ > β€œDescribe this satellite image.”
1878
+
1879
+ ```
1880
+ `caption`
1881
+ ```
1882
+
1883
+ ### Query
1884
+
1885
+ > οΏ½οΏ½What is the dominant land cover?”
1886
+
1887
+ ```
1888
+ `vqa`
1889
+ ```
1890
+
1891
+ ### Query
1892
+
1893
+ > β€œWhat changed?”
1894
+
1895
+ ```
1896
+ `change`
1897
+ ```
1898
+
1899
+ ### Query
1900
+
1901
+ > β€œCompare the radar and optical images to locate built-up areas.”
1902
+
1903
+ ```
1904
+ `optical\_sar`
1905
+ ```
1906
+
1907
+ ### Query
1908
+
1909
+ > β€œWhat changed and show me where?”
1910
+
1911
+ ```
1912
+ `change`
1913
+
1914
+ `+`
1915
+
1916
+ `spatial\_output=true`
1917
+ ```
1918
+
1919
+ One query can therefore result in a multi-step DAG.
1920
+
1921
+
1922
+ # 30. Dynamic Workflow DAG
1923
+
1924
+ The router outputs **intent attributes**, not a fixed tool call.
1925
+
1926
+ Example:
1927
+
1928
+ ```
1929
+ `\{`
1930
+
1931
+ ` "task": "change",`
1932
+
1933
+ ` "spatial\_output": true,`
1934
+
1935
+ ` "language\_output": true,`
1936
+
1937
+ ` "evidence\_required": true`
1938
+
1939
+ `\}`
1940
+ ```
1941
+
1942
+ The policy engine constructs:
1943
+
1944
+ ```
1945
+ `validate`
1946
+
1947
+ ` ↓`
1948
+
1949
+ `align`
1950
+
1951
+ ` ↓`
1952
+
1953
+ `change detector`
1954
+
1955
+ ` ↓`
1956
+
1957
+ `change regions`
1958
+
1959
+ ` ↓`
1960
+
1961
+ `VLM explanation`
1962
+
1963
+ ` ↓`
1964
+
1965
+ `confidence`
1966
+
1967
+ ` ↓`
1968
+
1969
+ `evidence`
1970
+ ```
1971
+
1972
+ This is more powerful than static routing while remaining deterministic.
1973
+
1974
+
1975
+ # 31. Dataset Engineering
1976
+
1977
+ ## BigEarthNet v2
1978
+
1979
+ Use for:
1980
+
1981
+ - remote-sensing adaptation
1982
+
1983
+ - land-cover understanding
1984
+
1985
+ - optical/SAR paired learning
1986
+
1987
+ - robustness experiments
1988
+
1989
+ The official BigEarthNet site reports 549,488 Sentinel-1/Sentinel-2 patch pairs in v2.0.
1990
+
1991
+ ### Subset
1992
+
1993
+ Start:
1994
+
1995
+ ```
1996
+ `50,000 training`
1997
+
1998
+ `5,000 validation`
1999
+ ```
2000
+
2001
+ Then scale if the Kaggle budget permits.
2002
+
2003
+
2004
+ # 32. VRSBench
2005
+
2006
+ Use for:
2007
+
2008
+ ```
2009
+ `captioning`
2010
+
2011
+ `VQA`
2012
+
2013
+ `grounding`
2014
+ ```
2015
+
2016
+ Current repository information reports:
2017
+
2018
+ ```
2019
+ `29,614 images`
2020
+
2021
+ `29,614 human-verified captions`
2022
+
2023
+ `52,472 object references`
2024
+
2025
+ `123,221 VQA pairs`
2026
+ ```
2027
+
2028
+ and provides evaluation code. ([GitHub](https://github.com/lx709/VRSBench?utm_source=chatgpt.com))
2029
+
2030
+
2031
+ # 33. RSVQA
2032
+
2033
+ Use for:
2034
+
2035
+ ```
2036
+ `single-image VQA`
2037
+ ```
2038
+
2039
+ Keep benchmark-specific preprocessing separate from generic EO preprocessing.
2040
+
2041
+ Never mix benchmark-specific conventions into the common image loader.
2042
+
2043
+
2044
+ # 34. CDVQA
2045
+
2046
+ Use for:
2047
+
2048
+ ```
2049
+ `change-based VQA`
2050
+
2051
+ `change description`
2052
+
2053
+ `temporal reasoning`
2054
+ ```
2055
+
2056
+ Keep its annotation schema in an adapter.
2057
+
2058
+
2059
+ # 35. LEVIR-CD
2060
+
2061
+ Use for:
2062
+
2063
+ ```
2064
+ `binary change-mask training`
2065
+ ```
2066
+
2067
+ Training patch size:
2068
+
2069
+ ```
2070
+ `256Γ—256`
2071
+ ```
2072
+
2073
+ in the standard STANet workflow.
2074
+
2075
+
2076
+ # 36. Data Leakage Prevention
2077
+
2078
+ This remains one of the hardest requirements.
2079
+
2080
+ ## Every sample gets:
2081
+
2082
+ ```
2083
+ `dataset\_id`
2084
+
2085
+ `scene\_id`
2086
+
2087
+ `sample\_id`
2088
+
2089
+ `source\_scene`
2090
+
2091
+ `geographic\_hash`
2092
+
2093
+ `acquisition\_date`
2094
+
2095
+ `sensor`
2096
+
2097
+ `sha256`
2098
+
2099
+ `split`
2100
+ ```
2101
+
2102
+ ### Split by scene
2103
+
2104
+ Never randomly split neighboring patches.
2105
+
2106
+ ```
2107
+ `scene`
2108
+
2109
+ ` ↓`
2110
+
2111
+ `train OR validation OR test`
2112
+ ```
2113
+
2114
+ not:
2115
+
2116
+ ```
2117
+ `patch`
2118
+
2119
+ ` ↓`
2120
+
2121
+ `random split`
2122
+ ```
2123
+
2124
+
2125
+ # 37. Test Set Firewall
2126
+
2127
+ Public test sets:
2128
+
2129
+ ```
2130
+ `evaluation/public\_test/`
2131
+ ```
2132
+
2133
+ are immutable.
2134
+
2135
+ Training code may not import them.
2136
+
2137
+ Evaluation code may read them only in:
2138
+
2139
+ ```
2140
+ `evaluation mode`
2141
+ ```
2142
+
2143
+ No cache of test answers is permitted.
2144
+
2145
+
2146
+ # 38. Hidden Set Firewall
2147
+
2148
+ There must be no:
2149
+
2150
+ ```
2151
+ `hidden\_test\_mode`
2152
+ ```
2153
+
2154
+ in training.
2155
+
2156
+ The hidden evaluation interface accepts:
2157
+
2158
+ ```
2159
+ `input imagery`
2160
+ ```
2161
+
2162
+ and produces:
2163
+
2164
+ ```
2165
+ `standard result schema`
2166
+ ```
2167
+
2168
+ with no knowledge of the hidden answer.
2169
+
2170
+
2171
+ # 39. GeoTIFF Contract
2172
+
2173
+ Validation:
2174
+
2175
+ ```
2176
+ `file`
2177
+
2178
+ ` ↓`
2179
+
2180
+ `TIFF/GeoTIFF parser`
2181
+
2182
+ ` ↓`
2183
+
2184
+ `dimensions`
2185
+
2186
+ ` ↓`
2187
+
2188
+ `bands`
2189
+
2190
+ ` ↓`
2191
+
2192
+ `dtype`
2193
+
2194
+ ` ↓`
2195
+
2196
+ `CRS`
2197
+
2198
+ ` ↓`
2199
+
2200
+ `transform`
2201
+
2202
+ ` ↓`
2203
+
2204
+ `bounds`
2205
+
2206
+ ` ↓`
2207
+
2208
+ `nodata`
2209
+
2210
+ ` ↓`
2211
+
2212
+ `modality`
2213
+
2214
+ ` ↓`
2215
+
2216
+ `temporal metadata`
2217
+
2218
+ ` ↓`
2219
+
2220
+ `pair compatibility`
2221
+ ```
2222
+
2223
+
2224
+ # 40. Pair Compatibility
2225
+
2226
+ For optical-SAR:
2227
+
2228
+ ```
2229
+ `CRS compatible?`
2230
+
2231
+ `bounds overlap?`
2232
+
2233
+ `dimensions sensible?`
2234
+
2235
+ `resolution known?`
2236
+
2237
+ `co-registration metadata?`
2238
+ ```
2239
+
2240
+ For temporal:
2241
+
2242
+ ```
2243
+ `T1 != T2`
2244
+
2245
+ `same geographic scene?`
2246
+
2247
+ `overlap?`
2248
+
2249
+ `resolution compatibility?`
2250
+
2251
+ `alignment quality?`
2252
+ ```
2253
+
2254
+
2255
+ # 41. Preprocessing Configuration
2256
+
2257
+ ```
2258
+ `image:`
2259
+
2260
+ ` max\_pixels: 25000000`
2261
+
2262
+ ` tile\_size: 512`
2263
+
2264
+ ` tile\_overlap: 128`
2265
+
2266
+ ` max\_tiles: 64`
2267
+
2268
+
2269
+ `optical:`
2270
+
2271
+ ` normalization: percentile`
2272
+
2273
+ ` lower\_percentile: 2`
2274
+
2275
+ ` upper\_percentile: 98`
2276
+
2277
+
2278
+ `sar:`
2279
+
2280
+ ` representation: db`
2281
+
2282
+ ` clip\_min\_db: -30`
2283
+
2284
+ ` clip\_max\_db: 5`
2285
+
2286
+
2287
+ `change:`
2288
+
2289
+ ` tile\_size: 256`
2290
+
2291
+ ` tile\_overlap: 32`
2292
+
2293
+
2294
+ `grounding:`
2295
+
2296
+ ` max\_candidates: 20`
2297
+
2298
+ ` nms\_iou: 0.5`
2299
+ ```
2300
+
2301
+
2302
+ # 42. Large Image Handling
2303
+
2304
+ ```
2305
+ `large image`
2306
+
2307
+ ` ↓`
2308
+
2309
+ `overview`
2310
+
2311
+ ` ↓`
2312
+
2313
+ `candidate tile scoring`
2314
+
2315
+ ` ↓`
2316
+
2317
+ `top 4 tiles`
2318
+
2319
+ ` ↓`
2320
+
2321
+ `specialist inference`
2322
+ ```
2323
+
2324
+ Candidate scoring can use cheap visual statistics:
2325
+
2326
+ ```
2327
+ `edge density`
2328
+
2329
+ `entropy`
2330
+
2331
+ `CLIP similarity`
2332
+
2333
+ `objectness`
2334
+ ```
2335
+
2336
+ No VLM inference for every tile.
2337
+
2338
+
2339
+ # 43. Remote-Sensing Adaptation
2340
+
2341
+ The official project requires adaptation using BigEarthNet or other allowed remote-sensing data.
2342
+
2343
+ ### VLM adaptation
2344
+
2345
+ Freeze:
2346
+
2347
+ ```
2348
+ `vision backbone`
2349
+
2350
+ `most language weights`
2351
+ ```
2352
+
2353
+ Train:
2354
+
2355
+ ```
2356
+ `LoRA adapters`
2357
+ ```
2358
+
2359
+ ### Initial configuration
2360
+
2361
+ ```
2362
+ `lora:`
2363
+
2364
+ ` rank: 16`
2365
+
2366
+ ` alpha: 32`
2367
+
2368
+ ` dropout: 0.05`
2369
+
2370
+
2371
+ `training:`
2372
+
2373
+ ` lr: 2e-4`
2374
+
2375
+ ` effective\_batch\_size: 16`
2376
+
2377
+ ` epochs: 1`
2378
+
2379
+ ` warmup\_ratio: 0.05`
2380
+
2381
+ ` weight\_decay: 0.01`
2382
+
2383
+ ` precision: bf16`
2384
+ ```
2385
+
2386
+
2387
+ # 44. VLM Training Data Construction
2388
+
2389
+ From BigEarthNet labels create controlled instruction pairs.
2390
+
2391
+ Example:
2392
+
2393
+ ```
2394
+ `Image:`
2395
+
2396
+ `BigEarthNet patch`
2397
+
2398
+
2399
+ `Question:`
2400
+
2401
+ `"What land-cover categories are visible?"`
2402
+
2403
+
2404
+ `Answer:`
2405
+
2406
+ `"Arable land, mixed forest and urban fabric."`
2407
+ ```
2408
+
2409
+ Also:
2410
+
2411
+ ```
2412
+ `"Is urban land cover visible?"`
2413
+
2414
+ `"Is water present?"`
2415
+
2416
+ `"Which category dominates?"`
2417
+
2418
+ `"Describe the scene using land-cover labels."`
2419
+ ```
2420
+
2421
+ Do not manufacture precise object counts that the source annotations do not support.
2422
+
2423
+
2424
+ # 45. Kaggle Budget
2425
+
2426
+ Current Kaggle documentation says free notebook GPU use includes T4Γ—2, and the P100 has now been retired effective September 15, 2026. Kaggle currently documents 12-hour CPU/GPU notebook sessions, and its efficient-GPU documentation describes a weekly GPU quota around 30 hours, subject to availability. ([Kaggle](https://www.kaggle.com/docs/notebooks?utm_source=chatgpt.com))
2427
+
2428
+ Therefore the plan must now assume:
2429
+
2430
+ **T4Γ—2 first.**
2431
+
2432
+ Not P100.
2433
+
2434
+ That is a material update from the earlier plan.
2435
+
2436
+
2437
+ # 46. Kaggle Training Schedule
2438
+
2439
+ | **Stage** | **Accelerator** | **Training** | **Target** |
2440
+ | :-: | :-: | :-: | :-: |
2441
+ | Router | CPU/T4 | ~2–3k queries | \<1 h |
2442
+ | VLM adaptation | T4Γ—2 | 40–60k samples | ≀8 h |
2443
+ | Grounding head | T4Γ—2 | VRSBench refs | ≀4 h |
2444
+ | Change detector | T4Γ—2 | LEVIR-CD | ≀5 h |
2445
+ | Change VQA | T4Γ—2 | CDVQA subset | ≀3 h |
2446
+ | CROMA fusion head | T4Γ—2 | BigEarthNet pair subset | ≀3 h |
2447
+ | Calibration | CPU | validation | \<1 h |
2448
+
2449
+ These are **budget targets**, not measured benchmark timings.
2450
+
2451
+ Every actual training run records:
2452
+
2453
+ ```
2454
+ `wall\_time`
2455
+
2456
+ `peak\_VRAM`
2457
+
2458
+ `peak\_RAM`
2459
+
2460
+ `samples\_per\_second`
2461
+
2462
+ `checkpoint\_size`
2463
+ ```
2464
+
2465
+
2466
+ # 47. Training Contingency
2467
+
2468
+ If a stage exceeds budget:
2469
+
2470
+ ```
2471
+ `1. reduce dataset subset`
2472
+
2473
+ `2. freeze additional weights`
2474
+
2475
+ `3. lower resolution`
2476
+
2477
+ `4. reduce batch`
2478
+
2479
+ `5. increase gradient accumulation`
2480
+
2481
+ `6. reduce LoRA rank`
2482
+
2483
+ `7. reduce epochs`
2484
+ ```
2485
+
2486
+ Never increase architectural complexity as the first response to a compute problem.
2487
+
2488
+
2489
+ # 48. Hugging Face Deployment
2490
+
2491
+ Hugging Face currently documents ZeroGPU Spaces as Gradio-only, with free personal accounts able to host up to two ZeroGPU Spaces, and a current free quota of 5 GPU minutes/day. The default `large` configuration provides 48 GB VRAM. ([Hugging Face](https://huggingface.co/docs/hub/spaces-zerogpu?utm_source=chatgpt.com))
2492
+
2493
+ For the competition prototype:
2494
+
2495
+ ```
2496
+ `One Gradio Space`
2497
+ ```
2498
+
2499
+ with:
2500
+
2501
+ ```
2502
+ `CPU preprocessing`
2503
+
2504
+ `+`
2505
+
2506
+ `GPU specialist execution when available`
2507
+ ```
2508
+
2509
+ ### Important
2510
+
2511
+ ZeroGPU is an accelerator.
2512
+
2513
+ It is not our architectural dependency.
2514
+
2515
+
2516
+ # 49. Lazy Model Loading
2517
+
2518
+ Use a simple model manager:
2519
+
2520
+ ```
2521
+ `request`
2522
+
2523
+ ` ↓`
2524
+
2525
+ `determine specialists`
2526
+
2527
+ ` ↓`
2528
+
2529
+ `load only those specialists`
2530
+
2531
+ ` ↓`
2532
+
2533
+ `execute`
2534
+
2535
+ ` ↓`
2536
+
2537
+ `release unused models`
2538
+ ```
2539
+
2540
+ Do not permanently place every model on GPU.
2541
+
2542
+
2543
+ # 50. Deployment Model Footprint Strategy
2544
+
2545
+ Keep:
2546
+
2547
+ ```
2548
+ `SmolVLM`
2549
+
2550
+ `RemoteCLIP ViT-B/32`
2551
+
2552
+ `STANet`
2553
+
2554
+ `CROMA-base`
2555
+
2556
+ `MiniLM router`
2557
+ ```
2558
+
2559
+ but load them according to workflow.
2560
+
2561
+ ### VQA
2562
+
2563
+ ```
2564
+ `MiniLM`
2565
+
2566
+ `SmolVLM`
2567
+ ```
2568
+
2569
+ ### Grounding
2570
+
2571
+ ```
2572
+ `MiniLM`
2573
+
2574
+ `RemoteCLIP`
2575
+
2576
+ `SmolVLM`
2577
+ ```
2578
+
2579
+ ### Change
2580
+
2581
+ ```
2582
+ `MiniLM`
2583
+
2584
+ `STANet`
2585
+
2586
+ `SmolVLM`
2587
+ ```
2588
+
2589
+ ### Optical-SAR
2590
+
2591
+ ```
2592
+ `MiniLM`
2593
+
2594
+ `CROMA`
2595
+
2596
+ `fusion head`
2597
+
2598
+ `SmolVLM`
2599
+ ```
2600
+
2601
+ This is why the modular specialist architecture matters.
2602
+
2603
+
2604
+ # 51. GUI
2605
+
2606
+ ## Header
2607
+
2608
+ ```
2609
+ `SatQuery AI`
2610
+
2611
+ `Interactive Remote-Sensing Intelligence`
2612
+ ```
2613
+
2614
+ ### Upload area
2615
+
2616
+ ```
2617
+ `Single Image`
2618
+
2619
+ `T1 / T2`
2620
+
2621
+ `Optical`
2622
+
2623
+ `SAR`
2624
+ ```
2625
+
2626
+ ### Viewer
2627
+
2628
+ Tabs:
2629
+
2630
+ ```
2631
+ `Original`
2632
+
2633
+ `Evidence`
2634
+
2635
+ `Grounding`
2636
+
2637
+ `Change`
2638
+
2639
+ `Optical`
2640
+
2641
+ `SAR`
2642
+
2643
+ `Fusion`
2644
+ ```
2645
+
2646
+ ### Query
2647
+
2648
+ ```
2649
+ `\[ Ask SatQuery AI... \]`
2650
+
2651
+
2652
+ `Detected task:`
2653
+
2654
+ `Optical + SAR`
2655
+
2656
+
2657
+ `Running:`
2658
+
2659
+ `βœ“ validation`
2660
+
2661
+ `βœ“ optical encoder`
2662
+
2663
+ `βœ“ SAR encoder`
2664
+
2665
+ `β†’ fusion`
2666
+ ```
2667
+
2668
+
2669
+ # 52. Result View
2670
+
2671
+ ```
2672
+ `Answer`
2673
+
2674
+ `--------------------------------`
2675
+
2676
+ `Built-up regions are concentrated`
2677
+
2678
+ `in the eastern portion of the image.`
2679
+
2680
+
2681
+ `Confidence`
2682
+
2683
+ `72%`
2684
+
2685
+
2686
+ `Evidence`
2687
+
2688
+ `\[highlighted image\]`
2689
+
2690
+
2691
+ `Optical evidence`
2692
+
2693
+ `\[SAR evidence\]`
2694
+
2695
+
2696
+ `Execution`
2697
+
2698
+ `\[details\]`
2699
+
2700
+
2701
+ `Download`
2702
+
2703
+ `\[JSON\] \[PDF\]`
2704
+ ```
2705
+
2706
+ The evaluator should be able to discover the entire architecture by using the application.
2707
+
2708
+
2709
+ # 53. Result Schema
2710
+
2711
+ Final schema:
2712
+
2713
+ ```
2714
+ `\{`
2715
+
2716
+ ` "task": "vqa|caption|grounding|change|optical\_sar",`
2717
+
2718
+ ` "answer": "...",`
2719
+
2720
+ ` "labels": \[\],`
2721
+
2722
+ ` "regions": \[\],`
2723
+
2724
+ ` "boxes": \[\],`
2725
+
2726
+ ` "masks": \[\],`
2727
+
2728
+ ` "change\_map": null,`
2729
+
2730
+ ` "evidence": \[\],`
2731
+
2732
+ ` "confidence": \{\},`
2733
+
2734
+ ` "geospatial": \{\},`
2735
+
2736
+ ` "execution\_trace": \{\}`
2737
+
2738
+ `\}`
2739
+ ```
2740
+
2741
+ The supplied project schema already establishes the core fields. Grounding adds `regions` to make spatial outputs first-class.
2742
+
2743
+
2744
+ # 54. API Contracts
2745
+
2746
+ ## Query
2747
+
2748
+ ```
2749
+ `\{`
2750
+
2751
+ ` "assets": \["asset\_001"\],`
2752
+
2753
+ ` "query": "Show the water body."`
2754
+
2755
+ `\}`
2756
+ ```
2757
+
2758
+ ## Intent
2759
+
2760
+ ```
2761
+ `\{`
2762
+
2763
+ ` "task": "grounding",`
2764
+
2765
+ ` "modality": "optical",`
2766
+
2767
+ ` "spatial\_output": true,`
2768
+
2769
+ ` "confidence": 0.92`
2770
+
2771
+ `\}`
2772
+ ```
2773
+
2774
+ ## Analysis
2775
+
2776
+ ```
2777
+ `\{`
2778
+
2779
+ ` "workflow": "grounding",`
2780
+
2781
+ ` "assets": \["asset\_001"\],`
2782
+
2783
+ ` "config\_hash": "..."`
2784
+
2785
+ `\}`
2786
+ ```
2787
+
2788
+ ## Result
2789
+
2790
+ Use the master result schema.
2791
+
2792
+ ## Trace
2793
+
2794
+ Use the execution trace schema.
2795
+
2796
+ ## Health
2797
+
2798
+ ```
2799
+ `\{`
2800
+
2801
+ ` "status": "ok",`
2802
+
2803
+ ` "models": \{`
2804
+
2805
+ ` "vlm": "ready",`
2806
+
2807
+ ` "grounding": "ready",`
2808
+
2809
+ ` "change": "ready",`
2810
+
2811
+ ` "fusion": "ready"`
2812
+
2813
+ ` \}`
2814
+
2815
+ `\}`
2816
+ ```
2817
+
2818
+
2819
+ # 55. Repository Structure
2820
+
2821
+ ```
2822
+ `satquery-ai/`
2823
+
2824
+ `β”‚`
2825
+
2826
+ `β”œβ”€β”€ app/`
2827
+
2828
+ `β”‚ β”œβ”€β”€ gradio\_app.py`
2829
+
2830
+ `β”‚ β”œβ”€β”€ ui\_state.py`
2831
+
2832
+ `β”‚ └── visualizers.py`
2833
+
2834
+ `β”‚`
2835
+
2836
+ `β”œβ”€β”€ core/`
2837
+
2838
+ `β”‚ β”œβ”€β”€ config.py`
2839
+
2840
+ `β”‚ β”œβ”€β”€ schemas.py`
2841
+
2842
+ `β”‚ β”œβ”€β”€ controller.py`
2843
+
2844
+ `β”‚ β”œβ”€β”€ planner.py`
2845
+
2846
+ `β”‚ β”œβ”€β”€ registry.py`
2847
+
2848
+ `β”‚ └── errors.py`
2849
+
2850
+ `β”‚`
2851
+
2852
+ `β”œβ”€β”€ router/`
2853
+
2854
+ `β”‚ β”œβ”€β”€ encoder.py`
2855
+
2856
+ `β”‚ β”œβ”€β”€ classifier.py`
2857
+
2858
+ `β”‚ β”œβ”€β”€ adapter.py`
2859
+
2860
+ `β”‚ β”œβ”€β”€ train.py`
2861
+
2862
+ `β”‚ └── dataset.py`
2863
+
2864
+ `β”‚`
2865
+
2866
+ `β”œβ”€β”€ specialists/`
2867
+
2868
+ `β”‚ β”œβ”€β”€ base.py`
2869
+
2870
+ `β”‚ β”‚`
2871
+
2872
+ `β”‚ β”œβ”€β”€ vqa/`
2873
+
2874
+ `β”‚ β”‚ β”œβ”€β”€ model.py`
2875
+
2876
+ `β”‚ β”‚ β”œβ”€β”€ inference.py`
2877
+
2878
+ `β”‚ β”‚ └── prompts.py`
2879
+
2880
+ `β”‚ β”‚`
2881
+
2882
+ `β”‚ β”œβ”€β”€ grounding/`
2883
+
2884
+ `β”‚ β”‚ β”œβ”€β”€ remoteclip.py`
2885
+
2886
+ `β”‚ β”‚ β”œβ”€β”€ head.py`
2887
+
2888
+ `β”‚ β”‚ β”œβ”€β”€ inference.py`
2889
+
2890
+ `β”‚ β”‚ └── postprocess.py`
2891
+
2892
+ `β”‚ β”‚`
2893
+
2894
+ `β”‚ β”œβ”€β”€ change/`
2895
+
2896
+ `β”‚ β”‚ β”œβ”€β”€ stanet.py`
2897
+
2898
+ `β”‚ β”‚ β”œβ”€β”€ inference.py`
2899
+
2900
+ `β”‚ β”‚ └── postprocess.py`
2901
+
2902
+ `β”‚ β”‚`
2903
+
2904
+ `β”‚ └── optical\_sar/`
2905
+
2906
+ `β”‚ β”œβ”€β”€ croma.py`
2907
+
2908
+ `β”‚ β”œβ”€β”€ fusion\_head.py`
2909
+
2910
+ `β”‚ └── inference.py`
2911
+
2912
+ `β”‚`
2913
+
2914
+ `β”œβ”€β”€ preprocessing/`
2915
+
2916
+ `β”‚ β”œβ”€β”€ raster.py`
2917
+
2918
+ `β”‚ β”œβ”€β”€ optical.py`
2919
+
2920
+ `β”‚ β”œβ”€β”€ sar.py`
2921
+
2922
+ `β”‚ β”œβ”€β”€ tiling.py`
2923
+
2924
+ `β”‚ β”œβ”€β”€ temporal.py`
2925
+
2926
+ `β”‚ └── sensor\_adapter.py`
2927
+
2928
+ `β”‚`
2929
+
2930
+ `β”œβ”€β”€ geospatial/`
2931
+
2932
+ `β”‚ β”œβ”€β”€ crs.py`
2933
+
2934
+ `β”‚ β”œβ”€β”€ alignment.py`
2935
+
2936
+ `β”‚ β”œβ”€β”€ transform.py`
2937
+
2938
+ `β”‚ └── overlap.py`
2939
+
2940
+ `β”‚`
2941
+
2942
+ `β”œβ”€β”€ evidence/`
2943
+
2944
+ `β”‚ β”œβ”€β”€ boxes.py`
2945
+
2946
+ `β”‚ β”œβ”€β”€ masks.py`
2947
+
2948
+ `β”‚ β”œβ”€β”€ crops.py`
2949
+
2950
+ `β”‚ β”œβ”€β”€ change\_maps.py`
2951
+
2952
+ `β”‚ └── confidence.py`
2953
+
2954
+ `β”‚`
2955
+
2956
+ `β”œβ”€β”€ reports/`
2957
+
2958
+ `β”‚ └── generator.py`
2959
+
2960
+ `β”‚`
2961
+
2962
+ `β”œβ”€β”€ evaluation/`
2963
+
2964
+ `β”‚ β”œβ”€β”€ manifests.py`
2965
+
2966
+ `β”‚ β”œβ”€β”€ leakage.py`
2967
+
2968
+ `β”‚ β”œβ”€β”€ metrics/`
2969
+
2970
+ `β”‚ β”œβ”€β”€ normalize.py`
2971
+
2972
+ `β”‚ β”œβ”€β”€ runner.py`
2973
+
2974
+ `β”‚ └── benchmark\_adapters/`
2975
+
2976
+ `β”‚`
2977
+
2978
+ `β”œβ”€β”€ training/`
2979
+
2980
+ `β”‚ β”œβ”€β”€ vlm/`
2981
+
2982
+ `β”‚ β”œβ”€β”€ grounding/`
2983
+
2984
+ `β”‚ β”œβ”€β”€ change/`
2985
+
2986
+ `β”‚ β”œβ”€β”€ fusion/`
2987
+
2988
+ `β”‚ └── calibration/`
2989
+
2990
+ `β”‚`
2991
+
2992
+ `β”œβ”€β”€ configs/`
2993
+
2994
+ `β”‚ β”œβ”€β”€ base.yaml`
2995
+
2996
+ `β”‚ β”œβ”€β”€ train.yaml`
2997
+
2998
+ `β”‚ β”œβ”€β”€ eval.yaml`
2999
+
3000
+ `β”‚ └── deploy.yaml`
3001
+
3002
+ `β”‚`
3003
+
3004
+ `β”œβ”€β”€ tests/`
3005
+
3006
+ `β”‚ β”œβ”€β”€ unit/`
3007
+
3008
+ `β”‚ β”œβ”€β”€ routing/`
3009
+
3010
+ `β”‚ β”œβ”€β”€ geospatial/`
3011
+
3012
+ `β”‚ β”œβ”€β”€ leakage/`
3013
+
3014
+ `β”‚ β”œβ”€β”€ model/`
3015
+
3016
+ `β”‚ └── e2e/`
3017
+
3018
+ `β”‚`
3019
+
3020
+ `β”œβ”€β”€ scripts/`
3021
+
3022
+ `β”‚ β”œβ”€β”€ prepare\_data.py`
3023
+
3024
+ `β”‚ β”œβ”€β”€ create\_manifest.py`
3025
+
3026
+ `β”‚ β”œβ”€β”€ train\_router.py`
3027
+
3028
+ `β”‚ β”œβ”€β”€ train\_vlm.py`
3029
+
3030
+ `β”‚ β”œβ”€β”€ train\_grounding.py`
3031
+
3032
+ `β”‚ β”œβ”€β”€ train\_change.py`
3033
+
3034
+ `β”‚ β”œβ”€β”€ train\_fusion.py`
3035
+
3036
+ `β”‚ └── evaluate.py`
3037
+
3038
+ `β”‚`
3039
+
3040
+ `└── README.md`
3041
+ ```
3042
+
3043
+
3044
+ # 56. Central Configuration Registry
3045
+
3046
+ ```
3047
+ `project:`
3048
+
3049
+ ` name: satquery-ai`
3050
+
3051
+ ` version: "1.0.0"`
3052
+
3053
+ ` seed: 42`
3054
+
3055
+
3056
+ `image:`
3057
+
3058
+ ` max\_pixels: 25000000`
3059
+
3060
+ ` tile\_size: 512`
3061
+
3062
+ ` tile\_overlap: 128`
3063
+
3064
+ ` max\_tiles: 64`
3065
+
3066
+
3067
+ `router:`
3068
+
3069
+ ` model: sentence-transformers/all-MiniLM-L6-v2`
3070
+
3071
+ ` max\_length: 128`
3072
+
3073
+ ` hidden\_dim: 128`
3074
+
3075
+ ` dropout: 0.10`
3076
+
3077
+ ` confidence\_threshold: 0.70`
3078
+
3079
+
3080
+ `vlm:`
3081
+
3082
+ ` checkpoint: HuggingFaceTB/SmolVLM-500M-Instruct`
3083
+
3084
+ ` max\_new\_tokens: 128`
3085
+
3086
+ ` temperature: 0.0`
3087
+
3088
+
3089
+ `grounding:`
3090
+
3091
+ ` checkpoint: chendelong/RemoteCLIP`
3092
+
3093
+ ` variant: ViT-B-32`
3094
+
3095
+ ` nms\_iou: 0.50`
3096
+
3097
+ ` max\_candidates: 20`
3098
+
3099
+ ` confidence\_threshold: 0.40`
3100
+
3101
+
3102
+ `change:`
3103
+
3104
+ ` tile\_size: 256`
3105
+
3106
+ ` tile\_overlap: 32`
3107
+
3108
+ ` threshold: 0.50`
3109
+
3110
+ ` min\_component\_pixels: 32`
3111
+
3112
+
3113
+ `croma:`
3114
+
3115
+ ` checkpoint: antofuller/CROMA`
3116
+
3117
+ ` variant: base`
3118
+
3119
+ ` image\_resolution: 120`
3120
+
3121
+ ` optical\_channels: 12`
3122
+
3123
+ ` sar\_channels: 2`
3124
+
3125
+
3126
+ `optical:`
3127
+
3128
+ ` normalization: percentile`
3129
+
3130
+ ` lower\_percentile: 2`
3131
+
3132
+ ` upper\_percentile: 98`
3133
+
3134
+
3135
+ `sar:`
3136
+
3137
+ ` representation: db`
3138
+
3139
+ ` clip\_min\_db: -30`
3140
+
3141
+ ` clip\_max\_db: 5`
3142
+
3143
+
3144
+ `training:`
3145
+
3146
+ ` precision: bf16`
3147
+
3148
+ ` vlm\_lr: 0.0002`
3149
+
3150
+ ` vlm\_batch\_size: 2`
3151
+
3152
+ ` vlm\_gradient\_accumulation: 8`
3153
+
3154
+ ` lora\_rank: 16`
3155
+
3156
+ ` lora\_alpha: 32`
3157
+
3158
+ ` lora\_dropout: 0.05`
3159
+
3160
+ ` weight\_decay: 0.01`
3161
+
3162
+ ` warmup\_ratio: 0.05`
3163
+
3164
+ ` epochs: 1`
3165
+
3166
+
3167
+ `grounding\_training:`
3168
+
3169
+ ` learning\_rate: 0.0001`
3170
+
3171
+ ` box\_loss\_weight: 0.5`
3172
+
3173
+ ` giou\_loss\_weight: 0.3`
3174
+
3175
+ ` confidence\_loss\_weight: 0.2`
3176
+
3177
+
3178
+ `confidence:`
3179
+
3180
+ ` temperature\_scaling: true`
3181
+
3182
+
3183
+ `runtime:`
3184
+
3185
+ ` max\_specialists: 4`
3186
+
3187
+ ` timeout\_seconds: 120`
3188
+
3189
+ ` unload\_after\_workflow: true`
3190
+
3191
+
3192
+ `evaluation:`
3193
+
3194
+ ` immutable\_public\_test: true`
3195
+
3196
+ ` hidden\_data\_access: false`
3197
+
3198
+ ` official\_aggregate\_weights: null`
3199
+ ```
3200
+
3201
+
3202
+ # 57. Failure Matrix
3203
+
3204
+ | **Failure** | **Detection** | **Recovery** | **User message** | **Trace** | **Fallback** |
3205
+ | :-: | :-: | :-: | :-: | :-: | :-: |
3206
+ | Router low confidence | probability threshold | ask classification fallback rules | β€œI’m not confident what task you requested.” | yes | deterministic keyword layer |
3207
+ | Router predicts invalid task | schema validator | reject | β€œThe uploaded inputs do not support this task.” | yes | none |
3208
+ | Corrupt TIFF | rasterio | reject | unreadable | yes | none |
3209
+ | Missing CRS | metadata parser | degraded | CRS missing | yes | non-geospatial mode |
3210
+ | Pair misalignment | alignment test | reject spatial analysis | images not sufficiently aligned | yes | textual only if safe |
3211
+ | Grounding head fails | runtime | retry CPU/GPU | grounding unavailable | yes | VQA answer without box |
3212
+ | Change model fails | runtime | fallback raw feature difference | change detector unavailable | yes | reduced mode |
3213
+ | CROMA unavailable | runtime | modality-specific analysis | joint fusion unavailable | yes | optical/SAR separate |
3214
+ | VLM OOM | exception | lower resolution | reduced reasoning mode | yes | template result |
3215
+ | Timeout | timer | abort specialist | processing timeout | yes | partial result |
3216
+ | Low confidence | calibration | cautious answer | low confidence | yes | uncertainty text |
3217
+
3218
+
3219
+ # 58. Fallback Router
3220
+
3221
+ Do not make the system entirely dependent on the trained intent model.
3222
+
3223
+ Use:
3224
+
3225
+ ```
3226
+ `Tiny NLP router`
3227
+
3228
+ ` ↓`
3229
+
3230
+ `if confidence β‰₯ threshold`
3231
+
3232
+ ` accept`
3233
+
3234
+ `else`
3235
+
3236
+ ` deterministic lexical fallback`
3237
+ ```
3238
+
3239
+ Examples:
3240
+
3241
+ ```
3242
+ `"where"`
3243
+
3244
+ `"locate"`
3245
+
3246
+ `"highlight"`
3247
+
3248
+ `"show region"`
3249
+ ```
3250
+
3251
+ β†’ grounding
3252
+
3253
+ ```
3254
+ `"changed"`
3255
+
3256
+ `"between"`
3257
+
3258
+ `"before and after"`
3259
+
3260
+ `"temporal"`
3261
+ ```
3262
+
3263
+ β†’ change
3264
+
3265
+ ```
3266
+ `"optical and SAR"`
3267
+
3268
+ `"radar and optical"`
3269
+
3270
+ `"both images"`
3271
+ ```
3272
+
3273
+ β†’ optical-SAR
3274
+
3275
+ This does **not** replace the learned router.
3276
+
3277
+ It protects against obvious failures.
3278
+
3279
+
3280
+ # 59. Testing
3281
+
3282
+ ## Router tests
3283
+
3284
+ At least:
3285
+
3286
+ ```
3287
+ `500 validation queries`
3288
+
3289
+ `100 hard negatives`
3290
+
3291
+ `50 unsupported queries`
3292
+ ```
3293
+
3294
+ Target:
3295
+
3296
+ ```
3297
+ `task accuracy β‰₯ 95%`
3298
+ ```
3299
+
3300
+ This is an engineering acceptance threshold, not a benchmark claim.
3301
+
3302
+
3303
+ ## Grounding tests
3304
+
3305
+ ```
3306
+ `IoU`
3307
+
3308
+ `Recall@0.5`
3309
+
3310
+ `coordinate conversion`
3311
+
3312
+ `NMS`
3313
+ ```
3314
+
3315
+ Test:
3316
+
3317
+ ```
3318
+ `object in center`
3319
+
3320
+ `object at boundary`
3321
+
3322
+ `multiple objects`
3323
+
3324
+ `small objects`
3325
+
3326
+ `large objects`
3327
+ ```
3328
+
3329
+
3330
+ ## Change tests
3331
+
3332
+ ```
3333
+ `perfect alignment`
3334
+
3335
+ `1-pixel shift`
3336
+
3337
+ `10-pixel shift`
3338
+
3339
+ `brightness change`
3340
+
3341
+ `no change`
3342
+
3343
+ `large construction change`
3344
+
3345
+ `small change`
3346
+ ```
3347
+
3348
+
3349
+ # 60. Adversarial Tests
3350
+
3351
+ The application must survive:
3352
+
3353
+ ```
3354
+ `blank image`
3355
+
3356
+ `all-zero image`
3357
+
3358
+ `extremely bright image`
3359
+
3360
+ `extremely dark image`
3361
+
3362
+ `noise image`
3363
+
3364
+ `unsupported TIFF`
3365
+
3366
+ `10000Γ—10000 image`
3367
+
3368
+ `missing metadata`
3369
+
3370
+ `wrong modality labels`
3371
+
3372
+ `same image uploaded twice`
3373
+
3374
+ `one temporal image`
3375
+
3376
+ `three images`
3377
+
3378
+ `optical/SAR size mismatch`
3379
+ ```
3380
+
3381
+
3382
+ # 61. Evaluation Harness
3383
+
3384
+ Architecture:
3385
+
3386
+ ```
3387
+ `Dataset Adapter`
3388
+
3389
+ ` ↓`
3390
+
3391
+ `Immutable Manifest`
3392
+
3393
+ ` ↓`
3394
+
3395
+ `Prediction Runner`
3396
+
3397
+ ` ↓`
3398
+
3399
+ `Schema Validator`
3400
+
3401
+ ` ↓`
3402
+
3403
+ `Metrics`
3404
+
3405
+ ` ↓`
3406
+
3407
+ `Normalization`
3408
+
3409
+ ` ↓`
3410
+
3411
+ `Report`
3412
+ ```
3413
+
3414
+
3415
+ # 62. Public/Hidden Evaluation Modes
3416
+
3417
+ ### Public mode
3418
+
3419
+ ```
3420
+ `benchmark test`
3421
+
3422
+ `frozen model`
3423
+
3424
+ `frozen config`
3425
+
3426
+ `frozen prompts`
3427
+ ```
3428
+
3429
+ ### Hidden-compatible mode
3430
+
3431
+ ```
3432
+ `unknown imagery`
3433
+
3434
+ `same preprocessing`
3435
+
3436
+ `same controller`
3437
+
3438
+ `same result schema`
3439
+
3440
+ `same model pipeline`
3441
+ ```
3442
+
3443
+ No benchmark-specific branching.
3444
+
3445
+
3446
+ # 63. Metric Normalisation
3447
+
3448
+ For each raw metric:
3449
+
3450
+ ```
3451
+ `raw`
3452
+
3453
+ ` ↓`
3454
+
3455
+ `task-specific transform`
3456
+
3457
+ ` ↓`
3458
+
3459
+ `0–1 normalized`
3460
+ ```
3461
+
3462
+ The framework supports:
3463
+
3464
+ ```
3465
+ `aggregate:`
3466
+
3467
+ ` official\_weights: null`
3468
+ ```
3469
+
3470
+ until the organizers publish them.
3471
+
3472
+ The uploaded specification explicitly prohibits inventing an official aggregate formula.
3473
+
3474
+
3475
+ # 64. Ablation Studies
3476
+
3477
+ Minimum set:
3478
+
3479
+ ### VLM
3480
+
3481
+ ```
3482
+ `base SmolVLM`
3483
+
3484
+ `vs`
3485
+
3486
+ `RS-adapted SmolVLM`
3487
+ ```
3488
+
3489
+ ### Fusion
3490
+
3491
+ ```
3492
+ `optical only`
3493
+
3494
+ `SAR only`
3495
+
3496
+ `CROMA joint`
3497
+ ```
3498
+
3499
+ ### Change
3500
+
3501
+ ```
3502
+ `raw difference`
3503
+
3504
+ `vs`
3505
+
3506
+ `STANet-style model`
3507
+ ```
3508
+
3509
+ ### Grounding
3510
+
3511
+ ```
3512
+ `RemoteCLIP zero-shot`
3513
+
3514
+ `vs`
3515
+
3516
+ `grounding head`
3517
+ ```
3518
+
3519
+ ### Router
3520
+
3521
+ ```
3522
+ `keyword baseline`
3523
+
3524
+ `vs`
3525
+
3526
+ `MiniLM router`
3527
+ ```
3528
+
3529
+ ### Agent
3530
+
3531
+ ```
3532
+ `hardcoded workflows`
3533
+
3534
+ `vs`
3535
+
3536
+ `NLP + deterministic policy`
3537
+ ```
3538
+
3539
+ This last ablation is particularly useful because it demonstrates that the β€œagentic” component is actually useful without pretending it is magic.
3540
+
3541
+
3542
+ # 65. Research/Architecture Search Plan
3543
+
3544
+ | **Candidate** | **Purpose** | **Decision** |
3545
+ | :-: | :-: | :-: |
3546
+ | SmolVLM-500M | compact VQA/caption | **Primary** |
3547
+ | PaliGemma2-3B | stronger VLM research baseline | Benchmark fallback |
3548
+ | RemoteCLIP ViT-B/32 | grounding | **Primary** |
3549
+ | RemoteCLIP ViT-L/14 | stronger grounding | Optional |
3550
+ | CROMA-base | optical-SAR | **Primary** |
3551
+ | Prithvi-EO-2.0-300M | optical EO features | Alternative |
3552
+ | STANet-style | change | **Primary** |
3553
+ | SatMAE | temporal representation | Reject for primary path |
3554
+ | SmolLM2-135M | NLP router | Not primary |
3555
+ | MiniLM-L6-v2 | NLP router | **Primary** |
3556
+
3557
+ PaliGemma 2 has out-of-the-box capabilities spanning VQA, captioning, object detection and segmentation, making it a useful research comparison, but its 3B-scale footprint and licensing/access conditions make it less attractive for the final compact competition runtime. ([Hugging Face](https://huggingface.co/google/paligemma2-3b-mix-224?utm_source=chatgpt.com))
3558
+
3559
+ SmolLM2-135M is Apache-2.0 and tiny, but it is a causal language model, so using it solely to classify a six-class intent space is needless generative machinery. ([Hugging Face](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct/blob/main/config.json?utm_source=chatgpt.com))
3560
+
3561
+
3562
+ # 66. Decision Log
3563
+
3564
+ | **Decision** | **Chosen option** | **Alternatives** | **Reason** |
3565
+ | :-: | :-: | :-: | :-: |
3566
+ | Router | MiniLM + classifier adapter | SmolLM2-135M | classification is simpler |
3567
+ | Orchestration | NLP + deterministic policy | autonomous agent | reproducibility |
3568
+ | VLM | SmolVLM-500M | PaliGemma2 | compact runtime |
3569
+ | Grounding | RemoteCLIP B/32 + head | VLM coordinate generation | actual spatial model |
3570
+ | Secondary task | captioning | grounding only | now grounding additionally exists |
3571
+ | Change | STANet-style | raw difference | learned change representation |
3572
+ | Fusion | CROMA | CNN concatenation | native optical-SAR representation |
3573
+ | Training | frozen backbones + heads/LoRA | full FT | compute |
3574
+ | Deployment | Gradio modular monolith | microservices | competition scope |
3575
+
3576
+
3577
+ # 67. Immutable Decisions
3578
+
3579
+ The implementation model must not redesign:
3580
+
3581
+ ```
3582
+ `\[ \] modular-monolith architecture`
3583
+
3584
+ `\[ \] tiny NLP intent router`
3585
+
3586
+ `\[ \] deterministic policy engine`
3587
+
3588
+ `\[ \] common specialist interface`
3589
+
3590
+ `\[ \] SmolVLM VLM layer`
3591
+
3592
+ `\[ \] RemoteCLIP grounding path`
3593
+
3594
+ `\[ \] STANet-style change path`
3595
+
3596
+ `\[ \] CROMA optical-SAR path`
3597
+
3598
+ `\[ \] shared evidence engine`
3599
+
3600
+ `\[ \] calibrated confidence`
3601
+
3602
+ `\[ \] common result schema`
3603
+
3604
+ `\[ \] execution trace`
3605
+
3606
+ `\[ \] leakage isolation`
3607
+ ```
3608
+
3609
+
3610
+ # 68. Variables Open for Tuning
3611
+
3612
+ ```
3613
+ `LoRA rank`
3614
+
3615
+ `VLM learning rate`
3616
+
3617
+ `router adapter dimension`
3618
+
3619
+ `router confidence threshold`
3620
+
3621
+ `grounding head architecture`
3622
+
3623
+ `grounding learning rate`
3624
+
3625
+ `change threshold`
3626
+
3627
+ `change loss weighting`
3628
+
3629
+ `CROMA fusion head width`
3630
+
3631
+ `tile size`
3632
+
3633
+ `tile overlap`
3634
+
3635
+ `top-K tile count`
3636
+
3637
+ `confidence calibration temperature`
3638
+ ```
3639
+
3640
+
3641
+ # 69. Variables Requiring Experimental Optimization
3642
+
3643
+ ## Router
3644
+
3645
+ ```
3646
+ `hidden dimension:`
3647
+
3648
+ `64–256`
3649
+
3650
+
3651
+ `dropout:`
3652
+
3653
+ `0–0.3`
3654
+
3655
+
3656
+ `confidence:`
3657
+
3658
+ `0.60–0.90`
3659
+ ```
3660
+
3661
+ ## Grounding
3662
+
3663
+ ```
3664
+ `head width:`
3665
+
3666
+ `256–1024`
3667
+
3668
+
3669
+ `learning rate:`
3670
+
3671
+ `5e-5–2e-4`
3672
+
3673
+
3674
+ `NMS:`
3675
+
3676
+ `0.4–0.6`
3677
+ ```
3678
+
3679
+ ## Change
3680
+
3681
+ ```
3682
+ `threshold:`
3683
+
3684
+ `0.30–0.70`
3685
+
3686
+
3687
+ `minimum component:`
3688
+
3689
+ `16–128 px`
3690
+ ```
3691
+
3692
+ ## CROMA
3693
+
3694
+ ```
3695
+ `fusion width:`
3696
+
3697
+ `256–1024`
3698
+
3699
+
3700
+ `dropout:`
3701
+
3702
+ `0–0.3`
3703
+ ```
3704
+
3705
+ Selection:
3706
+
3707
+ **validation only.**
3708
+
3709
+
3710
+ # 70. Risk Register
3711
+
3712
+ | **Risk** | Probability | Impact | **Mitigation** |
3713
+ | :-: | -: | -: | :-: |
3714
+ | VLM insufficient EO reasoning | Medium | High | RS adaptation + stronger benchmark |
3715
+ | Grounding weak on overhead imagery | Medium | High | VRSBench adaptation + RemoteCLIP |
3716
+ | CROMA Sentinel-to-Indian-sensor shift | High | Very High | sensor adapter + channel dropout |
3717
+ | Hidden test distribution shift | High | Very High | sensor-agnostic preprocessing |
3718
+ | Kaggle quota | Medium | High | staged training/checkpoints |
3719
+ | GPU OOM | Medium | High | frozen models + PEFT |
3720
+ | leakage | Low/Medium | Catastrophic | scene-level manifests |
3721
+ | router mistakes | Low | Medium | confidence + deterministic fallback |
3722
+ | model unavailable | Low | Medium | local HF cache + pinned revisions |
3723
+ | deployment memory | Medium | High | lazy loading |
3724
+
3725
+
3726
+ # 71. Implementation Order
3727
+
3728
+ This is the actual build order I would give the coding model.
3729
+
3730
+ ## Phase 0
3731
+
3732
+ Environment and repository skeleton.
3733
+
3734
+ ## Phase 1
3735
+
3736
+ Config registry + typed schemas.
3737
+
3738
+ ## Phase 2
3739
+
3740
+ GeoTIFF loading and validation.
3741
+
3742
+ ## Phase 3
3743
+
3744
+ Dataset manifests + leakage framework.
3745
+
3746
+ ## Phase 4
3747
+
3748
+ MiniLM intent router.
3749
+
3750
+ ## Phase 5
3751
+
3752
+ SmolVLM baseline.
3753
+
3754
+ ## Phase 6
3755
+
3756
+ BigEarthNet VLM adaptation.
3757
+
3758
+ ## Phase 7
3759
+
3760
+ RemoteCLIP grounding baseline.
3761
+
3762
+ ## Phase 8
3763
+
3764
+ Grounding head training.
3765
+
3766
+ ## Phase 9
3767
+
3768
+ STANet change detector.
3769
+
3770
+ ## Phase 10
3771
+
3772
+ CDVQA reasoning integration.
3773
+
3774
+ ## Phase 11
3775
+
3776
+ CROMA integration.
3777
+
3778
+ ## Phase 12
3779
+
3780
+ Optical-SAR fusion head.
3781
+
3782
+ ## Phase 13
3783
+
3784
+ Evidence engine.
3785
+
3786
+ ## Phase 14
3787
+
3788
+ Confidence calibration.
3789
+
3790
+ ## Phase 15
3791
+
3792
+ Deterministic workflow controller.
3793
+
3794
+ ## Phase 16
3795
+
3796
+ GUI.
3797
+
3798
+ ## Phase 17
3799
+
3800
+ Benchmark evaluation harness.
3801
+
3802
+ ## Phase 18
3803
+
3804
+ HF deployment.
3805
+
3806
+ ## Phase 19
3807
+
3808
+ Final hardening and demonstration.
3809
+
3810
+
3811
+ # 72. Stop/Go Gates
3812
+
3813
+ ## Gate 1
3814
+
3815
+ Proceed only when:
3816
+
3817
+ ```
3818
+ `\[ \] TIFF validator works`
3819
+
3820
+ `\[ \] CRS handling works`
3821
+
3822
+ `\[ \] leakage test passes`
3823
+ ```
3824
+
3825
+ ## Gate 2
3826
+
3827
+ Proceed to VLM adaptation only when:
3828
+
3829
+ ```
3830
+ `\[ \] baseline inference works`
3831
+
3832
+ `\[ \] dataset manifests frozen`
3833
+ ```
3834
+
3835
+ ## Gate 3
3836
+
3837
+ Proceed to grounding training when:
3838
+
3839
+ ```
3840
+ `\[ \] RemoteCLIP loads`
3841
+
3842
+ `\[ \] VRSBench boxes parse correctly`
3843
+
3844
+ `\[ \] coordinate conversions tested`
3845
+ ```
3846
+
3847
+ ## Gate 4
3848
+
3849
+ Proceed to controller integration when:
3850
+
3851
+ ```
3852
+ `\[ \] VQA works`
3853
+
3854
+ `\[ \] grounding works`
3855
+
3856
+ `\[ \] change works`
3857
+
3858
+ `\[ \] CROMA works`
3859
+ ```
3860
+
3861
+ ## Gate 5
3862
+
3863
+ Proceed to final evaluation when:
3864
+
3865
+ ```
3866
+ `\[ \] all prompts frozen`
3867
+
3868
+ `\[ \] all thresholds frozen`
3869
+
3870
+ `\[ \] model revisions frozen`
3871
+
3872
+ `\[ \] config hash recorded`
3873
+
3874
+ `\[ \] public test isolation verified`
3875
+ ```
3876
+
3877
+
3878
+ # 73. "DO NOT BUILD THIS"
3879
+
3880
+ Do not build:
3881
+
3882
+ ```
3883
+ `\[ \] Kubernetes`
3884
+
3885
+ `\[ \] Docker swarm`
3886
+
3887
+ `\[ \] Kafka`
3888
+
3889
+ `\[ \] Redis cluster`
3890
+
3891
+ `\[ \] vector database`
3892
+
3893
+ `\[ \] autonomous agent swarm`
3894
+
3895
+ `\[ \] multi-LLM debate`
3896
+
3897
+ `\[ \] seven VLMs`
3898
+
3899
+ `\[ \] full CROMA retraining`
3900
+
3901
+ `\[ \] full foundation-model retraining`
3902
+
3903
+ `\[ \] LLM-generated bounding boxes`
3904
+
3905
+ `\[ \] LLM-generated confidence`
3906
+
3907
+ `\[ \] benchmark-specific hidden-test branches`
3908
+
3909
+ `\[ \] cloud inference dependency`
3910
+ ```
3911
+
3912
+ The architecture is intentionally **competition-scale**.
3913
+
3914
+
3915
+ # 74. Production Readiness Boundary
3916
+
3917
+ You explicitly said **do not add production implementations**.
3918
+
3919
+ So this plan deliberately stops here:
3920
+
3921
+ ### Included
3922
+
3923
+ ```
3924
+ `modularity`
3925
+
3926
+ `stable interfaces`
3927
+
3928
+ `lazy model loading`
3929
+
3930
+ `error handling`
3931
+
3932
+ `config centralization`
3933
+
3934
+ `traceability`
3935
+
3936
+ `reproducibility`
3937
+
3938
+ `checkpointing`
3939
+
3940
+ `evaluation`
3941
+
3942
+ `basic resource limits`
3943
+ ```
3944
+
3945
+ ### Not included
3946
+
3947
+ ```
3948
+ `authentication`
3949
+
3950
+ `multi-tenant isolation`
3951
+
3952
+ `distributed queues`
3953
+
3954
+ `autoscaling`
3955
+
3956
+ `observability platform`
3957
+
3958
+ `Kubernetes`
3959
+
3960
+ `service mesh`
3961
+
3962
+ `distributed storage`
3963
+
3964
+ `horizontal worker orchestration`
3965
+
3966
+ `enterprise security`
3967
+
3968
+ `billing`
3969
+
3970
+ `SLA infrastructure`
3971
+ ```
3972
+
3973
+ The architecture is **scale-compatible**, but not a production implementation.
3974
+
3975
+ That's exactly what we want at this stage.
3976
+
3977
+
3978
+ # 75. Final Feasibility Audit
3979
+
3980
+ ### Can it train on current Kaggle free GPU?
3981
+
3982
+ **YES, conditionally.**
3983
+
3984
+ Current Kaggle documentation now points to T4Γ—2 as the available newer GPU path, with the P100 retired on September 15, 2026. ([Kaggle](https://www.kaggle.com/product-announcements/735239?utm_source=chatgpt.com))
3985
+
3986
+ ### Can it deploy on HF free hosting?
3987
+
3988
+ **YES, conditionally.**
3989
+
3990
+ HF currently supports free ZeroGPU Spaces for eligible personal accounts, with up to two Spaces and a free daily GPU quota. ([Hugging Face](https://huggingface.co/docs/hub/spaces-zerogpu?utm_source=chatgpt.com))
3991
+
3992
+ ### Remote-sensing adaptation?
3993
+
3994
+ **YES.**
3995
+
3996
+ ### Single-image VQA?
3997
+
3998
+ **YES.**
3999
+
4000
+ ### Captioning?
4001
+
4002
+ **YES.**
4003
+
4004
+ ### Grounding?
4005
+
4006
+ **YES.**
4007
+
4008
+ ### Temporal change?
4009
+
4010
+ **YES.**
4011
+
4012
+ ### Optical-SAR?
4013
+
4014
+ **YES.**
4015
+
4016
+ ### Agentic orchestration?
4017
+
4018
+ **YES.**
4019
+
4020
+ ### Natural-language understanding?
4021
+
4022
+ **YES. MiniLM adapter.**
4023
+
4024
+ ### Deterministic execution?
4025
+
4026
+ **YES.**
4027
+
4028
+ ### Evidence?
4029
+
4030
+ **YES.**
4031
+
4032
+ ### Confidence?
4033
+
4034
+ **YES.**
4035
+
4036
+ ### Execution trace?
4037
+
4038
+ **YES.**
4039
+
4040
+ ### Hidden ISRO/SAC compatibility?
4041
+
4042
+ **CONDITIONAL.**
4043
+
4044
+ The actual hidden distribution remains unavailable, so compatibility can be engineered but not empirically proven before official evaluation. The supplied contract explicitly says hidden annotations are not disclosed.
4045
+
4046
+
4047
+ # 76. Final Master Specification
4048
+
4049
+ ## Final models
4050
+
4051
+ ```
4052
+ `Router:`
4053
+
4054
+ `sentence-transformers/all-MiniLM-L6-v2`
4055
+
4056
+ `+ classifier adapter`
4057
+
4058
+
4059
+ `VLM:`
4060
+
4061
+ `HuggingFaceTB/SmolVLM-500M-Instruct`
4062
+
4063
+
4064
+ `Grounding:`
4065
+
4066
+ `chendelong/RemoteCLIP`
4067
+
4068
+ `RemoteCLIP-ViT-B-32`
4069
+
4070
+ `+ grounding head`
4071
+
4072
+
4073
+ `Change:`
4074
+
4075
+ `STANet-style Siamese model`
4076
+
4077
+
4078
+ `Optical-SAR:`
4079
+
4080
+ `antofuller/CROMA`
4081
+
4082
+ `CROMA\_base.pt`
4083
+
4084
+ `+ lightweight fusion head`
4085
+ ```
4086
+
4087
+ ## Final task map
4088
+
4089
+ ```
4090
+ `VQA`
4091
+
4092
+ `β†’ SmolVLM`
4093
+
4094
+
4095
+ `Caption`
4096
+
4097
+ `β†’ SmolVLM`
4098
+
4099
+
4100
+ `Grounding`
4101
+
4102
+ `β†’ MiniLM`
4103
+
4104
+ `β†’ RemoteCLIP`
4105
+
4106
+ `β†’ grounding head`
4107
+
4108
+ `β†’ SmolVLM explanation`
4109
+
4110
+
4111
+ `Change`
4112
+
4113
+ `β†’ STANet`
4114
+
4115
+ `β†’ SmolVLM explanation`
4116
+
4117
+
4118
+ `Optical-SAR`
4119
+
4120
+ `β†’ CROMA`
4121
+
4122
+ `β†’ fusion head`
4123
+
4124
+ `β†’ SmolVLM explanation`
4125
+ ```
4126
+
4127
+ ## Final control map
4128
+
4129
+ ```
4130
+ `User`
4131
+
4132
+ ` ↓`
4133
+
4134
+ `MiniLM`
4135
+
4136
+ ` ↓`
4137
+
4138
+ `Intent JSON`
4139
+
4140
+ ` ↓`
4141
+
4142
+ `Schema validation`
4143
+
4144
+ ` ↓`
4145
+
4146
+ `Deterministic workflow`
4147
+
4148
+ ` ↓`
4149
+
4150
+ `specialists`
4151
+
4152
+ ` ↓`
4153
+
4154
+ `evidence`
4155
+
4156
+ ` ↓`
4157
+
4158
+ `confidence`
4159
+
4160
+ ` ↓`
4161
+
4162
+ `result`
4163
+ ```
4164
+
4165
+
4166
+ # 77. IMPLEMENTATION MODEL HANDOFF
4167
+
4168
+ ## Immutable architecture decisions
4169
+
4170
+ The implementation model must **not redesign**:
4171
+
4172
+ 1. MiniLM intent router.
4173
+
4174
+ 2. Deterministic policy controller.
4175
+
4176
+ 3. Modular specialist interface.
4177
+
4178
+ 4. SmolVLM main VLM.
4179
+
4180
+ 5. RemoteCLIP grounding specialist.
4181
+
4182
+ 6. STANet-style change detector.
4183
+
4184
+ 7. CROMA-based optical-SAR path.
4185
+
4186
+ 8. Common evidence engine.
4187
+
4188
+ 9. Calibrated confidence.
4189
+
4190
+ 10. Unified result schema.
4191
+
4192
+ 11. Execution trace.
4193
+
4194
+ 12. Dataset leakage protections.
4195
+
4196
+ 13. Central configuration.
4197
+
4198
+ 14. Competition-scale Gradio deployment.
4199
+
4200
+
4201
+ ## Variables the implementation model may tune
4202
+
4203
+ ```
4204
+ `router adapter width`
4205
+
4206
+ `router threshold`
4207
+
4208
+ `LoRA rank`
4209
+
4210
+ `VLM LR`
4211
+
4212
+ `tile size`
4213
+
4214
+ `tile overlap`
4215
+
4216
+ `top-K tiles`
4217
+
4218
+ `grounding head width`
4219
+
4220
+ `grounding threshold`
4221
+
4222
+ `change threshold`
4223
+
4224
+ `component filtering`
4225
+
4226
+ `fusion head width`
4227
+
4228
+ `confidence calibration`
4229
+ ```
4230
+
4231
+
4232
+ ## Variables that must be experimentally optimized
4233
+
4234
+ ```
4235
+ `router threshold`
4236
+
4237
+ `grounding learning rate`
4238
+
4239
+ `grounding loss weights`
4240
+
4241
+ `change threshold`
4242
+
4243
+ `change loss weights`
4244
+
4245
+ `VLM learning rate`
4246
+
4247
+ `LoRA rank`
4248
+
4249
+ `CROMA fusion projection width`
4250
+
4251
+ `tile top-K`
4252
+ ```
4253
+
4254
+ All selection must happen on validation data.
4255
+
4256
+
4257
+ ## Unknowns
4258
+
4259
+ ### Exact official metrics
4260
+
4261
+ **UNVERIFIED**
4262
+
4263
+ Insert official definitions when published.
4264
+
4265
+ ### Exact hidden sensor band arrangement
4266
+
4267
+ **UNVERIFIED**
4268
+
4269
+ Resolve from official evaluation input metadata only.
4270
+
4271
+ ### Exact hidden task output format
4272
+
4273
+ **UNVERIFIED**
4274
+
4275
+ Resolve through the supplied evaluation specification when provided.
4276
+
4277
+ ### Exact HF quota at judging time
4278
+
4279
+ **VARIABLE**
4280
+
4281
+ CPU-compatible execution remains mandatory.
4282
+
4283
+
4284
+ # 78. Non-Negotiable Requirements
4285
+
4286
+ ```
4287
+ `\[ \] Natural language query understanding`
4288
+
4289
+ `\[ \] Learned NLP router`
4290
+
4291
+ `\[ \] Deterministic execution`
4292
+
4293
+ `\[ \] Single-image VQA`
4294
+
4295
+ `\[ \] Captioning`
4296
+
4297
+ `\[ \] Grounding`
4298
+
4299
+ `\[ \] Temporal change`
4300
+
4301
+ `\[ \] Optical-SAR analysis`
4302
+
4303
+ `\[ \] GeoTIFF`
4304
+
4305
+ `\[ \] Geo metadata`
4306
+
4307
+ `\[ \] Evidence`
4308
+
4309
+ `\[ \] Confidence`
4310
+
4311
+ `\[ \] Execution trace`
4312
+
4313
+ `\[ \] Remote-sensing adaptation`
4314
+
4315
+ `\[ \] Leakage prevention`
4316
+
4317
+ `\[ \] Public-test isolation`
4318
+
4319
+ `\[ \] Hidden-test compatible preprocessing`
4320
+
4321
+ `\[ \] Reproducible evaluation`
4322
+
4323
+ `\[ \] HF deployment`
4324
+
4325
+ `\[ \] Kaggle-compatible training`
4326
+ ```
4327
+
4328
+
4329
+ # 79. Completion Checklist
4330
+
4331
+ ```
4332
+ `ARCHITECTURE`
4333
+
4334
+ `\[ \] Specialist interface implemented`
4335
+
4336
+ `\[ \] Controller implemented`
4337
+
4338
+ `\[ \] Workflow DAG implemented`
4339
+
4340
+ `\[ \] Result schema implemented`
4341
+
4342
+
4343
+ `ROUTER`
4344
+
4345
+ `\[ \] MiniLM loaded`
4346
+
4347
+ `\[ \] Adapter trained`
4348
+
4349
+ `\[ \] Hard negatives tested`
4350
+
4351
+ `\[ \] Confidence threshold validated`
4352
+
4353
+ `\[ \] Fallback routing implemented`
4354
+
4355
+
4356
+ `VLM`
4357
+
4358
+ `\[ \] SmolVLM baseline`
4359
+
4360
+ `\[ \] BigEarthNet adaptation`
4361
+
4362
+ `\[ \] validation comparison`
4363
+
4364
+ `\[ \] frozen prompt set`
4365
+
4366
+
4367
+ `GROUNDING`
4368
+
4369
+ `\[ \] RemoteCLIP loaded`
4370
+
4371
+ `\[ \] VRSBench parser`
4372
+
4373
+ `\[ \] grounding head`
4374
+
4375
+ `\[ \] box conversion`
4376
+
4377
+ `\[ \] NMS`
4378
+
4379
+ `\[ \] confidence`
4380
+
4381
+
4382
+ `CHANGE`
4383
+
4384
+ `\[ \] STANet-style model`
4385
+
4386
+ `\[ \] LEVIR-CD training`
4387
+
4388
+ `\[ \] change map`
4389
+
4390
+ `\[ \] connected components`
4391
+
4392
+ `\[ \] CDVQA reasoning`
4393
+
4394
+
4395
+ `OPTICAL-SAR`
4396
+
4397
+ `\[ \] CROMA loaded`
4398
+
4399
+ `\[ \] sensor adapter`
4400
+
4401
+ `\[ \] optical preprocessing`
4402
+
4403
+ `\[ \] SAR preprocessing`
4404
+
4405
+ `\[ \] fusion head`
4406
+
4407
+ `\[ \] sensor-dropout testing`
4408
+
4409
+
4410
+ `EVIDENCE`
4411
+
4412
+ `\[ \] image crops`
4413
+
4414
+ `\[ \] boxes`
4415
+
4416
+ `\[ \] masks`
4417
+
4418
+ `\[ \] change maps`
4419
+
4420
+ `\[ \] modality evidence`
4421
+
4422
+
4423
+ `EVALUATION`
4424
+
4425
+ `\[ \] manifests`
4426
+
4427
+ `\[ \] leakage scans`
4428
+
4429
+ `\[ \] benchmark adapters`
4430
+
4431
+ `\[ \] metrics`
4432
+
4433
+ `\[ \] normalization`
4434
+
4435
+ `\[ \] immutable public test`
4436
+
4437
+ `\[ \] reproducible run manifest`
4438
+
4439
+
4440
+ `GUI`
4441
+
4442
+ `\[ \] upload`
4443
+
4444
+ `\[ \] viewer`
4445
+
4446
+ `\[ \] query`
4447
+
4448
+ `\[ \] task display`
4449
+
4450
+ `\[ \] grounding overlay`
4451
+
4452
+ `\[ \] change overlay`
4453
+
4454
+ `\[ \] optical/SAR comparison`
4455
+
4456
+ `\[ \] confidence`
4457
+
4458
+ `\[ \] trace`
4459
+
4460
+ `\[ \] report`
4461
+
4462
+
4463
+ `DEPLOYMENT`
4464
+
4465
+ `\[ \] HF Space`
4466
+
4467
+ `\[ \] CPU mode`
4468
+
4469
+ `\[ \] ZeroGPU mode`
4470
+
4471
+ `\[ \] lazy loading`
4472
+
4473
+ `\[ \] sequential request test`
4474
+
4475
+ `\[ \] cold-start test`
4476
+ ```
4477
+
4478
+
4479
+ # 80. The final conceptual architecture
4480
+
4481
+ This is the version I would now **freeze as the competition architecture**:
4482
+
4483
+ ```
4484
+ ` SATQUERY AI`
4485
+
4486
+ ` β”‚`
4487
+
4488
+ ` β–Ό`
4489
+
4490
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4491
+
4492
+ ` β”‚ Tiny NLP Router β”‚`
4493
+
4494
+ ` β”‚ MiniLM + Adapter β”‚`
4495
+
4496
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4497
+
4498
+ ` β”‚`
4499
+
4500
+ ` β–Ό`
4501
+
4502
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4503
+
4504
+ ` β”‚ Policy Controllerβ”‚`
4505
+
4506
+ ` β”‚ Deterministic β”‚`
4507
+
4508
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4509
+
4510
+ ` β”‚`
4511
+
4512
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4513
+
4514
+ ` οΏ½οΏ½οΏ½ β”‚ β”‚`
4515
+
4516
+ ` β–Ό β–Ό β–Ό`
4517
+
4518
+ ` VQA/CAP GROUNDING CHANGE`
4519
+
4520
+ ` SmolVLM RemoteCLIP + Head STANet`
4521
+
4522
+ ` β”‚ β”‚ β”‚`
4523
+
4524
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4525
+
4526
+ ` β”‚`
4527
+
4528
+ ` β–Ό`
4529
+
4530
+ ` OPTICAL + SAR`
4531
+
4532
+ ` CROMA`
4533
+
4534
+ ` +`
4535
+
4536
+ ` Fusion Head`
4537
+
4538
+ ` β”‚`
4539
+
4540
+ ` β–Ό`
4541
+
4542
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4543
+
4544
+ ` β”‚ Evidence Engine β”‚`
4545
+
4546
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4547
+
4548
+ ` β–Ό`
4549
+
4550
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4551
+
4552
+ ` β”‚ Confidence β”‚`
4553
+
4554
+ ` β”‚ Calibration β”‚`
4555
+
4556
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4557
+
4558
+ ` β–Ό`
4559
+
4560
+ ` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”`
4561
+
4562
+ ` β”‚ Result Contract β”‚`
4563
+
4564
+ ` β”‚ Answer + Spatial β”‚`
4565
+
4566
+ ` β”‚ Evidence + Trace β”‚`
4567
+
4568
+ ` β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜`
4569
+
4570
+ ` β–Ό`
4571
+
4572
+ ` Gradio GUI`
4573
+ ```
4574
+
4575
+ And the scaling principle remains:
4576
+
4577
+ ```
4578
+ `NOW`
4579
+
4580
+
4581
+ `Controller`
4582
+
4583
+ ` ↓`
4584
+
4585
+ `Python specialists`
4586
+
4587
+
4588
+
4589
+ `LATER`
4590
+
4591
+
4592
+ `Controller`
4593
+
4594
+ ` ↓`
4595
+
4596
+ `specialist interface`
4597
+
4598
+ ` ↓`
4599
+
4600
+ `β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”`
4601
+
4602
+ `VLM Grounding Change Fusion`
4603
+
4604
+ `worker worker worker worker`
4605
+ ```
4606
+
4607
+ Nothing fundamental changes.
4608
+
4609
+ That is the important architectural property we wanted.
4610
+
4611
+ ### Final recommendation
4612
+
4613
+ **Yes, use the tiny NLP adapter.**
4614
+
4615
+ More specifically:
4616
+
4617
+ > **MiniLM embedding encoder + tiny multi-head intent adapter + deterministic workflow policy.**
4618
+
4619
+ It gives SatQuery actual natural-language understanding while keeping orchestration auditable and predictable.
4620
+
4621
+ And **yes, add grounding as a first-class specialist**, with RemoteCLIP providing the remote-sensing image-text alignment and a lightweight learned grounding head producing the actual spatial localization. RemoteCLIP is already explicitly built for remote-sensing vision-language alignment and provides ViT-B/32, ViT-L/14 and RN50 pretrained checkpoints. ([GitHub](https://github.com/ChenDelong1999/RemoteCLIP?utm_source=chatgpt.com))
4622
+
4623
+ The major current infrastructure correction is also locked in: **Kaggle planning should now assume T4Γ—2, not P100**
4624
+
4625
+ This gives us a very clean progression:
4626
+
4627
+ ```
4628
+ `tiny NLP understanding`
4629
+
4630
+ ` +`
4631
+
4632
+ `deterministic agent`
4633
+
4634
+ ` +`
4635
+
4636
+ `specialist intelligence`
4637
+
4638
+ ` +`
4639
+
4640
+ `evidence`
4641
+
4642
+ ` +`
4643
+
4644
+ `confidence`
4645
+
4646
+ ` +`
4647
+
4648
+ `multimodal/temporal reasoning`
4649
+
4650
+ ` =`
4651
+
4652
+ `SatQuery AI`
4653
+ ```
4654
+
4655
+ And importantly, it remains a **competition prototype**, not a startup accidentally carrying Kubernetes on its back like a medieval peasant carrying a castle.
4656
+