thundercode commited on
Commit
3a34048
·
verified ·
1 Parent(s): b10b360

release: add MODEL_CARD.md

Browse files
Files changed (1) hide show
  1. MODEL_CARD.md +164 -0
MODEL_CARD.md ADDED
@@ -0,0 +1,164 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ license: other
4
+ library_name: pytorch
5
+ tags:
6
+ - remote-sensing
7
+ - satellite-imagery
8
+ - earth-observation
9
+ - change-detection
10
+ - visual-grounding
11
+ - image-captioning
12
+ - visual-question-answering
13
+ - optical-sar-fusion
14
+ - lora
15
+ - peft
16
+ pipeline_tag: image-to-text
17
+ config_hash: 78f1e3700da15aa1
18
+ ---
19
+
20
+ # Model Card — SatQuery AI
21
+
22
+ SatQuery AI answers natural-language questions about satellite imagery using a **router + specialists**
23
+ design. This card covers the **six trained artifacts** released by the project. It is deliberately
24
+ explicit about what is measured, what is not, and what was rejected.
25
+
26
+ > **The six trained artifacts are small modules on top of frozen, publicly-pinned backbones. No
27
+ > backbone weights are redistributed by this release** — they are fetched from the Hugging Face Hub at
28
+ > run time, pinned by revision.
29
+
30
+ Machine-readable identities (byte counts and sha256) are in
31
+ [`models/manifest.json`](models/manifest.json) and [`models/checksums.sha256`](models/checksums.sha256),
32
+ **generated by reading the files**. Where this card and the generated manifest disagree, the manifest
33
+ wins.
34
+
35
+ ---
36
+
37
+ ## 1. Artifacts in this release
38
+
39
+ | # | Task | Kind | File | Bytes | sha256 (first 16) |
40
+ |---|---|---|---|---|---|
41
+ | 1 | `change` | trained head | `head.pt` | 63,231,009 | `c5ef31277b67aa01` |
42
+ | 2 | `change_vqa` | trained head | `head.pt` | 5,822,809 | `cfae5e43b97ca930` |
43
+ | 3 | `optical_sar` | trained head | `head.pt` | 14,427,457 | `785815729a3a39fc` |
44
+ | 4 | `grounding` | trained head | `head.pt` | 12,639,041 | `93432f7034be91a8` |
45
+ | 5 | `router` | trained adapter | `adapter.pt` | 211,961 | `8527c3ed28a293e1` |
46
+ | 6 | `vlm` | LoRA adapter | `adapter_model.safetensors` | 34,798,048 | `07c76a75fa046248` |
47
+
48
+ Two of these hashes (`change_vqa`, `vlm`) **agree exactly** with hashes recorded independently at
49
+ promotion time — an external cross-check, not a self-consistency claim.
50
+
51
+ ## 2. Backbone dependencies (frozen, pinned by revision)
52
+
53
+ | Role | Repository | Revision |
54
+ |---|---|---|
55
+ | Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` |
56
+ | VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` |
57
+ | Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` |
58
+ | Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` |
59
+ | Change | STANet-style (ResNet-18 + PAM) | trained in-project |
60
+
61
+ ## 3. Intended use
62
+
63
+ - **Research and demonstration** of a modular, CPU-first remote-sensing QA system.
64
+ - **Routing and dispatch** of natural-language queries to the appropriate specialist.
65
+ - **Reproducible evaluation** of each specialist on its own documented split.
66
+
67
+ ## 4. Out-of-scope use
68
+
69
+ - **Safety-, legal- or life-critical decisions.** No accuracy, calibration or robustness guarantee is
70
+ offered for any high-stakes use.
71
+ - **Operational geospatial production** without independent validation.
72
+ - **Any use of the VLM adapter as a production model** — it is acceptance-rejected (§6).
73
+ - **Treating per-specialist metrics as system-level accuracy.** No end-to-end benchmark exists (§7).
74
+
75
+ ## 5. Measured performance
76
+
77
+ | Task | Metric | Value | Split / protocol |
78
+ |---|---|---|---|
79
+ | change | pooled IoU / macro IoU / pooled F1 | **0.8122 / 0.8457 / 0.8964** | LEVIR-CD-256 test, n = 2048 |
80
+ | grounding | mean best IoU / recall@0.5 | **0.2838 / 0.2198** canonical; **0.2566 / 0.1938** matched6 | VRSBench, n = 16159 |
81
+ | grounding | head-argmax / zero-shot baseline IoU | **0.1215 / 0.0972** | canonical |
82
+ | optical_sar | accuracy / macro F1 | **0.931 / 0.434161** | held-out test, n = 4000, 19 classes |
83
+ | change_vqa | accuracy / macro F1 | **0.697626 / 0.378373** (test); **0.651469 / 0.372309** (test2) | two test sets |
84
+ | vlm | exact_match / F1 | **0.963 / 0.96432** | frozen 1000-question subset |
85
+ | router | overall **ungated** accuracy | **0.965116** | val, n = 86 — **TEST NOT RUN** |
86
+
87
+ **Every value is checked against its artifact** by `tools/verify_readme_metrics.py`.
88
+
89
+ ### 5.1 Calibration — reported as a negative result
90
+
91
+ Temperature scaling is enabled (T = 0.9772732, fit on val n = 16,441). **ECE worsened**:
92
+ 0.013755 → **0.014929**. It is retained because it is part of the frozen configuration, not because it
93
+ helped.
94
+
95
+ ## 6. Acceptance status
96
+
97
+ | Artifact | Metrics | Acceptance |
98
+ |---|---|---|
99
+ | change | VERIFIED | accepted (shipped) |
100
+ | grounding | measured (2 protocols) | shipped |
101
+ | optical_sar | measured | **ruling OPEN** |
102
+ | change_vqa | measured (2 test sets) | **ruling OPEN** |
103
+ | router | measured (val only) | shipped; test NOT RUN |
104
+ | **vlm** | usable (exact_match 0.963) | **ACCEPTANCE-REJECTED** |
105
+
106
+ **`USABLE_VERIFIED` ≠ `ACCEPTANCE-ACCEPTED`.** The VLM adapter works and is not promoted; the deployed
107
+ caption/VQA path uses the **unadapted** model.
108
+
109
+ ## 7. Evaluation gaps (stated, not hidden)
110
+
111
+ - **No system-level end-to-end benchmark exists.** None is claimed.
112
+ - **Router test split: NOT RUN.**
113
+ - **Benchmark adapters: NOT RUN.**
114
+ - **Cross-dataset generalisation: NOT RUN.**
115
+ - **Human and robustness evaluation: NOT RUN.**
116
+
117
+ ## 8. Limitations
118
+
119
+ - Grounding absolute IoU is low (0.28) and protocol-sensitive.
120
+ - Optical-SAR accuracy is carried by common classes (macro-F1 0.434161).
121
+ - The BigEarthNet local subset is **100 % single-label** vs the official 1–11 multi-label scheme, so
122
+ its metrics are **not comparable** to published numbers.
123
+ - The optical-SAR service returns a bare class index, not a label.
124
+ - Known router residuals exist (e.g. *"What is the new runway?"* reads `change`).
125
+ - No `LICENSE` file exists in the source repository.
126
+
127
+ See [`docs/LIMITATIONS.md`](docs/LIMITATIONS.md) for the full catalogue.
128
+
129
+ ## 9. Training summary
130
+
131
+ Small modules on frozen backbones; seed 42; every artifact records the frozen config hash
132
+ `78f1e3700da15aa1`. Router and change/grounding/fusion heads train on CPU; the VLM LoRA adapter and
133
+ the change-VQA head were trained on **external GPUs** (the latter via a documented Kaggle run). Full
134
+ detail in [`docs/TRAINING.md`](docs/TRAINING.md).
135
+
136
+ ## 10. Provenance and verification
137
+
138
+ | Item | Location |
139
+ |---|---|
140
+ | Byte-exact manifest | `models/manifest.json` |
141
+ | Checksums | `models/checksums.sha256` |
142
+ | Metric verification tool | `tools/verify_readme_metrics.py` |
143
+ | Metric verification output | `tools/readme_metrics_report.txt` |
144
+ | Full documentation | `docs/` |
145
+ | Release manifest | `RELEASE_MANIFEST.md` |
146
+
147
+ ## 11. Licence
148
+
149
+ The project ships **no licence file**; a licence must be selected by the owner before public release
150
+ of the *code*. **Model weights carry the terms of their backbone licences** — consult each backbone's
151
+ Hugging Face page. Backbones are not redistributed here.
152
+
153
+ ## 12. Citation
154
+
155
+ If you use this work, cite the project repository:
156
+
157
+ ```bibtex
158
+ @misc{satquery_ai_2026,
159
+ title = {SatQuery AI: A Modular Router-and-Specialists System for Satellite Imagery Question Answering},
160
+ author = {SatQuery AI},
161
+ year = {2026},
162
+ note = {Public release: https://github.com/Anish-lab-blip/SatQuery-AI}
163
+ }
164
+ ```