Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/LIMITATIONS.md
Browse files- docs/LIMITATIONS.md +699 -62
docs/LIMITATIONS.md
CHANGED
|
@@ -4,82 +4,708 @@ An honest, exhaustive catalogue of everything SatQuery AI does **not** do, does
|
|
| 4 |
know. Negative results and open items are listed here rather than omitted, because a limitation that
|
| 5 |
is not written down is a limitation that will be discovered by someone else at the worst moment.
|
| 6 |
|
| 7 |
-
**Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN`.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
---
|
| 10 |
|
| 11 |
## 1. Model quality
|
| 12 |
|
| 13 |
-
|
| 14 |
-
|---|---|---|
|
| 15 |
-
| 1 | **Grounding IoU is low in absolute terms** | mean best IoU **0.2838** (canonical) / **0.2566** (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
|
| 16 |
-
| 2 | **Grounding is protocol-sensitive** | two protocols × two decode variants give very different numbers: head_threshold 0.2838, head_argmax **0.1215**, zero-shot 0.0972. An absolute value is meaningless without its protocol. |
|
| 17 |
-
| 3 | **Optical-SAR accuracy is carried by common classes** | accuracy **0.931** but macro-F1 **0.434161**. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. Never quote accuracy alone. |
|
| 18 |
-
| 4 | **Change-VQA is weak on rare classes** | accuracy 0.697626 / macro-F1 0.378373 (test) and 0.651469 / 0.372309 (test2). The wide accuracy–macroF1 gap is the signature of class imbalance. |
|
| 19 |
-
| 5 | **VQA is weak-but-related** | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
|
| 20 |
-
| 6 | **Optical-SAR returns a bare class index** | the live service returns `class_18`, not a human-readable CLC label. |
|
| 21 |
-
| 7 | **Calibration made things worse** | ECE **0.013755 → 0.014929** (`ece_improvement` −0.001174). Retained only because it is part of the frozen config — **not** because it helped. |
|
| 22 |
-
| 8 | **The VLM adapter is not accepted** | metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is **ACCEPTANCE-REJECTED**; the deployed caption/VQA path uses the **unadapted** model. |
|
| 23 |
-
| 9 | **The router number is not a test result** | 0.965116 is **validation**, **ungated**, **n = 86**, corpus-limited. The test split was **NOT RUN**. |
|
| 24 |
-
| 10 | **The router's routing is not perfect** | known residuals below (§2). |
|
| 25 |
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|---|---|---|---|
|
| 30 |
-
|
|
| 31 |
-
|
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
## 3. Evaluation gaps
|
| 35 |
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
|
| 49 |
## 4. Operational limitations
|
| 50 |
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
## 5. Packaging and licensing
|
| 64 |
|
| 65 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
|---|---|---|
|
| 67 |
-
|
|
| 68 |
-
|
|
| 69 |
-
|
|
| 70 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
## 6. Documentation caveats
|
| 73 |
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 81 |
|
| 82 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
These are things a reader might reasonably assume, which this project does **not** claim:
|
| 85 |
|
|
@@ -87,22 +713,33 @@ These are things a reader might reasonably assume, which this project does **not
|
|
| 87 |
- **No claim of production readiness for model quality.** The deployment runs; the models carry the
|
| 88 |
limitations above.
|
| 89 |
- **No claim that the trained heads generalise** beyond their training-family test splits.
|
| 90 |
-
- **No claim that calibration improves confidence** — it made ECE worse.
|
| 91 |
-
- **No claim that the VLM adapter is accepted** for production use.
|
| 92 |
- **No claim of an end-to-end accuracy number** — none exists.
|
| 93 |
-
- **No claim that the router is correct on all phrasings** — residuals exist.
|
|
|
|
|
|
|
| 94 |
- **No claim of robustness** to adversarial, corrupted, or out-of-distribution inputs.
|
| 95 |
- **No claim of geolocation accuracy** — grounding boxes are image-relative, not geodetic.
|
| 96 |
- **No claim that the system is a safety-, legal-, or life-critical tool.**
|
|
|
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
-
##
|
| 99 |
|
| 100 |
| Topic | Evidence |
|
| 101 |
|---|---|
|
| 102 |
| All measured metrics | `artifacts/**/*.json`, verified by `release/tools/verify_readme_metrics.py` |
|
| 103 |
| Metric honesty rules | [`BENCHMARKS.md`](BENCHMARKS.md), [`EVALUATION.md`](EVALUATION.md) |
|
| 104 |
| The grounding resolution rejection | `docs/PHASE7_RESOLUTION_DECISION.md` |
|
| 105 |
-
| The VLM rejection | `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`) |
|
| 106 |
-
|
|
| 107 |
-
|
|
| 108 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
know. Negative results and open items are listed here rather than omitted, because a limitation that
|
| 5 |
is not written down is a limitation that will be discovered by someone else at the worst moment.
|
| 6 |
|
| 7 |
+
**Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN` · `MEASURED`.
|
| 8 |
+
|
| 9 |
+
**How to read this document.** Limitations are numbered `L-01 …` and grouped by area. Each entry
|
| 10 |
+
states the limitation, the measured or observed detail behind it, and the file that records it. Where
|
| 11 |
+
a value is a status, it is stated exactly as the project's own records state it — a `REJECTED` is never
|
| 12 |
+
softened to "usable", an `OPEN` ruling is never presented as settled, and a validation number is never
|
| 13 |
+
promoted to a test result.
|
| 14 |
+
|
| 15 |
+
**Companion documents.** [`BENCHMARKS.md`](BENCHMARKS.md) and [`EVALUATION.md`](EVALUATION.md) hold
|
| 16 |
+
the metric honesty rules; [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) holds the findings behind several
|
| 17 |
+
entries here; [`DEPLOYMENT.md`](DEPLOYMENT.md) §8/§13 holds the operational blockers B-07 and B-02;
|
| 18 |
+
[`architecture/04-router.md`](architecture/04-router.md) documents the router residuals.
|
| 19 |
|
| 20 |
---
|
| 21 |
|
| 22 |
## 1. Model quality
|
| 23 |
|
| 24 |
+
### L-01 — Grounding IoU is low in absolute terms (`MEASURED`)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
+
The grounding head's mean best IoU is **0.2838** (canonical protocol) and **0.2566** (matched6
|
| 27 |
+
protocol) on VRSBench eval, n = 16,159. The trained head clearly beats the zero-shot baseline
|
| 28 |
+
(**0.0972**), but 0.28 is not "solved". Recall@0.5 is only **0.2198** (canonical) / **0.1938**
|
| 29 |
+
(matched6).
|
| 30 |
|
| 31 |
+
**Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`,
|
| 32 |
+
`eval_result_matched6.json`; `docs/PHASE7_RESOLUTION_DECISION.md`; `docs/FINAL_DELIVERY_TODO.md` §1.6.
|
| 33 |
+
|
| 34 |
+
### L-02 — Grounding is protocol-sensitive; an absolute value is meaningless without its protocol (`MEASURED`)
|
| 35 |
+
|
| 36 |
+
The same head reports very different numbers under different decode variants:
|
| 37 |
+
|
| 38 |
+
| Variant | Protocol | mean best IoU | recall@0.5 |
|
| 39 |
|---|---|---|---|
|
| 40 |
+
| `head_threshold` | canonical | 0.2838 | 0.2198 |
|
| 41 |
+
| `head_threshold` | matched6 | 0.2566 | 0.1938 |
|
| 42 |
+
| `head_argmax` | canonical | 0.1215 | 0.0795 |
|
| 43 |
+
| `zero_shot_matched` | canonical | 0.0972 | 0.0234 |
|
| 44 |
+
|
| 45 |
+
So a single grounding number quoted alone is misleading: the head/threshold versus head/argmax split
|
| 46 |
+
changes IoU by more than a factor of two. **Never quote one without the other.**
|
| 47 |
+
|
| 48 |
+
**Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json` and
|
| 49 |
+
`eval_result_matched6.json`; `release/DOCS_STYLE_GUIDE.md` §3.
|
| 50 |
+
|
| 51 |
+
### L-03 — The grounding head did not beat the zero-shot baseline on the validation curve (`MEASURED`)
|
| 52 |
+
|
| 53 |
+
The Benchmark page carries this as an honest label: the grounding head did not beat the baseline on
|
| 54 |
+
validation (`docs/FINAL_DELIVERY_TODO.md` §4 P6-T01). The head's advantage over zero-shot is a
|
| 55 |
+
**test-split** result (0.2838 vs 0.0972), not a validation result.
|
| 56 |
+
|
| 57 |
+
**Evidence:** `docs/FINAL_DELIVERY_TODO.md` §4 P6-T01; `docs/PHASE7_RESOLUTION_DECISION.md` ("the
|
| 58 |
+
Phase 8 head must beat 0.0972 to justify itself").
|
| 59 |
+
|
| 60 |
+
### L-04 — Optical-SAR accuracy is carried by common classes; macro-F1 is low (`MEASURED`, ruling `OPEN`)
|
| 61 |
+
|
| 62 |
+
The fusion head scores accuracy **0.931** but macro-F1 **0.434161** on a held-out test split of
|
| 63 |
+
n = 4,000 over 19 classes. **5 of the 19 classes are absent in the scored split** (`classes_absent:
|
| 64 |
+
[1, 11, 14, 15, 16]`) and contribute **0.0** to macro-F1 by construction
|
| 65 |
+
(`macro_f1_denominator: "all 19 classes (absent classes contribute 0.0)"`). The wide accuracy–macro-F1
|
| 66 |
+
gap is the signature of class imbalance. **Never quote accuracy without macro-F1.**
|
| 67 |
+
|
| 68 |
+
**Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`;
|
| 69 |
+
`docs/FINAL_DELIVERY_TODO.md` §1.6.
|
| 70 |
+
|
| 71 |
+
### L-05 — The optical-SAR ruling is `OPEN` (`OPEN`)
|
| 72 |
+
|
| 73 |
+
`pre_registered_115_metric.json` states plainly: *"This tool reports ONE head's held-out accuracy and
|
| 74 |
+
macro-F1. It selects no head, ranks nothing and compares no arms. Whether this constitutes a Phase 12
|
| 75 |
+
pass is the owner's ruling."* Phase 12 is **INCOMPLETE**; the pre-registered metric's gate criterion is
|
| 76 |
+
owner-gated (`docs/PHASE12_CURRENCY_CORRECTION.md` §0, §3; `docs/STEP7_BACKEND_CHAIN_REPORT.md` §17).
|
| 77 |
+
|
| 78 |
+
**Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`
|
| 79 |
+
(`is_deciding_statistic: false`); `docs/PHASE12_CURRENCY_CORRECTION.md`.
|
| 80 |
+
|
| 81 |
+
### L-06 — Change-VQA is weak on rare classes, on two test sets (`MEASURED`, ruling `OPEN`)
|
| 82 |
+
|
| 83 |
+
Change-VQA scores **accuracy 0.697626 / macro-F1 0.378373** on `Test` (n = 39,686) and
|
| 84 |
+
**0.651469 / 0.372309** on `Test2` (n = 31,036). The wide accuracy–macroF1 gap is the signature of
|
| 85 |
+
class imbalance. The ruling is **`OPEN`** (`docs/PHASE19_FINAL_HARDENING.md` §4.5: "the R-02
|
| 86 |
+
macro-F1/accuracy gap is unruled … not this work order's to rule"). Top-3 accuracy is 0.964698. The
|
| 87 |
+
head's metadata records `confidence_method: "uncalibrated"` and `test_splits_used: false` on the
|
| 88 |
+
training record itself.
|
| 89 |
+
|
| 90 |
+
**Evidence:** `artifacts/change_vqa/run/PROMOTION.json` (`test_accuracy`, `test_macro_f1`,
|
| 91 |
+
`test2_accuracy`); `artifacts/change_vqa/run/model_metadata.json`; `docs/PHASE19_FINAL_HARDENING.md`
|
| 92 |
+
§4.5.
|
| 93 |
+
|
| 94 |
+
### L-07 — VQA is weak-but-related (`MEASURED`)
|
| 95 |
+
|
| 96 |
+
The live VQA path answers broadly related content rather than a crisp class. Measured live: the case
|
| 97 |
+
A1 query answered **"Grassland"** for a scene where a more specific answer was expected
|
| 98 |
+
(`LIVE_VALIDATION_POSTFIX.md`, "Model-quality note"; `release/CURRENT_RELEASE_STATE.md` §6). This is a
|
| 99 |
+
model-quality limitation, not a deployment fault — the pipeline dispatches and returns a real answer.
|
| 100 |
+
|
| 101 |
+
**Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md`.
|
| 102 |
+
|
| 103 |
+
### L-08 — Optical-SAR returns a bare class index, not a human label (`MEASURED`)
|
| 104 |
+
|
| 105 |
+
The live service returns `class_18` rather than a human-readable CLC label. The live answer reads
|
| 106 |
+
`[optical_sar] Fused optical-SAR prediction: class_18 (margin 1.000; optical channels 4/12, SAR
|
| 107 |
+
channels 2/2)` — the modality accounting confirms the right channels reached the fusion head, but the
|
| 108 |
+
answer is not interpretable without a label map.
|
| 109 |
+
|
| 110 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6;
|
| 111 |
+
`.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md`.
|
| 112 |
+
|
| 113 |
+
### L-09 — Calibration made ECE worse (`MEASURED`)
|
| 114 |
+
|
| 115 |
+
Temperature scaling moved ECE from **0.013755 → 0.014929** (`ece_improvement: −0.001174` — **worse**)
|
| 116 |
+
while improving NLL marginally (0.689741 → 0.689631). It is **retained only because it is part of the
|
| 117 |
+
frozen config** — **not** because it helped. The fitted temperature is `0.9772731820958189`, fitted on
|
| 118 |
+
the Val split with n = 16,441, `effective: true`, `hit_bound: false`. The path is live
|
| 119 |
+
(`core/controller.py` calls `load_calibration(config)` and hands the artifact to `EvidenceEngine`), so
|
| 120 |
+
a deployed result carries a calibrated value — but the calibration is **not an improvement** and must
|
| 121 |
+
never be described as making confidence "more accurate".
|
| 122 |
+
|
| 123 |
+
**Evidence:** `artifacts/calibration_v001.json` (`metrics.ece_before`, `metrics.ece_after`,
|
| 124 |
+
`fit_diagnostics.temperature`); `docs/STEP7_BACKEND_CHAIN_REPORT.md` §7;
|
| 125 |
+
`docs/PHASE19_FINAL_HARDENING.md` §9 (calibration-success correction).
|
| 126 |
+
|
| 127 |
+
### L-10 — The VLM adapter is `ACCEPTANCE-REJECTED` (`REJECTED`)
|
| 128 |
+
|
| 129 |
+
The Phase-6 LoRA adapter's metrics are **usable** — test exact-match **0.963**, F1 **0.96432**,
|
| 130 |
+
aggregate test delta **+49.5 pp** — and the artifact is `USABLE_VERIFIED`. But its acceptance status is
|
| 131 |
+
**`REJECTED`**: v001 rejected it on val, and the independent-test rule v002 rejected it on the test
|
| 132 |
+
split (one class, `Mixed forest`, lost 4 questions at z = 2.1335). The closure record keeps the two
|
| 133 |
+
questions separate: *"'Verified' answers: is this artifact the one we trained, and does it work?
|
| 134 |
+
'Accepted' answers: did it clear the bar predeclared before we looked?"* The deployed caption/VQA path
|
| 135 |
+
uses the **unadapted** model by default; the adapter is attached only when `SATQUERY_VLM_ADAPTER` is
|
| 136 |
+
set. **USABLE ≠ ACCEPTED.**
|
| 137 |
+
|
| 138 |
+
**Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.acceptance_status`,
|
| 139 |
+
`why_acceptance_rejected`, `what_closure_does_not_claim`);
|
| 140 |
+
`docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md`.
|
| 141 |
+
|
| 142 |
+
### L-11 — The VLM acceptance-rule history is a post-hoc rule change (`MEASURED`, disclosed)
|
| 143 |
+
|
| 144 |
+
v001 was pre-registered (`declared_before_training: true`) and rejected the run on a per-class guardrail
|
| 145 |
+
whose 1.0 pp threshold sits **below the measurement resolution** of the data (0.20 questions at n = 20;
|
| 146 |
+
per-class SE 2.6–10.6 pp). v002 was declared **after** run 1 (`declared_before_training: false`) and is
|
| 147 |
+
documented as a **relaxation** of V2, justified measurement-theoretically, not by the outcome. The
|
| 148 |
+
record keeps v001 and its `REJECTED` verdict verbatim and states that v002 "is not a numerically
|
| 149 |
+
stricter bar" than the contract's ~3 SE figure. A reader must treat the acceptance verdict as resting
|
| 150 |
+
on 4 questions in one class of 33 — the "unfloored minimum-size exposure" the closure reports but does
|
| 151 |
+
not resolve.
|
| 152 |
+
|
| 153 |
+
**Evidence:** `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md` §4, §7.1, §7.4;
|
| 154 |
+
`artifacts/vlm/phase6_closure.json` (`thresholds_used.amendment`, `residual_risk`).
|
| 155 |
+
|
| 156 |
+
### L-12 — The router number is a validation result, ungated, n = 86 (`MEASURED`, `TEST NOT RUN`)
|
| 157 |
+
|
| 158 |
+
Router task accuracy is **0.965116** — measured on **validation**, **ungated**, **n = 86**,
|
| 159 |
+
**corpus-limited** (`corpus_total: 576`, `corpus_groups: 54`; `corpus_limited: true`). The **test
|
| 160 |
+
split was NOT RUN**. The number is indicative only and must never be quoted as a test result.
|
| 161 |
+
|
| 162 |
+
**Evidence:** `artifacts/router/router_adapter_v001/metadata.json` (`val_task_accuracy`);
|
| 163 |
+
`artifacts/router/threshold_sweep_val.json`; `docs/FINAL_DELIVERY_TODO.md` §1.6.
|
| 164 |
+
|
| 165 |
+
### L-13 — The router's routing is not perfect (known residuals) (`MEASURED`)
|
| 166 |
+
|
| 167 |
+
See §2. The router is a lexical/embedding classifier over a small corpus; residual misroutes exist and
|
| 168 |
+
are documented rather than hidden.
|
| 169 |
+
|
| 170 |
+
---
|
| 171 |
+
|
| 172 |
+
## 2. Router residuals (known misroutes)
|
| 173 |
+
|
| 174 |
+
### L-14 — *"What is the new runway?"* reads `change`, not `vqa` (`OPEN`)
|
| 175 |
+
|
| 176 |
+
The `new`-as-change heuristic fires on non-`where` questions. This is "strictly better than pre-fix,
|
| 177 |
+
where `new` was unconditionally temporal. A lexical router cannot cleanly separate 'the new X' from
|
| 178 |
+
'what's new'."
|
| 179 |
+
|
| 180 |
+
**Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
|
| 181 |
+
`release/CURRENT_RELEASE_STATE.md` §6.
|
| 182 |
+
|
| 183 |
+
### L-15 — *"How much built-up area was added?"* reads `vqa` (under-trigger) (`OPEN`)
|
| 184 |
+
|
| 185 |
+
`built` was dropped from the temporal set during the B-08 fix and `area` no longer matches inside
|
| 186 |
+
`areas`, so a change-style quantifier is not caught and the query under-triggers to `vqa`.
|
| 187 |
+
|
| 188 |
+
**Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
|
| 189 |
+
`docs/FINAL_DELIVERY_TODO.md` §5 B-08.
|
| 190 |
+
|
| 191 |
+
### L-16 — The `interpret()` / `chooseTask()` asymmetry is visually surprising (`RESOLVED`, intentional)
|
| 192 |
+
|
| 193 |
+
For *"What changed between the earlier and later image?"* with **one** asset attached, the console
|
| 194 |
+
**reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is **intentional** —
|
| 195 |
+
the reading is asset-count-blind (it describes the question's intent), while dispatch is
|
| 196 |
+
asset-count-aware (it respects what can actually be computed with the assets present) — but a reader
|
| 197 |
+
who sees the reading panel and the answer disagree may mistake it for a defect.
|
| 198 |
+
|
| 199 |
+
**Evidence:** [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §7;
|
| 200 |
+
`.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Discriminator note").
|
| 201 |
+
|
| 202 |
+
---
|
| 203 |
|
| 204 |
## 3. Evaluation gaps
|
| 205 |
|
| 206 |
+
### L-17 — No system-level end-to-end benchmark exists (`NOT RUN`)
|
| 207 |
+
|
| 208 |
+
**No end-to-end accuracy is claimed anywhere.** `STATUS.md` states there is no system-level E2E
|
| 209 |
+
benchmark, and the Benchmark page is required to label it `NOT RUN`.
|
| 210 |
+
|
| 211 |
+
**Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.7 item 7, §5 B-04; `docs/FINAL_DELIVERY_REPORT.md` §6.
|
| 212 |
+
|
| 213 |
+
### L-18 — The router test split was not run (`NOT RUN`)
|
| 214 |
+
|
| 215 |
+
See L-12. The test split exists but was never scored.
|
| 216 |
+
|
| 217 |
+
**Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.6; `docs/FINAL_DELIVERY_REPORT.md` §6.
|
| 218 |
+
|
| 219 |
+
### L-19 — The benchmark adapters were not run (`NOT RUN`)
|
| 220 |
+
|
| 221 |
+
No benchmark adapter is registered; the registry is empty by construction
|
| 222 |
+
(`docs/PHASE19_FINAL_HARDENING.md` §9, "Benchmark pass — Not claimed").
|
| 223 |
+
|
| 224 |
+
**Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §9; `docs/FINAL_DELIVERY_REPORT.md` §6.
|
| 225 |
+
|
| 226 |
+
### L-20 — No end-to-end latency benchmark (`NOT RUN`)
|
| 227 |
+
|
| 228 |
+
Per-specialist latency is recorded only incidentally (e.g. grounding encoder latency 2.205 ms/image at
|
| 229 |
+
224 in the canonical eval, 20.0 ms/image on a T4 per `docs/PHASE7_RESOLUTION_DECISION.md`). There is no
|
| 230 |
+
benchmark of the deployed request path across the four tiers.
|
| 231 |
+
|
| 232 |
+
**Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`;
|
| 233 |
+
`docs/PHASE7_RESOLUTION_DECISION.md`; [`PERFORMANCE.md`](PERFORMANCE.md).
|
| 234 |
+
|
| 235 |
+
### L-21 — No cross-dataset generalisation (`NOT RUN`)
|
| 236 |
+
|
| 237 |
+
Each specialist is evaluated only on its own training-family test split (LEVIR-CD for change,
|
| 238 |
+
VRSBench for grounding, a BigEarthNet-family split for fusion, CDVQA for change-VQA). Nothing measures
|
| 239 |
+
transfer to a different distribution.
|
| 240 |
+
|
| 241 |
+
**Evidence:** `docs/PHASE7_RESOLUTION_DECISION.md` ("Anything about hidden ISRO/SAC imagery … a
|
| 242 |
+
different distribution entirely"); the per-task evaluation sections of [`EVALUATION.md`](EVALUATION.md).
|
| 243 |
+
|
| 244 |
+
### L-22 — No human evaluation (`NOT RUN`)
|
| 245 |
+
|
| 246 |
+
No human study of answer quality, usefulness, or failure modes was performed.
|
| 247 |
+
|
| 248 |
+
**Evidence:** `docs/FINAL_DELIVERY_REPORT.md` §6 (unverified items); this document is the only
|
| 249 |
+
catalogue.
|
| 250 |
+
|
| 251 |
+
### L-23 — No robustness or adversarial evaluation (`NOT RUN`)
|
| 252 |
+
|
| 253 |
+
No evaluation of behaviour under adversarial, corrupted, or out-of-distribution inputs. The
|
| 254 |
+
change specialist carries invalid-data and registration-quality logic
|
| 255 |
+
(`docs/ARCHITECTURE_FREEZE.md` §23 false-change handling), but no robustness *evaluation* exists.
|
| 256 |
+
|
| 257 |
+
**Evidence:** `docs/FINAL_DELIVERY_REPORT.md` §6.
|
| 258 |
+
|
| 259 |
+
### L-24 — BigEarthNet label semantics make the local metrics non-comparable (`MEASURED`)
|
| 260 |
+
|
| 261 |
+
The local BigEarthNet subset is **100 % single-label** against the official **1–11 multi-label**
|
| 262 |
+
scheme, so metrics computed on it are **not comparable** to published multi-label numbers.
|
| 263 |
+
|
| 264 |
+
**Evidence:** [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §9; `docs/PHASE12_LABEL_POLICY_DECISION.md`.
|
| 265 |
+
|
| 266 |
+
### L-25 — The reliability diagram shipped is the pre-scaling curve (`MEASURED`)
|
| 267 |
+
|
| 268 |
+
The shipped reliability curve is the **pre-scaling** curve (labelled as such); the calibrated curve is
|
| 269 |
+
not plotted. The caption now reads "Measured" and the bins are transcribed from
|
| 270 |
+
`artifacts/calibration_v001.json`, but the remaining section-03 PR curves are still labelled
|
| 271 |
+
illustrative.
|
| 272 |
+
|
| 273 |
+
**Evidence:** `docs/FINAL_DELIVERY_TODO.md` §4 P6-T01 note; `artifacts/calibration_v001.json`.
|
| 274 |
+
|
| 275 |
+
### L-26 — Statistical significance exists for only one decision (`MEASURED`)
|
| 276 |
+
|
| 277 |
+
Only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval
|
| 278 |
+
(§[`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §3). Every other per-task number is a point estimate with no
|
| 279 |
+
significance test.
|
| 280 |
+
|
| 281 |
+
**Evidence:** `docs/PHASE7_RESOLUTION_DECISION.md` ("Paired analysis — independent confirmation").
|
| 282 |
+
|
| 283 |
+
### L-27 — The optical-SAR registry state is `degraded`, and the adapter never advertises `loaded` (`MEASURED`)
|
| 284 |
+
|
| 285 |
+
The registry resolves `optical_sar` to **`degraded`**, not `available`
|
| 286 |
+
(`docs/STEP7_BACKEND_CHAIN_REPORT.md` §17). More generally, the capability adapter derives contract
|
| 287 |
+
state from **artifact presence** and therefore **never emits `loaded` or `evicted`** — "a model is
|
| 288 |
+
resident" is unknowable without loading one, which the metadata path must not do
|
| 289 |
+
(`docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1). A consumer cannot learn from `/v1/capabilities` whether a
|
| 290 |
+
model is actually resident.
|
| 291 |
+
|
| 292 |
+
**Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §17; `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1.
|
| 293 |
+
|
| 294 |
+
### L-28 — The trace's registry block is a snapshot of the *previous* request (`MEASURED`)
|
| 295 |
+
|
| 296 |
+
`SpecialistRegistry.describe()` is assigned into the trace **immediately after planning** and **before
|
| 297 |
+
execution**, so on a cold process the client-visible `built` map is necessarily `{}` — including for
|
| 298 |
+
the request that is about to build the specialist. The same query returns two different traces
|
| 299 |
+
depending on how many requests the process has already served. The same trace is also internally
|
| 300 |
+
inconsistent about which clock it uses (`selected_models` is assigned inside execution, so it *does*
|
| 301 |
+
reflect the current request while `built` does not). **Severity: low — observability only.**
|
| 302 |
+
|
| 303 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.8 (finding F-19).
|
| 304 |
+
|
| 305 |
+
### L-29 — The uploaded-asset path surface has historically leaked server-side paths (`RESOLVED`, recorded)
|
| 306 |
+
|
| 307 |
+
A series of findings (F-13 … F-16) recorded that `trace.inputs`, `trace.steps[PARSE].detail`, and
|
| 308 |
+
`evidence[].artifact_ref` / `result.change_map` published server-side filesystem paths to an
|
| 309 |
+
unauthenticated client. All were fixed (paths reduced to basenames; refs set to `null` with an explicit
|
| 310 |
+
non-retrievable warning; no fabricated `artifact://` URIs). Two consequences remain worth recording:
|
| 311 |
+
the contract's §2.4 example once showed a fabricated `artifact://` URI the service cannot emit (now
|
| 312 |
+
corrected), and the ruling introduced a **signalling** change (F-16c): on a deployment that configures
|
| 313 |
+
`change.artifact_dir`, a no-change run now reports `degraded: true` where it previously reported
|
| 314 |
+
`degraded: false`.
|
| 315 |
+
|
| 316 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.3–§5.6; [`SECURITY.md`](SECURITY.md).
|
| 317 |
+
|
| 318 |
+
---
|
| 319 |
|
| 320 |
## 4. Operational limitations
|
| 321 |
|
| 322 |
+
### L-30 — B-07: transient tunnel gaps (`OPEN`)
|
| 323 |
+
|
| 324 |
+
A request can hang or return `504` when the tunnel agent is briefly absent. The patch is prepared,
|
| 325 |
+
**NOT deployed**.
|
| 326 |
+
|
| 327 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1; `docs/FINAL_DELIVERY_TODO.md` §5 B-07;
|
| 328 |
+
`release/CURRENT_RELEASE_STATE.md` §6.
|
| 329 |
+
|
| 330 |
+
### L-31 — B-07 root shape: the `auto`-mode fallthrough wastes the wake budget (`MEASURED`)
|
| 331 |
+
|
| 332 |
+
In `auto` transport mode a tunnel timeout **falls through** to the forward path, which then burns
|
| 333 |
+
`wake_timeout_s` (120 s) on a `302` → a worst case of ≈ **249 s** (150 + 120). Measured.
|
| 334 |
+
|
| 335 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6; [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §6;
|
| 336 |
+
[`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1.
|
| 337 |
+
|
| 338 |
+
### L-32 — Cold start is tens of seconds (`MEASURED`)
|
| 339 |
+
|
| 340 |
+
Render's free tier sleeps when idle, and the Codespace may be stopped (idle timeout 30 min). The first
|
| 341 |
+
request after idle waits for a wake. Documented, not hidden.
|
| 342 |
+
|
| 343 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8; `docs/DEPLOYMENT_TOPOLOGY.md` §2.
|
| 344 |
+
|
| 345 |
+
### L-33 — B-02: `codespace_name` trailing newline (`OPEN`, cosmetic)
|
| 346 |
+
|
| 347 |
+
The `/api/health` payload reports the raw `codespace_name` with a trailing `\n`. Cosmetic; the wake
|
| 348 |
+
path strips it.
|
| 349 |
+
|
| 350 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.2; `docs/FINAL_DELIVERY_TODO.md` §4 P2-T03.
|
| 351 |
+
|
| 352 |
+
### L-34 — No database, auth, or queue (`BY DESIGN`)
|
| 353 |
+
|
| 354 |
+
The gateway is stateless by design. There is no persistence of runs or users, no auth layer, and no
|
| 355 |
+
request queue (plan §73/§74; `docs/DEPLOYMENT_ARCHITECTURE.md` §6).
|
| 356 |
+
|
| 357 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2, §6.
|
| 358 |
+
|
| 359 |
+
### L-35 — The deployment repositories are private (`BY DESIGN`)
|
| 360 |
+
|
| 361 |
+
The three deploy repositories return `404` for an outside audience.
|
| 362 |
+
|
| 363 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.4; `docs/FINAL_DELIVERY_TODO.md` §4 P9-T01.
|
| 364 |
+
|
| 365 |
+
### L-36 — `deploy/` in the monorepo is stale and untracked (`OPEN` trap)
|
| 366 |
+
|
| 367 |
+
Edits there do not deploy; it is not the deployed source.
|
| 368 |
+
|
| 369 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §1.1; `docs/FINAL_DELIVERY_TODO.md` §1.1.
|
| 370 |
+
|
| 371 |
+
### L-37 — The rate limiter is a fairness control, not a protection control (`BY DESIGN`)
|
| 372 |
+
|
| 373 |
+
The per-IP limiter keys on the first hop of the client-supplied `X-Forwarded-For`; a caller that varies
|
| 374 |
+
the header is never throttled (measured 0/8 throttled with a fresh value per request, versus 5/8
|
| 375 |
+
without). It is explicitly **not** a security boundary.
|
| 376 |
+
|
| 377 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2; `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2.
|
| 378 |
+
|
| 379 |
+
### L-38 — No APM, distributed tracing, or cost accounting (`BY DESIGN`)
|
| 380 |
+
|
| 381 |
+
Observability is limited to the health payload and per-run traces. There is no APM, no distributed
|
| 382 |
+
tracing across the four tiers, and no cost accounting.
|
| 383 |
+
|
| 384 |
+
**Evidence:** [`architecture/10-observability-and-ops.md`](architecture/10-observability-and-ops.md);
|
| 385 |
+
[`OPERATIONS.md`](OPERATIONS.md).
|
| 386 |
+
|
| 387 |
+
### L-39 — Single-region, no HA (`BY DESIGN`)
|
| 388 |
+
|
| 389 |
+
One Render service, one Codespace. No redundancy, no failover, no multi-region deployment.
|
| 390 |
+
|
| 391 |
+
**Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §13.
|
| 392 |
+
|
| 393 |
+
### L-40 — A saturated asset store is indistinguishable from a misconfigured one (`OPEN`)
|
| 394 |
+
|
| 395 |
+
`POST /v1/assets` answers `503` both when the store is unconfigured and when it is full; the response
|
| 396 |
+
cannot tell them apart, and the counter that would have separated them was **removed** rather than given
|
| 397 |
+
a consumer. Treat the `503` on this route as ambiguous.
|
| 398 |
+
|
| 399 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 (finding F-11).
|
| 400 |
+
|
| 401 |
+
### L-41 — `change_vqa.artifact_dir` is a configured-but-inert key (`OPEN`, low severity)
|
| 402 |
+
|
| 403 |
+
`change_vqa.artifact_dir` is advertised in the same `optional_config_keys` table as the two keys that
|
| 404 |
+
work, but `specialists/change/vqa_specialist.py` **never reads it** (the attribute occurs exactly once,
|
| 405 |
+
as an assignment; the module contains no file-writing code). An operator who sets it receives no
|
| 406 |
+
artifacts and **no warning**.
|
| 407 |
+
|
| 408 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.6 (finding F-17).
|
| 409 |
+
|
| 410 |
+
### L-42 — The Anatomy page's plate image variant is cosmetic-wrong (`OPEN`, cosmetic)
|
| 411 |
+
|
| 412 |
+
The Anatomy plate points at the **720×720** variant of an image the recorded run analysed at
|
| 413 |
+
**730×730**. The content is identical and the canvas scales it, but the page's "the ACTUAL analysed
|
| 414 |
+
image" wording is very slightly loose.
|
| 415 |
+
|
| 416 |
+
**Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
|
| 417 |
+
`release/CURRENT_RELEASE_STATE.md` §6.
|
| 418 |
+
|
| 419 |
+
### L-43 — The `/api/capabilities` deployment block carried stale metadata (`RESOLVED`)
|
| 420 |
+
|
| 421 |
+
The capabilities `deployment` block once claimed `huggingface-spaces`/`zerogpu`. It was recorded as a
|
| 422 |
+
stale-metadata item (`docs/FINAL_DELIVERY_TODO.md` §1.7 item 6) and is reconciled in the shipped
|
| 423 |
+
contract; the frozen `configs/deploy.yaml` still describes the superseded target (§5 below).
|
| 424 |
+
|
| 425 |
+
**Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.7 item 6; [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
|
| 426 |
+
|
| 427 |
+
---
|
| 428 |
|
| 429 |
## 5. Packaging and licensing
|
| 430 |
|
| 431 |
+
### L-44 — There is no `LICENSE` file (`OPEN`)
|
| 432 |
+
|
| 433 |
+
**No `LICENSE` file exists** in the source repository. The README says "add a license file before
|
| 434 |
+
public release". A licence must be selected before public release of the code. This is a **release
|
| 435 |
+
blocker for the code**, not a model defect.
|
| 436 |
+
|
| 437 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6 ("**No `LICENSE` file exists** in the monorepo;
|
| 438 |
+
README says 'add a license file before public release'"); `release/DOCS_STYLE_GUIDE.md` §3.
|
| 439 |
+
|
| 440 |
+
### L-45 — Backbones are not redistributed (`BY DESIGN`)
|
| 441 |
+
|
| 442 |
+
The six trained artifacts are small modules; the backbones (SmolVLM, RemoteCLIP, MiniLM, CROMA,
|
| 443 |
+
STANet-style encoder) are fetched from their sources at run time, and their licences are their own.
|
| 444 |
+
Nothing here re-licenses a backbone.
|
| 445 |
+
|
| 446 |
+
**Evidence:** [`MODELS.md`](MODELS.md); `app/serving.py` (`resolve_checkpoint_path`, "offline first");
|
| 447 |
+
`docs/ARCHITECTURE_FREEZE.md` §2.
|
| 448 |
+
|
| 449 |
+
### L-46 — The released weights require their backbones (`MEASURED`)
|
| 450 |
+
|
| 451 |
+
A consumer of the six artifacts must also fetch the pinned backbones at the exact revisions; an
|
| 452 |
+
artifact alone is not runnable. The VLM adapter, for example, requires
|
| 453 |
+
`HuggingFaceTB/SmolVLM-500M-Instruct` (revision `a7da5b986cb5`) plus the tokenizer/processor files
|
| 454 |
+
listed in `phase6_closure.json`.
|
| 455 |
+
|
| 456 |
+
**Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.files_required_to_serve_standalone`);
|
| 457 |
+
`release/CURRENT_RELEASE_STATE.md` §3.
|
| 458 |
+
|
| 459 |
+
### L-47 — Artifact-tree bloat (`MEASURED`)
|
| 460 |
+
|
| 461 |
+
`artifacts/` totals roughly **3.7 GB**, almost all of it caches, duplicates and features rather than
|
| 462 |
+
released weights:
|
| 463 |
+
|
| 464 |
+
| Path | Size | Classification |
|
| 465 |
|---|---|---|
|
| 466 |
+
| `artifacts/grounding/remoteclip_grounding_v001/` | 1.4 GB | ARCHIVED (evidence) |
|
| 467 |
+
| `artifacts/grounding/remoteclip_grounding_v001.zip` | 774 MB | **DUPLICATE** of the directory above |
|
| 468 |
+
| `artifacts/optical_sar/fusion_head_v001/` | 276 MB | superseded by `fusion_head_production_v001` |
|
| 469 |
+
| `artifacts/optical_sar/fusion_features/` | 231 MB | reproducible cache |
|
| 470 |
+
| `artifacts/optical_sar/fusion_features_armB/` | 231 MB | reproducible cache |
|
| 471 |
+
| `artifacts/change/levir_change_cpu_probe_v001/` | 241 MB | **DUPLICATE** probe of `levir_change_v001` |
|
| 472 |
+
| `artifacts/change/levir_change_v001/` | 241 MB | ARCHIVED |
|
| 473 |
+
| `artifacts/phase12_selection/*.jsonl` | ~235 MB | data-selection manifests (seeds) |
|
| 474 |
+
|
| 475 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §3.
|
| 476 |
+
|
| 477 |
+
### L-48 — The VLM adapter is not committed and lives under `.scratch/` (`MEASURED`)
|
| 478 |
+
|
| 479 |
+
The adapter is **not** committed (`.gitignore` excludes `artifacts/`, `checkpoints/` and
|
| 480 |
+
`*.safetensors`) and its canonical path is under `.scratch/`, which a future cleanup could remove. The
|
| 481 |
+
closure record states that moving it to a non-scratch location is "a reasonable follow-up, not a
|
| 482 |
+
closure requirement", and that moving it would make the recorded evidence stale.
|
| 483 |
+
|
| 484 |
+
**Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.reconstruction.why_not_moved`).
|
| 485 |
+
|
| 486 |
+
### L-49 — The frozen `configs/deploy.yaml` describes a target that does not exist (`OPEN` paperwork)
|
| 487 |
+
|
| 488 |
+
`configs/deploy.yaml` still declares `platform: huggingface-spaces`, `sdk: gradio`, `zerogpu: true`,
|
| 489 |
+
and the `gpu_duration_*` values. It is inert (`registry: false`), never loaded by `core/config.py`, and
|
| 490 |
+
cannot be edited without either failing `scripts/validate_deploy_config.py` or moving `Config.hash`. It
|
| 491 |
+
is deliberately left undisturbed, but a reader who finds it will reasonably think the project targets a
|
| 492 |
+
Gradio ZeroGPU Space.
|
| 493 |
+
|
| 494 |
+
**Evidence:** `configs/deploy.yaml`; `docs/DEPLOYMENT_DECISION.md` §4; [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
|
| 495 |
+
|
| 496 |
+
---
|
| 497 |
|
| 498 |
## 6. Documentation caveats
|
| 499 |
|
| 500 |
+
### L-50 — `docs/FINAL_DELIVERY_REPORT.md` §6 is stale (`SUPERSEDED`)
|
| 501 |
+
|
| 502 |
+
It still lists the bundled EO change pair as **DEGRADED** (726² vs 736² → shape error) and **B-01** as
|
| 503 |
+
**BLOCKED**. Both were resolved on 2026-09-25: the EO pair is now same-shape (both 720×720, new `-720`
|
| 504 |
+
URLs) → **RESOLVED**; the Hugging Face link is live on all 11 pages → **B-01 CLOSED**.
|
| 505 |
+
|
| 506 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6 ("Documentation that this sprint supersedes");
|
| 507 |
+
`docs/FINAL_DELIVERY_REPORT.md` §6.
|
| 508 |
+
|
| 509 |
+
### L-51 — The monorepo `README.md` was materially stale (`SUPERSEDED`)
|
| 510 |
+
|
| 511 |
+
It described a hermetic frontend, an in-progress Render/Codespace, a `/v1/*` contract, omitted the
|
| 512 |
+
tunnel, and pointed at the stale `deploy/`. Superseded by this release's README.
|
| 513 |
+
|
| 514 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6.
|
| 515 |
+
|
| 516 |
+
### L-52 — The `hf/` docs were stale (`SUPERSEDED`)
|
| 517 |
+
|
| 518 |
+
`hf/SETUP.md` and `hf/README.md` asserted the project owns no weights and has no HF credentials. Both
|
| 519 |
+
were false at release time.
|
| 520 |
+
|
| 521 |
+
**Evidence:** `release/CURRENT_RELEASE_STATE.md` §6.
|
| 522 |
+
|
| 523 |
+
### L-53 — Stale-negative documentation is a structural hazard (`MEASURED`)
|
| 524 |
+
|
| 525 |
+
Three documents written after the Phase-12 A/B experiment still described it as **un-run**, even though
|
| 526 |
+
it had completed 34 hours earlier. The lesson recorded: "a stale negative claim is more dangerous than
|
| 527 |
+
a stale positive one … 'This has never been executed' is contradicted by *nothing* — no test fails, no
|
| 528 |
+
hash moves, no invariant breaks." The existing conformance tests check that documented things *exist*;
|
| 529 |
+
**no test can check that a documented absence is still absent**.
|
| 530 |
+
|
| 531 |
+
**Evidence:** `docs/PHASE12_CURRENCY_CORRECTION.md` §0, §6.
|
| 532 |
+
|
| 533 |
+
### L-54 — The original master plan describes a superseded deployment (`SUPERSEDED`)
|
| 534 |
+
|
| 535 |
+
The master plan specifies a Gradio GUI + HF Space + ZeroGPU + Railway, and a single-image workflow set.
|
| 536 |
+
The shipped system is a static frontend + Render + Codespace tunnel, serving JSON, with the change-VQA
|
| 537 |
+
and optical-SAR capabilities added later. The plan is a design document, not a description of the
|
| 538 |
+
shipped system.
|
| 539 |
+
|
| 540 |
+
**Evidence:** `docs/MASTER_ARCHITECTURE_PLAN.md`; `docs/DEPLOYMENT_DECISION.md` §4;
|
| 541 |
+
[`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
|
| 542 |
|
| 543 |
+
### L-55 — `mask_ref` was wrongly listed as a documented path surface (`CORRECTED`)
|
| 544 |
+
|
| 545 |
+
An audit note listed `Region.mask_ref` as a deliberate documented path surface; that was wrong —
|
| 546 |
+
`core/schemas.py` is a bare `str | None = None` with no description, and `grep -r mask_ref docs/` finds
|
| 547 |
+
nothing. Recorded so the correction is not lost.
|
| 548 |
+
|
| 549 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5 (note on `mask_ref`).
|
| 550 |
+
|
| 551 |
+
### L-56 — Documentation-validated-by-execution found nine factual errors (`MEASURED`)
|
| 552 |
+
|
| 553 |
+
Validating every JSON example against the real Pydantic models and every runbook claim against the
|
| 554 |
+
repository caught nine factual errors proof-reading had missed — including `CoordinateSystem`
|
| 555 |
+
documented as `normalized`/`geographic` when the real values are `normalized_0_1`/`geo`, and `Box`
|
| 556 |
+
documented with nested geometry when the real model is flat. The checks are now permanent tests (48
|
| 557 |
+
assertions). This is a caveat about how much confidence a *document* can carry: prose is not validated
|
| 558 |
+
by these tests.
|
| 559 |
+
|
| 560 |
+
**Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §3.6; `docs/STEP8_FINAL_CONFORMANCE_AUDIT.md` §13.
|
| 561 |
+
|
| 562 |
+
---
|
| 563 |
+
|
| 564 |
+
## 7. Interface, contract, and data-surface limitations
|
| 565 |
+
|
| 566 |
+
### L-57 — The contract promises forward-compatible reading; the schema forbids it (`OPEN`, contract contradiction)
|
| 567 |
+
|
| 568 |
+
`docs/API_CONTRACT.md` §1.1 says a consumer **"MUST tolerate unknown fields on read (forward
|
| 569 |
+
compatibility)"**, and §1 says additive changes do not bump the version — which only works if unknown
|
| 570 |
+
fields are ignorable. But `ResultEnvelope`, `SpecialistResult` and `ExecutionTrace` are all
|
| 571 |
+
`extra="forbid"`, so `model_validate` **rejects** a body carrying an unknown key. Both halves are
|
| 572 |
+
load-bearing and cannot both hold. The contract's own authority clause makes the **code** authoritative,
|
| 573 |
+
so the document's bullet is the inaccurate half — but the document was **not** corrected (a doc edit is
|
| 574 |
+
a reviewable change), so a reader of the contract is still told the wrong thing.
|
| 575 |
+
|
| 576 |
+
**Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §15 (finding C-2);
|
| 577 |
+
[`architecture/08-api-contract.md`](architecture/08-api-contract.md).
|
| 578 |
+
|
| 579 |
+
### L-58 — The capability adapter can never report `loaded` or `evicted` (`MEASURED`)
|
| 580 |
+
|
| 581 |
+
Because the metadata path must not construct a model, `app/deployment.py` derives contract state from
|
| 582 |
+
artifact **presence** and reconstructs the registry's word from that state — the inverse direction. One
|
| 583 |
+
observable consequence: **`loaded` and `degraded` are never emitted**, since "a model is resident" is
|
| 584 |
+
unknowable without loading one, and `evicted` is a runtime model-cache fact no static inspection can
|
| 585 |
+
observe. A consumer therefore cannot learn from `/v1/capabilities` whether a model is actually
|
| 586 |
+
resident.
|
| 587 |
+
|
| 588 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1;
|
| 589 |
+
`docs/STEP8_FINAL_CONFORMANCE_AUDIT.md` §3 (finding C-4).
|
| 590 |
+
|
| 591 |
+
### L-59 — Two subsystems use overlapping words with different meanings (`MEASURED`)
|
| 592 |
+
|
| 593 |
+
To the registry, `unavailable` means "no builder could be constructed"; to the contract it means
|
| 594 |
+
"present but broken, and **therefore a defect**". Transcribing one into the other without noticing
|
| 595 |
+
"would turn a deployment gap into a reported defect, or vice versa". The registry's `available` is not a
|
| 596 |
+
contract state at all, and `degraded` is *both* a whole-service status and a capability state — a reader
|
| 597 |
+
cannot tell from the word alone which is meant.
|
| 598 |
+
|
| 599 |
+
**Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §10 (findings H-2, M-2).
|
| 600 |
+
|
| 601 |
+
### L-60 — The capability source of truth was historically two producers (`RESOLVED`, recorded)
|
| 602 |
+
|
| 603 |
+
Before the owner ruling, `AnalysisController.health()` enumerated **6** capabilities from the registry
|
| 604 |
+
while the served `describe_deployment()` enumerated **2** from two `Path.exists()` calls. The registry
|
| 605 |
+
is now authoritative and `describe_deployment()` delegates to the single adapter — but the historical
|
| 606 |
+
divergence is recorded because "any table that exists in two places will drift, and this one already
|
| 607 |
+
had".
|
| 608 |
+
|
| 609 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1;
|
| 610 |
+
`docs/STEP7_BACKEND_CHAIN_REPORT.md` §10, §16 (finding M-1).
|
| 611 |
+
|
| 612 |
+
### L-61 — `mask_ref` is an undocumented optional field (`OPEN`, low)
|
| 613 |
+
|
| 614 |
+
`core/schemas.py` declares `mask_ref` as a bare `str | None = None` with no description, and
|
| 615 |
+
`grep -r mask_ref docs/` finds nothing. An earlier audit note wrongly listed it as a deliberate
|
| 616 |
+
documented path surface; that was corrected. It remains undocumented.
|
| 617 |
+
|
| 618 |
+
**Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5 (note on `mask_ref`).
|
| 619 |
+
|
| 620 |
+
### L-62 — VRSBench box coordinates are normalised to 0–100 (`MEASURED`)
|
| 621 |
+
|
| 622 |
+
VRSBench's repository notes that provided evaluation box coordinates are normalized to **0–100**, so the
|
| 623 |
+
evaluator adapter must explicitly convert between that convention and the project's internal **0–1**
|
| 624 |
+
representation rather than quietly treating the numbers as pixels. A consumer that reads a grounding box
|
| 625 |
+
without the coordinate-system field will misinterpret it.
|
| 626 |
+
|
| 627 |
+
**Evidence:** `docs/MASTER_ARCHITECTURE_PLAN.md` §13 (grounding head / coordinate convention);
|
| 628 |
+
`docs/STEP7_BACKEND_CHAIN_REPORT.md` §4 (`CoordinateSystem` is exactly `{normalized_0_1, pixel, geo}`).
|
| 629 |
+
|
| 630 |
+
### L-63 — The gateway's rate-limit and size-limit values are implementation choices, not plan facts (`OPEN`)
|
| 631 |
+
|
| 632 |
+
The plan specifies none of them. The defaults in `GatewayConfig` are choices the implementation made,
|
| 633 |
+
and the maintainer is asked to confirm them, because they bound one client's share of the daily budget.
|
| 634 |
+
No rate-limit value is specified by the plan.
|
| 635 |
+
|
| 636 |
+
**Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §4.4; `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.
|
| 637 |
+
|
| 638 |
+
### L-64 — `POST /v1/assets` was not in the original plan (`RESOLVED`, historically undefined)
|
| 639 |
+
|
| 640 |
+
The plan fixes three endpoints, yet `AnalysisRequest.assets` is a list of *handles*, which requires a
|
| 641 |
+
fourth. `docs/API_CONTRACT.md` §2.5 records the gap; the endpoint is now implemented, but the
|
| 642 |
+
fourth-endpoint decision was historically open and the gateway once answered `501` for it.
|
| 643 |
+
|
| 644 |
+
**Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §4.2;
|
| 645 |
+
`docs/STEP7_BACKEND_CHAIN_REPORT.md` §16.
|
| 646 |
+
|
| 647 |
+
---
|
| 648 |
+
|
| 649 |
+
## 8. Dataset and corpus limitations
|
| 650 |
+
|
| 651 |
+
### L-65 — The change model is evaluated only on LEVIR-CD-256 (`MEASURED`)
|
| 652 |
+
|
| 653 |
+
The change head's headline (pooled IoU 0.8122 / macro IoU 0.8457 / pooled F1 0.8964, n = 2048,
|
| 654 |
+
threshold 0.5) is measured on **LEVIR-CD-256** test. Nothing measures transfer to another change
|
| 655 |
+
dataset or to a different sensor pair.
|
| 656 |
+
|
| 657 |
+
**Evidence:** `artifacts/change/eval_test/eval_result.json`; [`EVALUATION.md`](EVALUATION.md).
|
| 658 |
+
|
| 659 |
+
### L-66 — The change evaluation records a large pixel imbalance (`MEASURED`)
|
| 660 |
+
|
| 661 |
+
The change test split is heavily imbalanced: pooled counts record `tp = 5,978,997`, `fp = 523,658`,
|
| 662 |
+
`fn = 858,407`, `tn = 126,856,666` over `n_pixels = 134,217,728`, with a mean change fraction of
|
| 663 |
+
**0.0509** and only **935 of 2,048** images containing change. A pooled IoU on a 5 % positive pixel rate
|
| 664 |
+
is not the same statistic as a balanced one, and the macro/pooled split (macro IoU 0.718 vs pooled IoU
|
| 665 |
+
0.8122) reflects that.
|
| 666 |
+
|
| 667 |
+
**Evidence:** `artifacts/change/eval_test/eval_result.json` (`metrics.pooled`, `metrics.macro`,
|
| 668 |
+
`mean_change_fraction`, `n_images_with_change`).
|
| 669 |
+
|
| 670 |
+
### L-67 — The router corpus is small and corpus-limited (`MEASURED`)
|
| 671 |
+
|
| 672 |
+
The router's training/evaluation corpus is **576** queries in **54** groups (`corpus_total: 576`,
|
| 673 |
+
`corpus_groups: 54`), with a by-task distribution of caption 91, change 115, grounding 128, optical_sar
|
| 674 |
+
50, unsupported 105, vqa 87. The validation split is **n = 86**. A 0.965116 validation accuracy over 86
|
| 675 |
+
examples, drawn from a 576-query corpus, is **indicative only**; it is not a benchmark result.
|
| 676 |
+
|
| 677 |
+
**Evidence:** `artifacts/router/router_adapter_v001/metadata.json` (`corpus`),
|
| 678 |
+
`artifacts/router/threshold_sweep_val.json`.
|
| 679 |
+
|
| 680 |
+
### L-68 — The change-VQA test and test2 splits share scenes (`MEASURED`)
|
| 681 |
+
|
| 682 |
+
The change-VQA dataset's integrity block lists `allowed_shared_pairs: [["Test", "Test2"]]` — the two
|
| 683 |
+
test splits legitimately share scenes, so they are **not independent draws** of the same population. A
|
| 684 |
+
number from one is not a confirmation of the other.
|
| 685 |
+
|
| 686 |
+
**Evidence:** `artifacts/change_vqa/run/run_record.json` (`dataset.integrity`).
|
| 687 |
+
|
| 688 |
+
### L-69 — The optical-SAR metric is on a 19-class space with five absent classes (`MEASURED`)
|
| 689 |
+
|
| 690 |
+
The fusion metric is defined over a **19-class** label space; the scored held-out split contains only 14
|
| 691 |
+
present classes (`classes_present: [0,2,3,4,5,6,7,8,9,10,12,13,17,18]`). The absent five contribute 0.0
|
| 692 |
+
to macro-F1 by construction, which is why accuracy (0.931) and macro-F1 (0.434161) diverge so sharply.
|
| 693 |
+
The metric is a single head's held-out result; it selects no head and compares no arms.
|
| 694 |
+
|
| 695 |
+
**Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`.
|
| 696 |
+
|
| 697 |
+
### L-70 — The grounding evaluation is a single dataset and a single expression family (`MEASURED`)
|
| 698 |
+
|
| 699 |
+
Grounding is measured only on VRSBench referring expressions (n = 16,159). Nothing measures grounding on
|
| 700 |
+
a different expression style or a different imagery distribution. The `matched6` protocol (top_k = 6) is
|
| 701 |
+
the only protocol variation recorded.
|
| 702 |
+
|
| 703 |
+
**Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`,
|
| 704 |
+
`eval_result_matched6.json`.
|
| 705 |
+
|
| 706 |
+
---
|
| 707 |
+
|
| 708 |
+
## 9. Explicit non-claims
|
| 709 |
|
| 710 |
These are things a reader might reasonably assume, which this project does **not** claim:
|
| 711 |
|
|
|
|
| 713 |
- **No claim of production readiness for model quality.** The deployment runs; the models carry the
|
| 714 |
limitations above.
|
| 715 |
- **No claim that the trained heads generalise** beyond their training-family test splits.
|
| 716 |
+
- **No claim that calibration improves confidence** — it made ECE worse (0.013755 → 0.014929).
|
| 717 |
+
- **No claim that the VLM adapter is accepted** for production use — it is `ACCEPTANCE-REJECTED`.
|
| 718 |
- **No claim of an end-to-end accuracy number** — none exists.
|
| 719 |
+
- **No claim that the router is correct on all phrasings** — residuals exist (§2).
|
| 720 |
+
- **No claim that the router number is a test result** — it is validation, ungated, n = 86.
|
| 721 |
+
- **No claim that the optical-SAR or change-VQA rulings are settled** — both are `OPEN`.
|
| 722 |
- **No claim of robustness** to adversarial, corrupted, or out-of-distribution inputs.
|
| 723 |
- **No claim of geolocation accuracy** — grounding boxes are image-relative, not geodetic.
|
| 724 |
- **No claim that the system is a safety-, legal-, or life-critical tool.**
|
| 725 |
+
- **No claim that the code is licensed** — no `LICENSE` file exists.
|
| 726 |
+
- **No claim that the release repositories are public** — three of four are private by design.
|
| 727 |
+
|
| 728 |
+
---
|
| 729 |
|
| 730 |
+
## 10. Where the evidence lives
|
| 731 |
|
| 732 |
| Topic | Evidence |
|
| 733 |
|---|---|
|
| 734 |
| All measured metrics | `artifacts/**/*.json`, verified by `release/tools/verify_readme_metrics.py` |
|
| 735 |
| Metric honesty rules | [`BENCHMARKS.md`](BENCHMARKS.md), [`EVALUATION.md`](EVALUATION.md) |
|
| 736 |
| The grounding resolution rejection | `docs/PHASE7_RESOLUTION_DECISION.md` |
|
| 737 |
+
| The VLM rejection | `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`), `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md` |
|
| 738 |
+
| The optical-SAR ruling | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
|
| 739 |
+
| The change-VQA ruling | `artifacts/change_vqa/run/PROMOTION.json`, `docs/PHASE19_FINAL_HARDENING.md` §4.5 |
|
| 740 |
+
| The calibration result | `artifacts/calibration_v001.json`, `docs/STEP7_BACKEND_CHAIN_REPORT.md` §7 |
|
| 741 |
+
| B-07 / B-02 | [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1, §8.2; `docs/FINAL_DELIVERY_TODO.md` §5 |
|
| 742 |
+
| Router residuals | [`architecture/04-router.md`](architecture/04-router.md), `LIVE_VALIDATION_POSTFIX.md` |
|
| 743 |
+
| Environment traps | [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md), [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §8 |
|
| 744 |
+
| Stale documentation | `release/CURRENT_RELEASE_STATE.md` §6, `docs/PHASE12_CURRENCY_CORRECTION.md` |
|
| 745 |
+
| The stale-claim class | `docs/PHASE12_CURRENCY_CORRECTION.md` §6 |
|