File size: 7,191 Bytes
c2283ee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
# Evaluation

This document describes **how** each number in [`BENCHMARKS.md`](BENCHMARKS.md) was produced, and
enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was
not measured is tagged `NOT RUN`, and a metric that got worse is shown getting worse.

**Status tags:** `VERIFIED` · `MEASURED` · `NOT RUN` · `OPEN` · `REJECTED`.

---

## 1. Evaluation-honesty rules (enforced by convention and by tooling)

1. **Evidence before claims.** Every reported number has an artifact path. A number with no artifact
   does not appear in the docs.
2. **Two protocols are never collapsed.** Grounding is reported under canonical **and** matched6.
3. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
4. **accuracy never travels without macro-F1** for imbalanced multi-class heads.
5. **Validation is not test.** The router figure is labelled validation, ungated, n = 86.
6. **A negative result stays negative.** Calibration ECE worsened; it is shown worsening.
7. **USABLE ≠ ACCEPTED.** The VLM adapter's metrics are real; its status is acceptance-rejected.
8. **No composite/vanity score.** There is no single headline accuracy for the system, and none is
   invented by averaging the per-task numbers.

## 2. Per-task evaluation protocols

### 2.1 Change detection — `VERIFIED`

- **Split:** LEVIR-CD-256 test, n = 2,048 (immutable public split).
- **Threshold:** 0.50, frozen in config.
- **Metrics:** pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can
  be recomputed.
- **Why both pooled and macro:** the test split is only ≈ 5 % changed pixels
  (mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different
  questions about that imbalance.
- **Artifact:** `artifacts/change/eval_test/eval_result.json`.

### 2.2 Grounding — `MEASURED`, two protocols × two decode variants

- **Split:** VRSBench, n = 16,159.
- **Box convention:** VRSBench 0–100 → project 0–1, via declared `benchmark_box_scale: 100.0`.
- **Protocols:** *canonical* and *matched6* — both reported.
- **Decode variants:** *head threshold* (the shipped decode), *head argmax*, and *zero-shot
  matched* (baseline).
- **Resolution:** frozen at 224; 448 rejected by a pre-registered paired test.

| Protocol | decode | mean best IoU | recall@0.5 |
|---|---|---|---|
| canonical | head threshold | 0.2838 | 0.2198 |
| canonical | head argmax | 0.1215 | — |
| canonical | zero-shot matched | 0.0972 | — |
| matched6 | head threshold | 0.2566 | 0.1938 |

**Artifacts:** `…/eval_result_canonical.json`, `…/eval_result_matched6.json`.

### 2.3 Optical-SAR fusion — `MEASURED`, ruling `OPEN`

- **Split:** held-out test, n = 4,000.
- **Label space:** 19 CLC classes; 14 present, 5 absent in the scored split.
- **macro-F1 denominator:** all 19 classes (absent classes contribute 0.0) — recorded explicitly.
- **Pre-registration:** the metric is a pre-registered protocol (`pre_registered_11.5`);
  `is_deciding_statistic: False`.
- **Artifact:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`.

### 2.4 Change-VQA — `MEASURED`, ruling `OPEN`

- **Splits:** `test` (n = 39,686) and `test2`.
- **Selection:** epoch 8, chosen on val answer accuracy 0.700018.
- **Artifact:** `artifacts/change_vqa/run/PROMOTION.json`.

### 2.5 VLM adapter — `MEASURED`, `ACCEPTANCE-REJECTED`

- **Split:** frozen 1,000-question subset.
- **Metrics:** exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline).
- **Decision:** the artifact is **not accepted** for promotion; the record of *why* is stored in
  `artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`.
- **Artifact:** `artifacts/vlm/phase6_closure.json` (`status: CLOSED`).

### 2.6 Router — `MEASURED`, `TEST NOT RUN`

- **Split scored:** validation, n = 86, corpus-limited.
- **Metric:** overall **ungated** accuracy 0.965116.
- **Not done:** the test split was never scored. The val number is indicative only.

### 2.7 Calibration — `MEASURED`, not an improvement

- **Fit split:** val, n = 16,441. Temperature **T = 0.9772732**.
- **Result:** ECE 0.013755 → **0.014929** (`ece_improvement = −0.001174`).
- **Retention rationale:** part of the frozen config, **not** because it helped.
- **Artifact:** `artifacts/calibration_v001.json`.

## 3. Behavioural evaluation (live validation)

Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser,
one upload per case, with per-case screenshots and recorded run ids:

| Property | Result |
|---|---|
| Independent full passes | **3** |
| Cases per pass | 8 (6 regression + 2 router-defect) |
| Passes at 8/8 | **3 of 3** |
| Live runs | **24** |
| Correct dispatches | **24** |
| Mock-node contamination | **0** |
| Trace fill | **94.4444 %** |
| Frontend regression suite | **106 passed** |

### 3.1 Harness integrity — a bug that was caught

An earlier harness revision typed queries with **synthetic CDP key events**, which Chrome **silently
drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query
and still recorded a "result" — a false pass.

The current harness **asserts form state before dispatch**: that the query box really holds the
intended query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. The
earlier 8/8 run was independently checked and confirmed **not** to have been infected (its answers
were query-specific and the query text was embedded in the answers). This failure mode is recorded
here because it is exactly the kind of silent false-positive an evaluation harness must not have.

## 4. Test suite

| Suite | Result | Notes |
|---|---|---|
| `tests/unit/test_frontend_live_wiring.py` | **106 passed** | re-run this session |
| Doc/frontend suite (5 files) | **183 passed** | |
| Full `tests/unit` | 5–6 failures, **all environmental/ordering** | 4× `test_safe_delete_shim` (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) |
| Affected files re-run together | **137 passed** | confirms the failures are not real regressions |

**The environmental failures are not hidden.** They are attributable to the sandbox delete guard and
test ordering, not to the code under test.

## 5. What is NOT evaluated

| Item | State |
|---|---|
| System-level end-to-end accuracy | **NOT RUN — none exists** |
| Router test split | **NOT RUN** |
| Benchmark adapters | **NOT RUN** |
| End-to-end latency benchmark | **NOT RUN** |
| Cross-dataset generalisation | **NOT RUN** |
| Human evaluation | **NOT RUN** |
| Adversarial / robustness evaluation | **NOT RUN** |

## 6. Reproducing the metric check

```bash
python release/tools/verify_readme_metrics.py
```

This walks every claim in [`../README.md`](../README.md) and [`BENCHMARKS.md`](BENCHMARKS.md) to its
source artifact and compares values at the printed precision. It prints `ALL CLAIMS VERIFIED` (exit 0)
only when **all 20** numeric claims match and the status assertions hold. Committed output:
`release/tools/readme_metrics_report.txt`.