Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 7,191 Bytes
c2283ee | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | # Evaluation
This document describes **how** each number in [`BENCHMARKS.md`](BENCHMARKS.md) was produced, and
enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was
not measured is tagged `NOT RUN`, and a metric that got worse is shown getting worse.
**Status tags:** `VERIFIED` · `MEASURED` · `NOT RUN` · `OPEN` · `REJECTED`.
---
## 1. Evaluation-honesty rules (enforced by convention and by tooling)
1. **Evidence before claims.** Every reported number has an artifact path. A number with no artifact
does not appear in the docs.
2. **Two protocols are never collapsed.** Grounding is reported under canonical **and** matched6.
3. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
4. **accuracy never travels without macro-F1** for imbalanced multi-class heads.
5. **Validation is not test.** The router figure is labelled validation, ungated, n = 86.
6. **A negative result stays negative.** Calibration ECE worsened; it is shown worsening.
7. **USABLE ≠ ACCEPTED.** The VLM adapter's metrics are real; its status is acceptance-rejected.
8. **No composite/vanity score.** There is no single headline accuracy for the system, and none is
invented by averaging the per-task numbers.
## 2. Per-task evaluation protocols
### 2.1 Change detection — `VERIFIED`
- **Split:** LEVIR-CD-256 test, n = 2,048 (immutable public split).
- **Threshold:** 0.50, frozen in config.
- **Metrics:** pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can
be recomputed.
- **Why both pooled and macro:** the test split is only ≈ 5 % changed pixels
(mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different
questions about that imbalance.
- **Artifact:** `artifacts/change/eval_test/eval_result.json`.
### 2.2 Grounding — `MEASURED`, two protocols × two decode variants
- **Split:** VRSBench, n = 16,159.
- **Box convention:** VRSBench 0–100 → project 0–1, via declared `benchmark_box_scale: 100.0`.
- **Protocols:** *canonical* and *matched6* — both reported.
- **Decode variants:** *head threshold* (the shipped decode), *head argmax*, and *zero-shot
matched* (baseline).
- **Resolution:** frozen at 224; 448 rejected by a pre-registered paired test.
| Protocol | decode | mean best IoU | recall@0.5 |
|---|---|---|---|
| canonical | head threshold | 0.2838 | 0.2198 |
| canonical | head argmax | 0.1215 | — |
| canonical | zero-shot matched | 0.0972 | — |
| matched6 | head threshold | 0.2566 | 0.1938 |
**Artifacts:** `…/eval_result_canonical.json`, `…/eval_result_matched6.json`.
### 2.3 Optical-SAR fusion — `MEASURED`, ruling `OPEN`
- **Split:** held-out test, n = 4,000.
- **Label space:** 19 CLC classes; 14 present, 5 absent in the scored split.
- **macro-F1 denominator:** all 19 classes (absent classes contribute 0.0) — recorded explicitly.
- **Pre-registration:** the metric is a pre-registered protocol (`pre_registered_11.5`);
`is_deciding_statistic: False`.
- **Artifact:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`.
### 2.4 Change-VQA — `MEASURED`, ruling `OPEN`
- **Splits:** `test` (n = 39,686) and `test2`.
- **Selection:** epoch 8, chosen on val answer accuracy 0.700018.
- **Artifact:** `artifacts/change_vqa/run/PROMOTION.json`.
### 2.5 VLM adapter — `MEASURED`, `ACCEPTANCE-REJECTED`
- **Split:** frozen 1,000-question subset.
- **Metrics:** exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline).
- **Decision:** the artifact is **not accepted** for promotion; the record of *why* is stored in
`artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`.
- **Artifact:** `artifacts/vlm/phase6_closure.json` (`status: CLOSED`).
### 2.6 Router — `MEASURED`, `TEST NOT RUN`
- **Split scored:** validation, n = 86, corpus-limited.
- **Metric:** overall **ungated** accuracy 0.965116.
- **Not done:** the test split was never scored. The val number is indicative only.
### 2.7 Calibration — `MEASURED`, not an improvement
- **Fit split:** val, n = 16,441. Temperature **T = 0.9772732**.
- **Result:** ECE 0.013755 → **0.014929** (`ece_improvement = −0.001174`).
- **Retention rationale:** part of the frozen config, **not** because it helped.
- **Artifact:** `artifacts/calibration_v001.json`.
## 3. Behavioural evaluation (live validation)
Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser,
one upload per case, with per-case screenshots and recorded run ids:
| Property | Result |
|---|---|
| Independent full passes | **3** |
| Cases per pass | 8 (6 regression + 2 router-defect) |
| Passes at 8/8 | **3 of 3** |
| Live runs | **24** |
| Correct dispatches | **24** |
| Mock-node contamination | **0** |
| Trace fill | **94.4444 %** |
| Frontend regression suite | **106 passed** |
### 3.1 Harness integrity — a bug that was caught
An earlier harness revision typed queries with **synthetic CDP key events**, which Chrome **silently
drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query
and still recorded a "result" — a false pass.
The current harness **asserts form state before dispatch**: that the query box really holds the
intended query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. The
earlier 8/8 run was independently checked and confirmed **not** to have been infected (its answers
were query-specific and the query text was embedded in the answers). This failure mode is recorded
here because it is exactly the kind of silent false-positive an evaluation harness must not have.
## 4. Test suite
| Suite | Result | Notes |
|---|---|---|
| `tests/unit/test_frontend_live_wiring.py` | **106 passed** | re-run this session |
| Doc/frontend suite (5 files) | **183 passed** | |
| Full `tests/unit` | 5–6 failures, **all environmental/ordering** | 4× `test_safe_delete_shim` (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) |
| Affected files re-run together | **137 passed** | confirms the failures are not real regressions |
**The environmental failures are not hidden.** They are attributable to the sandbox delete guard and
test ordering, not to the code under test.
## 5. What is NOT evaluated
| Item | State |
|---|---|
| System-level end-to-end accuracy | **NOT RUN — none exists** |
| Router test split | **NOT RUN** |
| Benchmark adapters | **NOT RUN** |
| End-to-end latency benchmark | **NOT RUN** |
| Cross-dataset generalisation | **NOT RUN** |
| Human evaluation | **NOT RUN** |
| Adversarial / robustness evaluation | **NOT RUN** |
## 6. Reproducing the metric check
```bash
python release/tools/verify_readme_metrics.py
```
This walks every claim in [`../README.md`](../README.md) and [`BENCHMARKS.md`](BENCHMARKS.md) to its
source artifact and compares values at the printed precision. It prints `ALL CLAIMS VERIFIED` (exit 0)
only when **all 20** numeric claims match and the status assertions hold. Committed output:
`release/tools/readme_metrics_report.txt`.
|