SatQuery / docs /LIMITATIONS.md
thundercode's picture
release: add docs/LIMITATIONS.md
d452c40 verified
|
Raw History Blame
5.75 kB

Limitations

An honest catalogue of everything this system does not do, does poorly, or does not know. Negative results and open items are listed here rather than omitted.

Status tags: OPEN · NOT RUN · REJECTED · DEFERRED · BY DESIGN.


1. Model quality

# Limitation Detail
1 Grounding IoU is low in absolute terms mean best IoU 0.2838 (canonical) / 0.2566 (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved".
2 Grounding depends on the box convention Two protocols and two decode variants produce very different numbers (0.2838 vs 0.1215 argmax). Absolute values are protocol-sensitive.
3 Optical-SAR accuracy is carried by common classes accuracy 0.931 but macro-F1 0.434161. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction.
4 Change-VQA is weak on rare classes accuracy 0.697626 / macro-F1 0.378373 (and 0.651469 / 0.372309 on the second test set).
5 VQA is weak-but-related the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class.
6 Optical-SAR returns a bare class index the live service returns class_18, not a human-readable label.
7 Calibration made things worse ECE 0.013755 → 0.014929. Retained only because it is part of the frozen config.
8 The VLM adapter is not accepted metrics usable (exact_match 0.963), but the artifact is acceptance-rejected; the deployed path uses the unadapted model.
9 The router number is not a test result 0.965116 is validation, ungated, n = 86. The test split was NOT RUN.

2. Router residuals (known misroutes)

# Query Behaviour Note
10 "What is the new runway?" reads change, not vqa residual ambiguity — "new" triggers a change reading
11 "How much built-up area was added?" reads vqa (under-trigger) a change-style quantifier not caught by the router
12 "What changed between the earlier and later image?" with one asset console reads change, dispatches change_vqa intentional (asset-count-aware dispatch) but visually surprising — documented in ARCHITECTURE.md §4

3. Evaluation gaps

# Gap State
13 No system-level end-to-end benchmark NOT RUN — none exists. No end-to-end accuracy is claimed.
14 Router test split NOT RUN
15 Benchmark adapters NOT RUN
16 End-to-end latency benchmark NOT RUN (per-specialist latency only, incidentally)
17 Cross-dataset generalisation NOT RUN — each specialist is evaluated on its own training-family split only
18 Human evaluation NOT RUN
19 Robustness / adversarial evaluation NOT RUN
20 BigEarthNet label semantics the local subset is 100 % single-label vs the official 1–11 multi-label scheme, so metrics are not comparable to published numbers
21 Reliability diagram the pre-scaling diagram is labelled as such; the calibrated curve is not plotted

4. Operational limitations

# Limitation State
22 B-07: transient tunnel gaps a request can hang or return 504. Patch prepared, NOT deployed. OPEN
23 B-07 root shape in auto mode a tunnel timeout falls through to the forward path, burning wake_timeout_s (120 s) on a 302; worst case ≈ 249 s. Measured.
24 Cold start is tens of seconds Render free tier sleeps; the Codespace may be stopped. Documented, not hidden.
25 B-02: codespace_name trailing newline cosmetic; the wake path strips it. OPEN (cosmetic)
26 No database, auth, or queue stateless gateway by design — no persistence of runs or users.
27 Deployment repos are private their links 404 for an outside audience. BY DESIGN
28 deploy/ in the monorepo is stale/untracked not the deployed source. Trap.

5. Packaging and licensing

# Limitation State
29 No LICENSE file none exists in the source repository. OPEN — a licence must be selected before public release of the code.
30 Backbones are not redistributed fetched from the Hub at run time; their licences are their own.
31 Artifact bloat artifacts/ is ~3.7 GB, mostly reproducible caches, duplicate ZIPs and superseded checkpoints rather than released weights.

6. Documentation caveats

# Caveat State
32 docs/FINAL_DELIVERY_REPORT.md §6 is stale it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25.
33 The monorepo README.md was materially stale it described a hermetic frontend, an in-progress Render/Codespace, and a /v1/* contract. Superseded by this release's README.
34 hf/ docs were stale they asserted the project owns no weights and has no HF credentials. Both were false at release time.
35 Anatomy plate image variant points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic.

7. Explicit non-claims

  • No claim of state-of-the-art performance on any benchmark.
  • No claim of production readiness for the model quality — the deployment runs, but the models carry the limitations above.
  • No claim that the trained heads generalise beyond their training-family test splits.
  • No claim that calibration improves confidence.
  • No claim that the VLM adapter is accepted for production use.