SatQuery / docs /LIMITATIONS.md
thundercode's picture
release: add docs/LIMITATIONS.md
18f4386 verified
|
Raw History Blame
8.36 kB

Limitations

An honest, exhaustive catalogue of everything SatQuery AI does not do, does poorly, or does not know. Negative results and open items are listed here rather than omitted, because a limitation that is not written down is a limitation that will be discovered by someone else at the worst moment.

Status tags: OPEN · NOT RUN · REJECTED · DEFERRED · BY DESIGN.


1. Model quality

# Limitation Detail
1 Grounding IoU is low in absolute terms mean best IoU 0.2838 (canonical) / 0.2566 (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved".
2 Grounding is protocol-sensitive two protocols × two decode variants give very different numbers: head_threshold 0.2838, head_argmax 0.1215, zero-shot 0.0972. An absolute value is meaningless without its protocol.
3 Optical-SAR accuracy is carried by common classes accuracy 0.931 but macro-F1 0.434161. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. Never quote accuracy alone.
4 Change-VQA is weak on rare classes accuracy 0.697626 / macro-F1 0.378373 (test) and 0.651469 / 0.372309 (test2). The wide accuracy–macroF1 gap is the signature of class imbalance.
5 VQA is weak-but-related the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class.
6 Optical-SAR returns a bare class index the live service returns class_18, not a human-readable CLC label.
7 Calibration made things worse ECE 0.013755 → 0.014929 (ece_improvement −0.001174). Retained only because it is part of the frozen config — not because it helped.
8 The VLM adapter is not accepted metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is ACCEPTANCE-REJECTED; the deployed caption/VQA path uses the unadapted model.
9 The router number is not a test result 0.965116 is validation, ungated, n = 86, corpus-limited. The test split was NOT RUN.
10 The router's routing is not perfect known residuals below (§2).

2. Router residuals (known misroutes)

# Query Behaviour Note
11 "What is the new runway?" reads change, not vqa the word "new" triggers a change reading
12 "How much built-up area was added?" reads vqa (under-trigger) a change-style quantifier the router does not catch
13 "What changed between the earlier and later image?" with one asset console reads change, dispatches change_vqa intentional — reading is asset-count-blind, dispatch is asset-count-aware — but visually surprising. Documented in architecture/04-router.md.

3. Evaluation gaps

# Gap State
14 No system-level end-to-end benchmark NOT RUN — none exists. No end-to-end accuracy is claimed anywhere.
15 Router test split NOT RUN
16 Benchmark adapters NOT RUN
17 End-to-end latency benchmark NOT RUN (per-specialist latency is recorded only incidentally)
18 Cross-dataset generalisation NOT RUN — each specialist is evaluated only on its own training-family split
19 Human evaluation NOT RUN
20 Robustness / adversarial evaluation NOT RUN
21 BigEarthNet label semantics the local subset is 100 % single-label vs the official 1–11 multi-label scheme, so its metrics are not comparable to published numbers
22 Reliability diagram the shipped diagram is the pre-scaling curve (labelled as such); the calibrated curve is not plotted
23 Statistical significance for most metrics only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval. Other per-task numbers are point estimates.

4. Operational limitations

# Limitation State
24 B-07: transient tunnel gaps a request can hang or return 504. Patch prepared, NOT deployed. OPEN
25 B-07 root shape in auto mode a tunnel timeout falls through to the forward path, burning wake_timeout_s (120 s) on a 302; worst case ≈ 249 s (150 + 120). Measured.
26 Cold start is tens of seconds Render free tier sleeps; the Codespace may be stopped. Documented, not hidden.
27 B-02: codespace_name trailing newline cosmetic; the wake path strips it. OPEN (cosmetic)
28 No database, auth, or queue stateless gateway by design — no persistence of runs or users.
29 Deployment repos are private their links 404 for an outside audience. BY DESIGN
30 deploy/ in the monorepo is stale/untracked not the deployed source. Trap.
31 No APM, distributed tracing, or cost accounting observability is limited to the health payload and per-run traces.
32 Single-region, no HA one Render service, one Codespace.

5. Packaging and licensing

# Limitation State
33 No LICENSE file none exists in the source repository. OPEN — a licence must be selected before public release of the code.
34 Backbones are not redistributed fetched from the Hub at run time; their licences are their own.
35 Artifact bloat artifacts/ is ~3.7 GB, mostly reproducible caches, a duplicate 774 MB ZIP, a duplicate 241 MB probe, two 231 MB feature caches, a superseded 276 MB head, and ~235 MB of selection manifests — not released weights.
36 The released weights require their backbones the six artifacts are small modules; a consumer must also fetch the pinned backbones.

6. Documentation caveats

# Caveat State
37 docs/FINAL_DELIVERY_REPORT.md §6 is stale it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25.
38 The monorepo README.md was materially stale it described a hermetic frontend, an in-progress Render/Codespace, and a /v1/* contract. Superseded by this release's README.
39 hf/ docs were stale they asserted the project owns no weights and has no HF credentials. Both were false at release time.
40 Anatomy plate image variant points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic.
41 The original master plan describes a superseded deployment it specifies a Gradio GUI + HF Space + ZeroGPU + Railway. The shipped system is a static frontend + Render + Codespace tunnel, serving JSON.

7. Explicit non-claims

These are things a reader might reasonably assume, which this project does not claim:

  • No claim of state-of-the-art performance on any benchmark.
  • No claim of production readiness for model quality. The deployment runs; the models carry the limitations above.
  • No claim that the trained heads generalise beyond their training-family test splits.
  • No claim that calibration improves confidence — it made ECE worse.
  • No claim that the VLM adapter is accepted for production use.
  • No claim of an end-to-end accuracy number — none exists.
  • No claim that the router is correct on all phrasings — residuals exist.
  • No claim of robustness to adversarial, corrupted, or out-of-distribution inputs.
  • No claim of geolocation accuracy — grounding boxes are image-relative, not geodetic.
  • No claim that the system is a safety-, legal-, or life-critical tool.

8. Where the evidence lives

Topic Evidence
All measured metrics artifacts/**/*.json, verified by release/tools/verify_readme_metrics.py
Metric honesty rules BENCHMARKS.md, EVALUATION.md
The grounding resolution rejection docs/PHASE7_RESOLUTION_DECISION.md
The VLM rejection artifacts/vlm/phase6_closure.json (why_acceptance_rejected)
B-07 / B-02 DEPLOYMENT.md §8.1, §5
Router residuals architecture/04-router.md
Environment traps REPRODUCIBILITY.md