Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/LIMITATIONS.md
Browse files- docs/LIMITATIONS.md +65 -39
docs/LIMITATIONS.md
CHANGED
|
@@ -1,7 +1,8 @@
|
|
| 1 |
# Limitations
|
| 2 |
|
| 3 |
-
An honest catalogue of everything
|
| 4 |
-
results and open items are listed here rather than omitted
|
|
|
|
| 5 |
|
| 6 |
**Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN`.
|
| 7 |
|
|
@@ -11,72 +12,97 @@ results and open items are listed here rather than omitted.
|
|
| 11 |
|
| 12 |
| # | Limitation | Detail |
|
| 13 |
|---|---|---|
|
| 14 |
-
| 1 | **Grounding IoU is low in absolute terms** | mean best IoU 0.2838 (canonical) / 0.2566 (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
|
| 15 |
-
| 2 | **Grounding
|
| 16 |
-
| 3 | **Optical-SAR accuracy is carried by common classes** | accuracy 0.931 but macro-F1 0.434161. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. |
|
| 17 |
-
| 4 | **Change-VQA is weak on rare classes** | accuracy 0.697626 / macro-F1 0.378373 (and 0.651469 / 0.372309
|
| 18 |
| 5 | **VQA is weak-but-related** | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
|
| 19 |
-
| 6 | **Optical-SAR returns a bare class index** | the live service returns `class_18`, not a human-readable label. |
|
| 20 |
-
| 7 | **Calibration made things worse** | ECE 0.013755 → 0.014929. Retained only because it is part of the frozen config. |
|
| 21 |
-
| 8 | **The VLM adapter is not accepted** | metrics usable (exact_match 0.963
|
| 22 |
-
| 9 | **The router number is not a test result** | 0.965116 is validation, **ungated**, n = 86. The test split was **NOT RUN**. |
|
|
|
|
| 23 |
|
| 24 |
## 2. Router residuals (known misroutes)
|
| 25 |
|
| 26 |
| # | Query | Behaviour | Note |
|
| 27 |
|---|---|---|---|
|
| 28 |
-
|
|
| 29 |
-
|
|
| 30 |
-
|
|
| 31 |
|
| 32 |
## 3. Evaluation gaps
|
| 33 |
|
| 34 |
| # | Gap | State |
|
| 35 |
|---|---|---|
|
| 36 |
-
|
|
| 37 |
-
|
|
| 38 |
-
|
|
| 39 |
-
|
|
| 40 |
-
|
|
| 41 |
-
|
|
| 42 |
-
|
|
| 43 |
-
|
|
| 44 |
-
|
|
|
|
|
| 45 |
|
| 46 |
## 4. Operational limitations
|
| 47 |
|
| 48 |
| # | Limitation | State |
|
| 49 |
|---|---|---|
|
| 50 |
-
|
|
| 51 |
-
|
|
| 52 |
-
|
|
| 53 |
-
|
|
| 54 |
-
|
|
| 55 |
-
|
|
| 56 |
-
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
## 5. Packaging and licensing
|
| 59 |
|
| 60 |
| # | Limitation | State |
|
| 61 |
|---|---|---|
|
| 62 |
-
|
|
| 63 |
-
|
|
| 64 |
-
|
|
|
|
|
| 65 |
|
| 66 |
## 6. Documentation caveats
|
| 67 |
|
| 68 |
| # | Caveat | State |
|
| 69 |
|---|---|---|
|
| 70 |
-
|
|
| 71 |
-
|
|
| 72 |
-
|
|
| 73 |
-
|
|
|
|
|
| 74 |
|
| 75 |
## 7. Explicit non-claims
|
| 76 |
|
|
|
|
|
|
|
| 77 |
- **No claim of state-of-the-art performance** on any benchmark.
|
| 78 |
-
- **No claim of production readiness
|
| 79 |
-
|
| 80 |
- **No claim that the trained heads generalise** beyond their training-family test splits.
|
| 81 |
-
- **No claim that calibration improves confidence
|
| 82 |
- **No claim that the VLM adapter is accepted** for production use.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Limitations
|
| 2 |
|
| 3 |
+
An honest, exhaustive catalogue of everything SatQuery AI does **not** do, does poorly, or does not
|
| 4 |
+
know. Negative results and open items are listed here rather than omitted, because a limitation that
|
| 5 |
+
is not written down is a limitation that will be discovered by someone else at the worst moment.
|
| 6 |
|
| 7 |
**Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN`.
|
| 8 |
|
|
|
|
| 12 |
|
| 13 |
| # | Limitation | Detail |
|
| 14 |
|---|---|---|
|
| 15 |
+
| 1 | **Grounding IoU is low in absolute terms** | mean best IoU **0.2838** (canonical) / **0.2566** (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
|
| 16 |
+
| 2 | **Grounding is protocol-sensitive** | two protocols × two decode variants give very different numbers: head_threshold 0.2838, head_argmax **0.1215**, zero-shot 0.0972. An absolute value is meaningless without its protocol. |
|
| 17 |
+
| 3 | **Optical-SAR accuracy is carried by common classes** | accuracy **0.931** but macro-F1 **0.434161**. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. Never quote accuracy alone. |
|
| 18 |
+
| 4 | **Change-VQA is weak on rare classes** | accuracy 0.697626 / macro-F1 0.378373 (test) and 0.651469 / 0.372309 (test2). The wide accuracy–macroF1 gap is the signature of class imbalance. |
|
| 19 |
| 5 | **VQA is weak-but-related** | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
|
| 20 |
+
| 6 | **Optical-SAR returns a bare class index** | the live service returns `class_18`, not a human-readable CLC label. |
|
| 21 |
+
| 7 | **Calibration made things worse** | ECE **0.013755 → 0.014929** (`ece_improvement` −0.001174). Retained only because it is part of the frozen config — **not** because it helped. |
|
| 22 |
+
| 8 | **The VLM adapter is not accepted** | metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is **ACCEPTANCE-REJECTED**; the deployed caption/VQA path uses the **unadapted** model. |
|
| 23 |
+
| 9 | **The router number is not a test result** | 0.965116 is **validation**, **ungated**, **n = 86**, corpus-limited. The test split was **NOT RUN**. |
|
| 24 |
+
| 10 | **The router's routing is not perfect** | known residuals below (§2). |
|
| 25 |
|
| 26 |
## 2. Router residuals (known misroutes)
|
| 27 |
|
| 28 |
| # | Query | Behaviour | Note |
|
| 29 |
|---|---|---|---|
|
| 30 |
+
| 11 | *"What is the new runway?"* | reads `change`, not `vqa` | the word "new" triggers a change reading |
|
| 31 |
+
| 12 | *"How much built-up area was added?"* | reads `vqa` (under-trigger) | a change-style quantifier the router does not catch |
|
| 32 |
+
| 13 | *"What changed between the earlier and later image?"* with **one** asset | console **reads** `change`, **dispatches** `change_vqa` | **intentional** — reading is asset-count-blind, dispatch is asset-count-aware — but visually surprising. Documented in [`architecture/04-router.md`](architecture/04-router.md). |
|
| 33 |
|
| 34 |
## 3. Evaluation gaps
|
| 35 |
|
| 36 |
| # | Gap | State |
|
| 37 |
|---|---|---|
|
| 38 |
+
| 14 | **No system-level end-to-end benchmark** | **NOT RUN — none exists.** No end-to-end accuracy is claimed anywhere. |
|
| 39 |
+
| 15 | **Router test split** | **NOT RUN** |
|
| 40 |
+
| 16 | **Benchmark adapters** | **NOT RUN** |
|
| 41 |
+
| 17 | **End-to-end latency benchmark** | **NOT RUN** (per-specialist latency is recorded only incidentally) |
|
| 42 |
+
| 18 | **Cross-dataset generalisation** | **NOT RUN** — each specialist is evaluated only on its own training-family split |
|
| 43 |
+
| 19 | **Human evaluation** | **NOT RUN** |
|
| 44 |
+
| 20 | **Robustness / adversarial evaluation** | **NOT RUN** |
|
| 45 |
+
| 21 | **BigEarthNet label semantics** | the local subset is **100 % single-label** vs the official 1–11 multi-label scheme, so its metrics are **not comparable** to published numbers |
|
| 46 |
+
| 22 | **Reliability diagram** | the shipped diagram is the **pre-scaling** curve (labelled as such); the calibrated curve is not plotted |
|
| 47 |
+
| 23 | **Statistical significance for most metrics** | only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval. Other per-task numbers are point estimates. |
|
| 48 |
|
| 49 |
## 4. Operational limitations
|
| 50 |
|
| 51 |
| # | Limitation | State |
|
| 52 |
|---|---|---|
|
| 53 |
+
| 24 | **B-07: transient tunnel gaps** | a request can hang or return 504. Patch prepared, **NOT deployed**. **OPEN** |
|
| 54 |
+
| 25 | **B-07 root shape** | in `auto` mode a tunnel timeout falls through to the forward path, burning `wake_timeout_s` (120 s) on a `302`; worst case ≈ **249 s** (150 + 120). Measured. |
|
| 55 |
+
| 26 | **Cold start is tens of seconds** | Render free tier sleeps; the Codespace may be stopped. Documented, not hidden. |
|
| 56 |
+
| 27 | **B-02: `codespace_name` trailing newline** | cosmetic; the wake path strips it. **OPEN (cosmetic)** |
|
| 57 |
+
| 28 | **No database, auth, or queue** | stateless gateway **by design** — no persistence of runs or users. |
|
| 58 |
+
| 29 | **Deployment repos are private** | their links 404 for an outside audience. **BY DESIGN** |
|
| 59 |
+
| 30 | **`deploy/` in the monorepo is stale/untracked** | not the deployed source. Trap. |
|
| 60 |
+
| 31 | **No APM, distributed tracing, or cost accounting** | observability is limited to the health payload and per-run traces. |
|
| 61 |
+
| 32 | **Single-region, no HA** | one Render service, one Codespace. |
|
| 62 |
|
| 63 |
## 5. Packaging and licensing
|
| 64 |
|
| 65 |
| # | Limitation | State |
|
| 66 |
|---|---|---|
|
| 67 |
+
| 33 | **No `LICENSE` file** | none exists in the source repository. **OPEN** — a licence must be selected before public release of the code. |
|
| 68 |
+
| 34 | **Backbones are not redistributed** | fetched from the Hub at run time; their licences are their own. |
|
| 69 |
+
| 35 | **Artifact bloat** | `artifacts/` is ~3.7 GB, mostly reproducible caches, a duplicate 774 MB ZIP, a duplicate 241 MB probe, two 231 MB feature caches, a superseded 276 MB head, and ~235 MB of selection manifests — not released weights. |
|
| 70 |
+
| 36 | **The released weights require their backbones** | the six artifacts are small modules; a consumer must also fetch the pinned backbones. |
|
| 71 |
|
| 72 |
## 6. Documentation caveats
|
| 73 |
|
| 74 |
| # | Caveat | State |
|
| 75 |
|---|---|---|
|
| 76 |
+
| 37 | **`docs/FINAL_DELIVERY_REPORT.md` §6 is stale** | it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25. |
|
| 77 |
+
| 38 | **The monorepo `README.md` was materially stale** | it described a hermetic frontend, an in-progress Render/Codespace, and a `/v1/*` contract. Superseded by this release's README. |
|
| 78 |
+
| 39 | **`hf/` docs were stale** | they asserted the project owns no weights and has no HF credentials. Both were false at release time. |
|
| 79 |
+
| 40 | **Anatomy plate image variant** | points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic. |
|
| 80 |
+
| 41 | **The original master plan describes a superseded deployment** | it specifies a Gradio GUI + HF Space + ZeroGPU + Railway. The shipped system is a static frontend + Render + Codespace tunnel, serving JSON. |
|
| 81 |
|
| 82 |
## 7. Explicit non-claims
|
| 83 |
|
| 84 |
+
These are things a reader might reasonably assume, which this project does **not** claim:
|
| 85 |
+
|
| 86 |
- **No claim of state-of-the-art performance** on any benchmark.
|
| 87 |
+
- **No claim of production readiness for model quality.** The deployment runs; the models carry the
|
| 88 |
+
limitations above.
|
| 89 |
- **No claim that the trained heads generalise** beyond their training-family test splits.
|
| 90 |
+
- **No claim that calibration improves confidence** — it made ECE worse.
|
| 91 |
- **No claim that the VLM adapter is accepted** for production use.
|
| 92 |
+
- **No claim of an end-to-end accuracy number** — none exists.
|
| 93 |
+
- **No claim that the router is correct on all phrasings** — residuals exist.
|
| 94 |
+
- **No claim of robustness** to adversarial, corrupted, or out-of-distribution inputs.
|
| 95 |
+
- **No claim of geolocation accuracy** — grounding boxes are image-relative, not geodetic.
|
| 96 |
+
- **No claim that the system is a safety-, legal-, or life-critical tool.**
|
| 97 |
+
|
| 98 |
+
## 8. Where the evidence lives
|
| 99 |
+
|
| 100 |
+
| Topic | Evidence |
|
| 101 |
+
|---|---|
|
| 102 |
+
| All measured metrics | `artifacts/**/*.json`, verified by `release/tools/verify_readme_metrics.py` |
|
| 103 |
+
| Metric honesty rules | [`BENCHMARKS.md`](BENCHMARKS.md), [`EVALUATION.md`](EVALUATION.md) |
|
| 104 |
+
| The grounding resolution rejection | `docs/PHASE7_RESOLUTION_DECISION.md` |
|
| 105 |
+
| The VLM rejection | `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`) |
|
| 106 |
+
| B-07 / B-02 | [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1, §5 |
|
| 107 |
+
| Router residuals | [`architecture/04-router.md`](architecture/04-router.md) |
|
| 108 |
+
| Environment traps | [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) |
|