Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/LIMITATIONS.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 8.36 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/394471d520abaa95b85b18eddbdbd65d55f7d996/docs/LIMITATIONS.md
- Command line
-
hf download hf://thundercode/SatQuery@394471d520abaa95b85b18eddbdbd65d55f7d996/docs/LIMITATIONS.md
-
curl -L -o LIMITATIONS.md https://huggingface.co/thundercode/SatQuery/resolve/394471d520abaa95b85b18eddbdbd65d55f7d996/docs/LIMITATIONS.md
8.36 kB
Limitations
An honest, exhaustive catalogue of everything SatQuery AI does not do, does poorly, or does not know. Negative results and open items are listed here rather than omitted, because a limitation that is not written down is a limitation that will be discovered by someone else at the worst moment.
Status tags: OPEN · NOT RUN · REJECTED · DEFERRED · BY DESIGN.
1. Model quality
| # | Limitation | Detail |
|---|---|---|
| 1 | Grounding IoU is low in absolute terms | mean best IoU 0.2838 (canonical) / 0.2566 (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
| 2 | Grounding is protocol-sensitive | two protocols × two decode variants give very different numbers: head_threshold 0.2838, head_argmax 0.1215, zero-shot 0.0972. An absolute value is meaningless without its protocol. |
| 3 | Optical-SAR accuracy is carried by common classes | accuracy 0.931 but macro-F1 0.434161. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. Never quote accuracy alone. |
| 4 | Change-VQA is weak on rare classes | accuracy 0.697626 / macro-F1 0.378373 (test) and 0.651469 / 0.372309 (test2). The wide accuracy–macroF1 gap is the signature of class imbalance. |
| 5 | VQA is weak-but-related | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
| 6 | Optical-SAR returns a bare class index | the live service returns class_18, not a human-readable CLC label. |
| 7 | Calibration made things worse | ECE 0.013755 → 0.014929 (ece_improvement −0.001174). Retained only because it is part of the frozen config — not because it helped. |
| 8 | The VLM adapter is not accepted | metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is ACCEPTANCE-REJECTED; the deployed caption/VQA path uses the unadapted model. |
| 9 | The router number is not a test result | 0.965116 is validation, ungated, n = 86, corpus-limited. The test split was NOT RUN. |
| 10 | The router's routing is not perfect | known residuals below (§2). |
2. Router residuals (known misroutes)
| # | Query | Behaviour | Note |
|---|---|---|---|
| 11 | "What is the new runway?" | reads change, not vqa |
the word "new" triggers a change reading |
| 12 | "How much built-up area was added?" | reads vqa (under-trigger) |
a change-style quantifier the router does not catch |
| 13 | "What changed between the earlier and later image?" with one asset | console reads change, dispatches change_vqa |
intentional — reading is asset-count-blind, dispatch is asset-count-aware — but visually surprising. Documented in architecture/04-router.md. |
3. Evaluation gaps
| # | Gap | State |
|---|---|---|
| 14 | No system-level end-to-end benchmark | NOT RUN — none exists. No end-to-end accuracy is claimed anywhere. |
| 15 | Router test split | NOT RUN |
| 16 | Benchmark adapters | NOT RUN |
| 17 | End-to-end latency benchmark | NOT RUN (per-specialist latency is recorded only incidentally) |
| 18 | Cross-dataset generalisation | NOT RUN — each specialist is evaluated only on its own training-family split |
| 19 | Human evaluation | NOT RUN |
| 20 | Robustness / adversarial evaluation | NOT RUN |
| 21 | BigEarthNet label semantics | the local subset is 100 % single-label vs the official 1–11 multi-label scheme, so its metrics are not comparable to published numbers |
| 22 | Reliability diagram | the shipped diagram is the pre-scaling curve (labelled as such); the calibrated curve is not plotted |
| 23 | Statistical significance for most metrics | only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval. Other per-task numbers are point estimates. |
4. Operational limitations
| # | Limitation | State |
|---|---|---|
| 24 | B-07: transient tunnel gaps | a request can hang or return 504. Patch prepared, NOT deployed. OPEN |
| 25 | B-07 root shape | in auto mode a tunnel timeout falls through to the forward path, burning wake_timeout_s (120 s) on a 302; worst case ≈ 249 s (150 + 120). Measured. |
| 26 | Cold start is tens of seconds | Render free tier sleeps; the Codespace may be stopped. Documented, not hidden. |
| 27 | B-02: codespace_name trailing newline |
cosmetic; the wake path strips it. OPEN (cosmetic) |
| 28 | No database, auth, or queue | stateless gateway by design — no persistence of runs or users. |
| 29 | Deployment repos are private | their links 404 for an outside audience. BY DESIGN |
| 30 | deploy/ in the monorepo is stale/untracked |
not the deployed source. Trap. |
| 31 | No APM, distributed tracing, or cost accounting | observability is limited to the health payload and per-run traces. |
| 32 | Single-region, no HA | one Render service, one Codespace. |
5. Packaging and licensing
| # | Limitation | State |
|---|---|---|
| 33 | No LICENSE file |
none exists in the source repository. OPEN — a licence must be selected before public release of the code. |
| 34 | Backbones are not redistributed | fetched from the Hub at run time; their licences are their own. |
| 35 | Artifact bloat | artifacts/ is ~3.7 GB, mostly reproducible caches, a duplicate 774 MB ZIP, a duplicate 241 MB probe, two 231 MB feature caches, a superseded 276 MB head, and ~235 MB of selection manifests — not released weights. |
| 36 | The released weights require their backbones | the six artifacts are small modules; a consumer must also fetch the pinned backbones. |
6. Documentation caveats
| # | Caveat | State |
|---|---|---|
| 37 | docs/FINAL_DELIVERY_REPORT.md §6 is stale |
it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25. |
| 38 | The monorepo README.md was materially stale |
it described a hermetic frontend, an in-progress Render/Codespace, and a /v1/* contract. Superseded by this release's README. |
| 39 | hf/ docs were stale |
they asserted the project owns no weights and has no HF credentials. Both were false at release time. |
| 40 | Anatomy plate image variant | points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic. |
| 41 | The original master plan describes a superseded deployment | it specifies a Gradio GUI + HF Space + ZeroGPU + Railway. The shipped system is a static frontend + Render + Codespace tunnel, serving JSON. |
7. Explicit non-claims
These are things a reader might reasonably assume, which this project does not claim:
- No claim of state-of-the-art performance on any benchmark.
- No claim of production readiness for model quality. The deployment runs; the models carry the limitations above.
- No claim that the trained heads generalise beyond their training-family test splits.
- No claim that calibration improves confidence — it made ECE worse.
- No claim that the VLM adapter is accepted for production use.
- No claim of an end-to-end accuracy number — none exists.
- No claim that the router is correct on all phrasings — residuals exist.
- No claim of robustness to adversarial, corrupted, or out-of-distribution inputs.
- No claim of geolocation accuracy — grounding boxes are image-relative, not geodetic.
- No claim that the system is a safety-, legal-, or life-critical tool.
8. Where the evidence lives
| Topic | Evidence |
|---|---|
| All measured metrics | artifacts/**/*.json, verified by release/tools/verify_readme_metrics.py |
| Metric honesty rules | BENCHMARKS.md, EVALUATION.md |
| The grounding resolution rejection | docs/PHASE7_RESOLUTION_DECISION.md |
| The VLM rejection | artifacts/vlm/phase6_closure.json (why_acceptance_rejected) |
| B-07 / B-02 | DEPLOYMENT.md §8.1, §5 |
| Router residuals | architecture/04-router.md |
| Environment traps | REPRODUCIBILITY.md |