Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/LIMITATIONS.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 5.75 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/LIMITATIONS.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/docs/LIMITATIONS.md
-
curl -L -o LIMITATIONS.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/LIMITATIONS.md
5.75 kB
Limitations
An honest catalogue of everything this system does not do, does poorly, or does not know. Negative results and open items are listed here rather than omitted.
Status tags: OPEN · NOT RUN · REJECTED · DEFERRED · BY DESIGN.
1. Model quality
| # | Limitation | Detail |
|---|---|---|
| 1 | Grounding IoU is low in absolute terms | mean best IoU 0.2838 (canonical) / 0.2566 (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
| 2 | Grounding depends on the box convention | Two protocols and two decode variants produce very different numbers (0.2838 vs 0.1215 argmax). Absolute values are protocol-sensitive. |
| 3 | Optical-SAR accuracy is carried by common classes | accuracy 0.931 but macro-F1 0.434161. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. |
| 4 | Change-VQA is weak on rare classes | accuracy 0.697626 / macro-F1 0.378373 (and 0.651469 / 0.372309 on the second test set). |
| 5 | VQA is weak-but-related | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
| 6 | Optical-SAR returns a bare class index | the live service returns class_18, not a human-readable label. |
| 7 | Calibration made things worse | ECE 0.013755 → 0.014929. Retained only because it is part of the frozen config. |
| 8 | The VLM adapter is not accepted | metrics usable (exact_match 0.963), but the artifact is acceptance-rejected; the deployed path uses the unadapted model. |
| 9 | The router number is not a test result | 0.965116 is validation, ungated, n = 86. The test split was NOT RUN. |
2. Router residuals (known misroutes)
| # | Query | Behaviour | Note |
|---|---|---|---|
| 10 | "What is the new runway?" | reads change, not vqa |
residual ambiguity — "new" triggers a change reading |
| 11 | "How much built-up area was added?" | reads vqa (under-trigger) |
a change-style quantifier not caught by the router |
| 12 | "What changed between the earlier and later image?" with one asset | console reads change, dispatches change_vqa |
intentional (asset-count-aware dispatch) but visually surprising — documented in ARCHITECTURE.md §4 |
3. Evaluation gaps
| # | Gap | State |
|---|---|---|
| 13 | No system-level end-to-end benchmark | NOT RUN — none exists. No end-to-end accuracy is claimed. |
| 14 | Router test split | NOT RUN |
| 15 | Benchmark adapters | NOT RUN |
| 16 | End-to-end latency benchmark | NOT RUN (per-specialist latency only, incidentally) |
| 17 | Cross-dataset generalisation | NOT RUN — each specialist is evaluated on its own training-family split only |
| 18 | Human evaluation | NOT RUN |
| 19 | Robustness / adversarial evaluation | NOT RUN |
| 20 | BigEarthNet label semantics | the local subset is 100 % single-label vs the official 1–11 multi-label scheme, so metrics are not comparable to published numbers |
| 21 | Reliability diagram | the pre-scaling diagram is labelled as such; the calibrated curve is not plotted |
4. Operational limitations
| # | Limitation | State |
|---|---|---|
| 22 | B-07: transient tunnel gaps | a request can hang or return 504. Patch prepared, NOT deployed. OPEN |
| 23 | B-07 root shape | in auto mode a tunnel timeout falls through to the forward path, burning wake_timeout_s (120 s) on a 302; worst case ≈ 249 s. Measured. |
| 24 | Cold start is tens of seconds | Render free tier sleeps; the Codespace may be stopped. Documented, not hidden. |
| 25 | B-02: codespace_name trailing newline |
cosmetic; the wake path strips it. OPEN (cosmetic) |
| 26 | No database, auth, or queue | stateless gateway by design — no persistence of runs or users. |
| 27 | Deployment repos are private | their links 404 for an outside audience. BY DESIGN |
| 28 | deploy/ in the monorepo is stale/untracked |
not the deployed source. Trap. |
5. Packaging and licensing
| # | Limitation | State |
|---|---|---|
| 29 | No LICENSE file |
none exists in the source repository. OPEN — a licence must be selected before public release of the code. |
| 30 | Backbones are not redistributed | fetched from the Hub at run time; their licences are their own. |
| 31 | Artifact bloat | artifacts/ is ~3.7 GB, mostly reproducible caches, duplicate ZIPs and superseded checkpoints rather than released weights. |
6. Documentation caveats
| # | Caveat | State |
|---|---|---|
| 32 | docs/FINAL_DELIVERY_REPORT.md §6 is stale |
it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25. |
| 33 | The monorepo README.md was materially stale |
it described a hermetic frontend, an in-progress Render/Codespace, and a /v1/* contract. Superseded by this release's README. |
| 34 | hf/ docs were stale |
they asserted the project owns no weights and has no HF credentials. Both were false at release time. |
| 35 | Anatomy plate image variant | points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic. |
7. Explicit non-claims
- No claim of state-of-the-art performance on any benchmark.
- No claim of production readiness for the model quality — the deployment runs, but the models carry the limitations above.
- No claim that the trained heads generalise beyond their training-family test splits.
- No claim that calibration improves confidence.
- No claim that the VLM adapter is accepted for production use.