SatQuery / MODEL_CARD.md
thundercode's picture
release: add MODEL_CARD.md
3a34048 verified
|
Raw History Blame
7 kB
---
language: en
license: other
library_name: pytorch
tags:
- remote-sensing
- satellite-imagery
- earth-observation
- change-detection
- visual-grounding
- image-captioning
- visual-question-answering
- optical-sar-fusion
- lora
- peft
pipeline_tag: image-to-text
config_hash: 78f1e3700da15aa1
---
# Model Card — SatQuery AI
SatQuery AI answers natural-language questions about satellite imagery using a **router + specialists**
design. This card covers the **six trained artifacts** released by the project. It is deliberately
explicit about what is measured, what is not, and what was rejected.
> **The six trained artifacts are small modules on top of frozen, publicly-pinned backbones. No
> backbone weights are redistributed by this release** — they are fetched from the Hugging Face Hub at
> run time, pinned by revision.
Machine-readable identities (byte counts and sha256) are in
[`models/manifest.json`](models/manifest.json) and [`models/checksums.sha256`](models/checksums.sha256),
**generated by reading the files**. Where this card and the generated manifest disagree, the manifest
wins.
---
## 1. Artifacts in this release
| # | Task | Kind | File | Bytes | sha256 (first 16) |
|---|---|---|---|---|---|
| 1 | `change` | trained head | `head.pt` | 63,231,009 | `c5ef31277b67aa01` |
| 2 | `change_vqa` | trained head | `head.pt` | 5,822,809 | `cfae5e43b97ca930` |
| 3 | `optical_sar` | trained head | `head.pt` | 14,427,457 | `785815729a3a39fc` |
| 4 | `grounding` | trained head | `head.pt` | 12,639,041 | `93432f7034be91a8` |
| 5 | `router` | trained adapter | `adapter.pt` | 211,961 | `8527c3ed28a293e1` |
| 6 | `vlm` | LoRA adapter | `adapter_model.safetensors` | 34,798,048 | `07c76a75fa046248` |
Two of these hashes (`change_vqa`, `vlm`) **agree exactly** with hashes recorded independently at
promotion time — an external cross-check, not a self-consistency claim.
## 2. Backbone dependencies (frozen, pinned by revision)
| Role | Repository | Revision |
|---|---|---|
| Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` |
| VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` |
| Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` |
| Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` |
| Change | STANet-style (ResNet-18 + PAM) | trained in-project |
## 3. Intended use
- **Research and demonstration** of a modular, CPU-first remote-sensing QA system.
- **Routing and dispatch** of natural-language queries to the appropriate specialist.
- **Reproducible evaluation** of each specialist on its own documented split.
## 4. Out-of-scope use
- **Safety-, legal- or life-critical decisions.** No accuracy, calibration or robustness guarantee is
offered for any high-stakes use.
- **Operational geospatial production** without independent validation.
- **Any use of the VLM adapter as a production model** — it is acceptance-rejected (§6).
- **Treating per-specialist metrics as system-level accuracy.** No end-to-end benchmark exists (§7).
## 5. Measured performance
| Task | Metric | Value | Split / protocol |
|---|---|---|---|
| change | pooled IoU / macro IoU / pooled F1 | **0.8122 / 0.8457 / 0.8964** | LEVIR-CD-256 test, n = 2048 |
| grounding | mean best IoU / recall@0.5 | **0.2838 / 0.2198** canonical; **0.2566 / 0.1938** matched6 | VRSBench, n = 16159 |
| grounding | head-argmax / zero-shot baseline IoU | **0.1215 / 0.0972** | canonical |
| optical_sar | accuracy / macro F1 | **0.931 / 0.434161** | held-out test, n = 4000, 19 classes |
| change_vqa | accuracy / macro F1 | **0.697626 / 0.378373** (test); **0.651469 / 0.372309** (test2) | two test sets |
| vlm | exact_match / F1 | **0.963 / 0.96432** | frozen 1000-question subset |
| router | overall **ungated** accuracy | **0.965116** | val, n = 86 — **TEST NOT RUN** |
**Every value is checked against its artifact** by `tools/verify_readme_metrics.py`.
### 5.1 Calibration — reported as a negative result
Temperature scaling is enabled (T = 0.9772732, fit on val n = 16,441). **ECE worsened**:
0.013755 → **0.014929**. It is retained because it is part of the frozen configuration, not because it
helped.
## 6. Acceptance status
| Artifact | Metrics | Acceptance |
|---|---|---|
| change | VERIFIED | accepted (shipped) |
| grounding | measured (2 protocols) | shipped |
| optical_sar | measured | **ruling OPEN** |
| change_vqa | measured (2 test sets) | **ruling OPEN** |
| router | measured (val only) | shipped; test NOT RUN |
| **vlm** | usable (exact_match 0.963) | **ACCEPTANCE-REJECTED** |
**`USABLE_VERIFIED` ≠ `ACCEPTANCE-ACCEPTED`.** The VLM adapter works and is not promoted; the deployed
caption/VQA path uses the **unadapted** model.
## 7. Evaluation gaps (stated, not hidden)
- **No system-level end-to-end benchmark exists.** None is claimed.
- **Router test split: NOT RUN.**
- **Benchmark adapters: NOT RUN.**
- **Cross-dataset generalisation: NOT RUN.**
- **Human and robustness evaluation: NOT RUN.**
## 8. Limitations
- Grounding absolute IoU is low (0.28) and protocol-sensitive.
- Optical-SAR accuracy is carried by common classes (macro-F1 0.434161).
- The BigEarthNet local subset is **100 % single-label** vs the official 1–11 multi-label scheme, so
its metrics are **not comparable** to published numbers.
- The optical-SAR service returns a bare class index, not a label.
- Known router residuals exist (e.g. *"What is the new runway?"* reads `change`).
- No `LICENSE` file exists in the source repository.
See [`docs/LIMITATIONS.md`](docs/LIMITATIONS.md) for the full catalogue.
## 9. Training summary
Small modules on frozen backbones; seed 42; every artifact records the frozen config hash
`78f1e3700da15aa1`. Router and change/grounding/fusion heads train on CPU; the VLM LoRA adapter and
the change-VQA head were trained on **external GPUs** (the latter via a documented Kaggle run). Full
detail in [`docs/TRAINING.md`](docs/TRAINING.md).
## 10. Provenance and verification
| Item | Location |
|---|---|
| Byte-exact manifest | `models/manifest.json` |
| Checksums | `models/checksums.sha256` |
| Metric verification tool | `tools/verify_readme_metrics.py` |
| Metric verification output | `tools/readme_metrics_report.txt` |
| Full documentation | `docs/` |
| Release manifest | `RELEASE_MANIFEST.md` |
## 11. Licence
The project ships **no licence file**; a licence must be selected by the owner before public release
of the *code*. **Model weights carry the terms of their backbone licences** — consult each backbone's
Hugging Face page. Backbones are not redistributed here.
## 12. Citation
If you use this work, cite the project repository:
```bibtex
@misc{satquery_ai_2026,
title = {SatQuery AI: A Modular Router-and-Specialists System for Satellite Imagery Question Answering},
author = {SatQuery AI},
year = {2026},
note = {Public release: https://github.com/Anish-lab-blip/SatQuery-AI}
}
```