We’re exploring a simple idea: before asking a model to become a harbor specialist, help it build the skills that the job needs. That’s what Filling Before Advancing, or FBA, is about.
We study this with LLaVA-v1.5 and Qwen3-VL, using harbor imagery from RGB, SAR, PAN, and NIR sensors.
A small update: we’ve put the project page and examples online first. Checkpoints and inference instructions will follow after paper acceptance.
我们先把项目介绍和样例放在这里。论文录用后,会陆续补上模型权重和使用方法,欢迎回来看看。
Pages and files
Updated 2026-10-04. The website and HF preview pages are online; resource files are a separate release. Both HF repositories currently contain only README.md and .gitattributes.
| Artifact | File availability | Public version |
|---|---|---|
| Training & evaluation code | Not released · package in preparation | — |
| FBA / LLaVA-v1.5-7B weights | Not released · checkpoint package in preparation | — |
| FBA / Qwen3-VL-8B weights | Not released · checkpoint package in preparation | — |
| CPRS full dataset | Not released · source permissions under review | — |
| HarborEval evaluation package | Not released · public evaluation package in preparation | — |
— means no public resource version yet. Base-model names are not FBA release versions. Releases are planned progressively after paper acceptance. Public examples remain illustrative; HarborEval reference answers and scoring notes are currently private.
Shared checklist · License scope
How it works
The route has three steps:
- RS-Anchor — learn to read overhead imagery and connect it with language.
- Bridge-Conv — learn from related coastal scenes across different sensors.
- Scenario-EG — answer harbor questions, locate objects with a grid, and write reports using visible evidence.
A quick look at the results
On HarborEval, with scores on a 0–100 scale:
| Training route | LLaVA-v1.5 | Qwen3-VL |
|---|---|---|
| Direct-SFT | 57.95 | 81.09 |
| FBA | 70.29 | 83.37 |
HarborEval includes T5: Sensor-aware observability and T6: Evidence judgment. T5 asks what a sensor can support observing; T6 asks whether the visual evidence supports a claim.
That’s +12.34 and +2.28 score points over Direct-SFT. The GitHub results include the full comparisons and the areas where there’s still room to improve.
Take a look around
The project website is a good place to start: it has a short overview and eight interactive grid examples. You can also browse the CPRS dataset page or follow updates on GitHub.
We’ll add download and setup notes alongside the checkpoints. Base-model and source-data terms still apply; each release will explain the relevant licenses.
A note on licenses
We haven’t assigned one license to every part of this project. Original code and website text, each model variant, source images, annotations, and evaluation records have separate scopes. Model and data releases will include the applicable terms; upstream licenses still apply. See license scope.
当前不以一个许可覆盖所有内容。代码、网页文字、两种权重、源图像与标注将分别说明条件,上游材料的许可继续适用。
Cite the paper
@article{zong2026fba,
title = {Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs},
author = {Zong, Yuheng and Wang, Minghua and Zhao, Xin and Zhan, Zhi-Hui and Plaza, Antonio and Benediktsson, Jon Atli},
journal = {arXiv preprint arXiv:2607.22205},
year = {2026},
doi = {10.48550/arXiv.2607.22205},
url = {https://arxiv.org/abs/2607.22205}
}
