REMOTE SENSING / VISION–LANGUAGE POST-TRAINING
FillingBeforeAdvancing
Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs
RGB / CPRS SOURCESpecialization starts with the foundations.
General aerial understanding is only the beginning. Reliable harbor reasoning also depends on overhead-view semantics, sensor-aware observations, and transferable domain knowledge.
FBA fills these prerequisite capability gaps in order, before tuning evidence-grounded target behavior. We study this route through multi-source coastal harbor understanding.
盈科而后进Water fills each hollow before flowing onward.
Inspired by Mencius · from breadth to focus
- ≈810K
- CPRS supervision records
- 1,245
- HarborEval diagnostic items
- 4 sensors
- RGB · SAR · PAN · NIR
Three stages. Each with a purpose.
Supervision follows capability dependencies, from broad remote sensing semantics to focused, evidence-grounded behavior.
- S1
RS-Anchor
Anchor the semantics
Establish overhead-view visual–language alignment and broad RGB remote sensing understanding.
569,853 image–caption pairs - S2
Bridge-Conv
Bridge domains and sensors
Learn shared priors across coastal, port, water, dock, and ship scenes in RGB, SAR, PAN, and NIR.
187,296 SFT records - S3
Scenario-EG
Ground the specialization
Tune perception, spatial reasoning, localization, uncertainty, rejection, and grounded reporting.
53,000 SFT records
From pixels to structured supervision.
A closer look at the curated CPRS source collection: clear imagery, grid-based dialogue, multi-cell grounding, and explicit object relations.
- RGB 6,659
- source records · 19 object categories
- SAR 1,348
- source records · 11 object categories
- PAN 193
- source records · 12 object categories
- NIR 100
- source records · 9 object categories
This inspected source snapshot contains 8,300 image–dialogue records, with 2–4 turns per record. These counts describe the source collection, not the ≈810K stage-wise CPRS supervision corpus. Sensor examples are separate scenes, not paired acquisitions.
Explore eight selected cases.
Choose a sensor and a case. Inspect the dialogue, switch on the 3×3 grid, and follow the visible object relations.
The main vessel crosses two adjacent cells; the warehouse cluster provides a separate grounding target.
Grid names follow the existing project convention: top/middle/bottom × left/center/right. The colored cells show the first grid answer; a target may occupy multiple cells.
Illustrative dialogue adapted from the selected images using the project’s grid format. It is not verbatim source dialogue or an official benchmark annotation.
3×3 GRID / ILLUSTRATIVE DIALOGUE
Which grid cells visibly contain the main large vessel beside the central quay? Select all supported cells.
middle-center, middle-right.
View selected dialogue (2 turns)
A stronger route to harbor understanding.
Controlled route comparison on HarborEval, across two backbone families. Scores are on a 0–100 scale.
LLaVA-v1.5 vs. Direct-SFT
Qwen3-VL vs. Direct-SFT
| Training route | LLaVA-v1.5 | Qwen3-VL |
|---|---|---|
| Official / base | 46.22 | 70.37 |
| Direct-SFT | 57.95 | 81.09 |
| Strongest Collapsed-SFT | 55.74 | 79.36 |
| Filling Before Advancing | 70.29 | 83.37 |
Gains are overall score points, not relative percentages. Qwen3-VL improves across all eight components; LLaVA-v1.5 gains especially in T5 observability, T6 evidence judgment, and T8 rejection, while object recognition and grid grounding remain below Direct-SFT.
View T1–T8 diagnostic results
| Track | Diagnostic component | LLaVA-v1.5 | Qwen3-VL |
|---|---|---|---|
| T1 | Object/scene VQA | 75.07 → 73.47 | 87.99 → 92.42 |
| T2 | Functional-zone interpretation | 66.67 → 67.22 | 78.33 → 81.11 |
| T3 | Spatial-relation reasoning | 60.67 → 69.10 | 81.46 → 83.15 |
| T4 | Multi-cell grid grounding | 48.37 → 43.82 | 66.06 → 68.61 |
| T5 | Sensor-aware observability | 50.88 → 80.12 | 81.29 → 82.46 |
| T6 | Evidence judgment | 52.03 → 69.11 | 78.05 → 79.67 |
| T7 | Evidence-grounded report generation | 72.60 → 73.70 | 80.23 → 81.32 |
| T8 | Non-harbor and near-domain rejection | 37.28 → 85.80 | 95.27 → 98.22 |
Scores show Direct-SFT → FBA. T5 tests sensor-aware observability, not sensor classification; T6 tests visual claim support and uncertainty.
Continue with the research.
Pages are online. Files will follow.
Updated 2026-10-04 · The website and HF preview pages are online; resource files are a separate release. Both HF repositories currently contain only README.md and .gitattributes.
| Artifact | File availability | Public version |
|---|---|---|
| Training & evaluation code | Not released · package in preparation | — |
| FBA / LLaVA-v1.5-7B weights | Not released · checkpoint package in preparation | — |
| FBA / Qwen3-VL-8B weights | Not released · checkpoint package in preparation | — |
| CPRS full dataset | Not released · source permissions under review | — |
| HarborEval evaluation package | Not released · public evaluation package in preparation | — |
— means no public resource version yet. Releases are planned after paper acceptance. HarborEval reference answers and scoring notes currently remain private.
Citation
@article{zong2026fba,
title = {Filling Before Advancing: Capability-Gap-Driven Post-Training
for Scenario-Specialized Remote Sensing MLLMs},
author = {Zong, Yuheng and Wang, Minghua and Zhao, Xin and Zhan, Zhi-Hui
and Plaza, Antonio and Benediktsson, Jon Atli},
journal = {arXiv preprint arXiv:2607.22205},
year = {2026},
doi = {10.48550/arXiv.2607.22205},
url = {https://arxiv.org/abs/2607.22205}
}