FBA.

REMOTE SENSING / VISION–LANGUAGE POST-TRAINING

FillingBeforeAdvancing

Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

Yuheng Zong · Minghua Wang · Xin Zhao
Zhi-Hui Zhan · Antonio Plaza · Jon Atli Benediktsson

Clear RGB aerial image of a harbor with ships, piers, and warehouses, from the CPRS source showcaseRGB / CPRS SOURCE
THE PRINCIPLEFill the gaps. Then advance.
From broad foundations to grounded specializationExplore the work ↓
01 / THE IDEA

Specialization starts with the foundations.

General aerial understanding is only the beginning. Reliable harbor reasoning also depends on overhead-view semantics, sensor-aware observations, and transferable domain knowledge.

FBA fills these prerequisite capability gaps in order, before tuning evidence-grounded target behavior. We study this route through multi-source coastal harbor understanding.

盈科而后进

Water fills each hollow before flowing onward.

Inspired by Mencius · from breadth to focus
≈810K
CPRS supervision records
1,245
HarborEval diagnostic items
4 sensors
RGB · SAR · PAN · NIR
02 / THE ROUTE

Three stages. Each with a purpose.

Supervision follows capability dependencies, from broad remote sensing semantics to focused, evidence-grounded behavior.

  1. S1

    RS-Anchor

    Anchor the semantics

    Establish overhead-view visual–language alignment and broad RGB remote sensing understanding.

    569,853 image–caption pairs
  2. S2

    Bridge-Conv

    Bridge domains and sensors

    Learn shared priors across coastal, port, water, dock, and ship scenes in RGB, SAR, PAN, and NIR.

    187,296 SFT records
  3. S3

    Scenario-EG

    Ground the specialization

    Tune perception, spatial reasoning, localization, uncertainty, rejection, and grounded reporting.

    53,000 SFT records
View the full training diagram
The three ordered FBA training stages and their supervision roles
The original research diagram · open image for a closer look.
03 / INSIDE THE DATA

From pixels to structured supervision.

A closer look at the curated CPRS source collection: clear imagery, grid-based dialogue, multi-cell grounding, and explicit object relations.

RGB 6,659
source records · 19 object categories
SAR 1,348
source records · 11 object categories
PAN 193
source records · 12 object categories
NIR 100
source records · 9 object categories

This inspected source snapshot contains 8,300 image–dialogue records, with 2–4 turns per record. These counts describe the source collection, not the ≈810K stage-wise CPRS supervision corpus. Sensor examples are separate scenes, not paired acquisitions.

Choose a sensor and a case. Inspect the dialogue, switch on the 3×3 grid, and follow the visible object relations.

The main vessel crosses two adjacent cells; the warehouse cluster provides a separate grounding target.

Grid names follow the existing project convention: top/middle/bottom × left/center/right. The colored cells show the first grid answer; a target may occupy multiple cells.

Illustrative dialogue adapted from the selected images using the project’s grid format. It is not verbatim source dialogue or an official benchmark annotation.

RGB harbor image from the CPRS source showcase
RGB / CPRS · 00003

3×3 GRID / ILLUSTRATIVE DIALOGUE

Which grid cells visibly contain the main large vessel beside the central quay? Select all supported cells.

middle-center, middle-right.

OBJECT RELATIONS
ship → docked_at → pierwarehouse → next_to → warehouseship → near → warehouse
View selected dialogue (2 turns)
View sample record ↗
04 / CORE RESULTS

A stronger route to harbor understanding.

Controlled route comparison on HarborEval, across two backbone families. Scores are on a 0–100 scale.

+12.34

LLaVA-v1.5 vs. Direct-SFT

+2.28

Qwen3-VL vs. Direct-SFT

HarborEval controlled training route comparison
Training routeLLaVA-v1.5Qwen3-VL
Official / base46.2270.37
Direct-SFT57.9581.09
Strongest Collapsed-SFT55.7479.36
Filling Before Advancing70.2983.37

Gains are overall score points, not relative percentages. Qwen3-VL improves across all eight components; LLaVA-v1.5 gains especially in T5 observability, T6 evidence judgment, and T8 rejection, while object recognition and grid grounding remain below Direct-SFT.

View T1–T8 diagnostic results
HarborEval tracks: Direct-SFT to FBA
TrackDiagnostic componentLLaVA-v1.5Qwen3-VL
T1Object/scene VQA75.07 → 73.4787.99 → 92.42
T2Functional-zone interpretation66.67 → 67.2278.33 → 81.11
T3Spatial-relation reasoning60.67 → 69.1081.46 → 83.15
T4Multi-cell grid grounding48.37 → 43.8266.06 → 68.61
T5Sensor-aware observability50.88 → 80.1281.29 → 82.46
T6Evidence judgment52.03 → 69.1178.05 → 79.67
T7Evidence-grounded report generation72.60 → 73.7080.23 → 81.32
T8Non-harbor and near-domain rejection37.28 → 85.8095.27 → 98.22

Scores show Direct-SFT → FBA. T5 tests sensor-aware observability, not sensor classification; T6 tests visual claim support and uncertainty.

Explore all results and ablations ↗
WHERE WE GO NEXT

Gravity and Flight

From filling capability gaps to using strong capabilities reliably.

BACKGROUND
Stronger base models already recognize much of a scene. Limited, high-quality supervision should help them use that knowledge reliably.
THE QUESTION
Two harbors can contain similar ships and quays yet support different judgments. Which evidence should change the conclusion, and which should leave it intact?
THE GOAL
Keep what the evidence still supports, revise judgments when decisive details change, and become less specific when support is missing. This is the focus of our follow-up work.
Follow Gravity and Flight on GitHub
05 / EXPLORE & CITE

Continue with the research.

Pages are online. Files will follow.

Updated 2026-10-04 · The website and HF preview pages are online; resource files are a separate release. Both HF repositories currently contain only README.md and .gitattributes.

Resource file availability and public versions
ArtifactFile availabilityPublic version
Training & evaluation codeNot released · package in preparation—
FBA / LLaVA-v1.5-7B weightsNot released · checkpoint package in preparation—
FBA / Qwen3-VL-8B weightsNot released · checkpoint package in preparation—
CPRS full datasetNot released · source permissions under review—
HarborEval evaluation packageNot released · public evaluation package in preparation—

— means no public resource version yet. Releases are planned after paper acceptance. HarborEval reference answers and scoring notes currently remain private.

Citation

@article{zong2026fba,
  title   = {Filling Before Advancing: Capability-Gap-Driven Post-Training
             for Scenario-Specialized Remote Sensing MLLMs},
  author  = {Zong, Yuheng and Wang, Minghua and Zhao, Xin and Zhan, Zhi-Hui
             and Plaza, Antonio and Benediktsson, Jon Atli},
  journal = {arXiv preprint arXiv:2607.22205},
  year    = {2026},
  doi     = {10.48550/arXiv.2607.22205},
  url     = {https://arxiv.org/abs/2607.22205}
}