cs.CVOct 7, 2026

SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models

Authors: Jeonghyeok Do, Munchurl Kim

Organizations: Korea Advanced Institute of Science and Technology (KAIST)

Abstract

Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR--EO observations on tasks that benefit from complementary sensing.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cross-modal learning for SAR target recognition using optical vision foundation models

    Sep 7, 2026Lucas Hirsch, James R. Hopgood, Javid Khan +2Cross-Modal AlignmentTransfer Learning

  2. TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

    May 12, 2026Zhuoyu Cai, Dou Quan, Ning Huyan +3Remote SensingSynthetic Aperture Radar