SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models
Organizations: Korea Advanced Institute of Science and Technology (KAIST)
Abstract
Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR--EO observations on tasks that benefit from complementary sensing.
Figures & tables
| Method | Venue | Backbone | FUSAR-Ship | MSTAR | SAR-ACD | |||
| 40-shot | 30% | 40-shot | 30% | 40-shot | 30% | |||
| ResNet-50 He et al. (2016) | CVPR’16 | ResNet-50 | – | 58.41 | – | 89.94 | – | 59.70 |
| Swin Transformer Liu et al. (2021) | ICCV’21 | Swin-B | – | 60.79 | – | 82.97 | – | 67.50 |
| BEiT Bao et al. (2021) | ICLR’22 | ViT-B | 59.70 | 71.13 | 40.70 | 69.75 | – | 79.77 |
| CROMA Fuller et al. (2023) | NeurIPS’23 | ViT-B | 83.71 | – | – | – | – | 88.99 |
| SAR-JEPA Li et al. (2023) | ISPRS JPRS’24 | ViT-B | 85.80 | – | 91.60 | – | 75.50 | – |
| Method / Feature | SAR | EO | ||||||
| 1-shot | 5-shot | 10-shot | 40-shot | 1-shot | 5-shot | 10-shot | 40-shot | |
| DINOv3 (teacher) | 38.17 | 50.72 | 56.01 | 64.35 | 59.51 | 75.96 | 83.34 | 90.88 |
| SARMAE ∗ | 37.22 | 48.64 | 53.97 | 64.72 | 43.56 | 59.57 | 66.49 | 78.36 |
| CoDe-MAE | – | – | 59.88 † | – | – | – | 81.18 † | – |
| SAREO-FM ( ) | 39.07 | 53.54 | 59.36 | 69.24 | 64.41 | 81.09 | 86.76 | 92.58 |
| Input | Method | Feature | Benchmark | |
| So2Sat LCZ42 | BigEarthNet-MM | |||
| SAR+EO | Random initialization | 40.56 | 54.29 | |
| EO | DINOv3 (teacher) | 54.25 | 64.73 | |
| EO | SAREO-FM | 58.14 | 64.42 | |
| SAR | SARMAE ∗ | 28.96 | 57.44 | |
| SAR | SAREO-FM | 29.59 | 60.47 | |
| Pretraining configuration | MSTAR (OA ) | FUSAR (OA ) | Avg. (OA ) | |
| Random initialization | 26.26 | 47.86 | 37.06 | – |
| Plain MAE He et al. (2022) | 66.16 | 76.46 | 71.31 | |
| DINOv3 supervision | 73.49 | 79.90 | 76.70 | |
| Decoupled semantic supervision | 77.11 | 82.55 | 79.83 | |
| Raw EO modality ( SAREO-FM ) | 78.41 | 84.00 | 81.21 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Section | Contents |
| Section A | Main-paper results and additional controls |
| Section B | Masking modes and training objectives |
| Section C | Pre-training corpus, datasets, and evaluation scope |
| Section D | Architecture and pre-training details |
| Section E | Feature readouts and downstream protocols |
| Section F | Inference cost |
| Dataset | Method / Feature | 1-shot | 2-shot | 5-shot | 10-shot | 20-shot | 40-shot |
| SARMAE ∗ | 54.19 | 64.95 | 70.03 | 78.18 | 83.44 | 88.35 | |
| FUSAR-Ship | SAREO-FM ( ) | 55.56 | 64.79 | 72.84 | 80.46 | 85.32 | 89.81 |
| SAREO-FM ( ) | 55.79 | 68.54 | 75.74 | 81.65 | 85.50 | 89.56 | |
| SARMAE ∗ | 35.09 | 39.56 | 47.12 | 51.68 | 60.30 | 70.11 | |
| SAREO-FM ( ) | 33.81 | 36.31 | 44.78 | 51.87 | 60.69 | 71.35 | |
| SAR-ACD | SAREO-FM ( ) | 33.64 | 36.87 | 48.10 | 53.65 | 61.65 | 71.72 |
| Configuration | SAR | EO | Queries | Total |
| Independent/shared masking | 64 | 64 | 256 | 384 |
| Complete EO dropping | 64 | 0 | 256 | 320 |
| Complete SAR dropping | 0 | 64 | 256 | 320 |
| Unpaired SAR training | 64 | 0 | 256 | 320 |
| SAR-only inference | 256 | 0 | 256 | 512 |
| EO-only inference | 0 | 256 | 256 | 512 |
| Sample / mode | Prob. | Student encoder image streams | Teacher input | Active objective |
| Paired, independent masks | 0.35 | Visible SAR + visible EO | Clean EO | |
| Paired, shared mask | 0.35 | Visible SAR + visible EO | Clean EO | |
| Paired, complete EO dropping | 0.15 | Visible SAR only | Clean EO, teacher only | |
| Paired, complete SAR dropping | 0.15 | Visible EO only | Clean EO | |
| Unpaired SAR | – | Visible SAR only | Unavailable | |
| Unpaired EO | – | Not sampled in this work | – | – |
| Dataset | Input | Classes | Metric | Protocol | SAR-1M relation | Role in this work |
| MSTAR | SAR | 10 | OA | 1–40-shot / 30% | In corpus | In-corpus target transfer |
| FUSAR-Ship | SAR | 10 | OA | 1–40-shot / 30% | In corpus | In-corpus ship transfer |
| SAR-ACD | SAR | 5 | OA | 1–40-shot / 30% | Out of corpus | Primary out-of-corpus SAR test |
| EuroSAT protocol | EO or SAR | 10 | OA | 1/5/10/40-shot | No out-of-corpus claim | Unimodal scene transfer |
| NWPU-RESISC45 | EO | 45 | OA | Few-shot / higher supervision | No out-of-corpus claim | EO scene transfer |
| AID | EO | 30 | OA | Few-shot / higher supervision | No out-of-corpus claim | EO scene transfer |
| Component | Specification |
| Input resolution | |
| Patch size / grid | / |
| SAR / EO tokens | 256 per available modality |
| Shared encoder | ViT-B/16, 12 blocks, |
| Attention / MLP | 12 heads / ratio 4 |
| Semantic queries | 256 learnable, spatially indexed |
| Component | Pretraining | Downstream |
| Shared SAREO encoder | Trainable | Used |
| Semantic queries | Trainable | Used |
| SAR/EO patch embeddings | Trainable | Used as available |
| SAR/EO reconstruction decoders | Trainable | Discarded |
| Semantic projector | Trainable | Discarded |
| DINOv3-7B teacher | Frozen | Discarded |
| Available input | |||
| SAR or EO | 768 | 768 | 1,536 |
| SAR + EO | 1,536 | 768 | 2,304 |