Multimodal Remote Sensing Image Registration: A Comprehensive Review, Challenges and Prospects
Organizations: Faculty of Geosciences and Engineering, Southwest Jiaotong University. Chengdu, 611756, China. · Yunnan Key Laboratory of Quantitative Remote Sensing / Yunnan International Joint Laboratory for Integrated Sky-Ground Intelligent Monitoring of Mountain Hazards, Faculty of Land Resources Engineering, Kunming University of Science and Technology. Kunming, 650093, China. · School of Software Engineering, Beijing Jiaotong University. Beijing, 100044, China.
Abstract
Multimodal remote sensing image registration is a crucial prerequisite for the collaborative processing and downstream application of remote sensing data, such as image fusion, change detection, and target recognition. However, significant variations in radiometry, geometry, scale, viewpoint, and time often exist between multimodal images. These differences, driven by varying sensor geometries, physical radiation mechanisms, imaging platforms, and environmental disturbances, pose severe challenges to achieving high-precision, robust registration. This paper systematically reviews the progress of mainstream multimodal remote sensing image registration methods. Based on their registration pipelines, existing approaches are categorized into three main types: region-based, feature-based, and deep learning-based methods. We detail the core principles, representative algorithms, advantages, and limitations of each category. Additionally, we summarize publicly available multimodal image datasets in the remote sensing domain, analyzing their specific characteristics and applicable scenarios. Finally, we highlight current bottlenecks in high-precision registration research and outline future development trends. This review aims to provide a comprehensive reference and valuable insights for researchers in related fields.
Figures & tables
| Representative algorithm | Core ideas |
| Chen and Varshney (2000) | Image pyramids & MI; Robust to nonlinear radiative differences |
| Inglada (2002) | Evaluation of MI; Good comprehensive performance |
| Chen et al. (2003) | GPVE optimization; Solves interpolation artifacts |
| Cole-Rhodes et al. (2003) | SPSA gradient policy optimization; Accelerated convergence |
| Suri and Reinartz (2010) | Feature selection for high-resolution images; Improved performance |
| NCMI ( Parmehr et al., 2014 ) | Histogram-based non-parametric PDF estimation; Improved performance |
| Method type | Representative algorithm | Core ideas | Acquisition channels |
| Sparse structural features | HOG NCC ( Li et al., 2013 ) | HOG & NCC combination | - |
| DLSS ( Ye et al., 2017b ) | Dense local self-similar description; Shape attributes | [Code] | |
| HOPC ( Ye et al., 2017a ) | Histogram of Oriented Phase Consistency; Global satellite image processing | [Code] | |
| HOMPC ( Fu et al., 2018 ) | Log-Gabor filtering; Local amplitude & phase consistency | - | |
| RLSS ( Xiong et al., 2020 ) | Spearman rank correlation & LSS; Local shape attribute description | - | |
| Dense structural features | CFOG ( Ye et al., 2019 ) | Pixel-by-pixel structural matching; HOG, LSS & CFOG framework | [Code] |
| Method type | Representative algorithm | Core ideas | Acquisition channels |
| Gradient optimization | SAR-SIFT ( Dellinger et al., 2015 ) | Ratio gradient via ROEWA operator; Robust to SAR speckle noise | [Code] |
| PSO-SIFT ( Ma et al., 2017 ) | Gradient derivative in Gaussian scale space; Robust to intensity variations | [Code] | |
| OS-SIFT ( Xiang et al., 2018 ) | Multi-scale Sobel & ROEWA operators; Consistent optical-SAR gradients | [Code] | |
| CoFSM ( Yao et al., 2022 ) | Scale space via co-occurrence filtering; Butterworth filter gradient optimization | [Code] | |
| Local self-similarity | DOBSS ( Sedaghat and Ebadi, 2015 ) | Related value histogram; Internal geometric layout capture | - |
| HOSS ( Sedaghat and Mohammadi, 2019 ) | Adaptive log-extreme spatial structure & max correlation rotation index map | [Code] |
| Representative algorithm | Core ideas | Acquisition channels |
| Merkle et al. (2017) | Siamese FCN; Dot product similarity for optical-SAR offsets | - |
| Hughes et al. (2018) | Independent convolutional streams (Pseudo-Siamese); Accommodates image heterogeneity | - |
| SFCNet ( Zhang et al., 2019a ) | Siamese FCN with convolutional connections; Novel feature distance loss | - |
| Li et al. (2021a) | Semantic template matching framework | [code] |
| MSF SiamUNet-7 ( Liu et al., 2021 ) | MSF SiamUNet-7; Phase-structure features; Local texture & global structure fusion | - |
| OSMNet ( Zhang et al., 2022a ) | Multi-frequency channel attention; adaptive weighted loss; Structural consistency | [Code] |
| Representative algorithm | Core ideas | Acquisition channels |
| Ye et al. (2018a) | CNN features (FC & conv layers); SIFT | - |
| Dong et al. (2019) | DescNet deep feature vectors; Robust keypoint descriptors | - |
| Ma et al. (2019b) | Deep & handcrafted feature combination; Fine matching & transformation | - |
| CMM-Net ( Lan et al., 2021 ) | CNN high-dimensional feature maps; Keypoint detection & description | [Code] |
| MAP-Net ( Cui et al., 2022 ) | SPAP module context integration; Attention-weighted dense features | - |
| Sim-CSPNet ( Xiang et al., 2022a ) | Sim-CSPNet architecture; Rapid deep & shallow feature extraction | - |
| Representative algorithm | Core ideas |
| Arar et al. (2020) | Special loss function; Strict geometric position preservation; Texture style change |
| Zhao et al. (2022) | Smooth CycleGAN; Novel loss function; Improved generation quality |
| Chen et al. (2022) | Modified CycleGAN; Identity mapping loss for color constraint |
| CDA-GAN ( Wu et al., 2024 ) | Attention mechanism; High-frequency texture retention; Pseudo-SAR generation |
| CSTNet ( Bi et al., 2025 ) | Visible-to-pseudo-infrared conversion; Unified modal features; Multi-scale cascaded network |
| Seg-CycleGAN ( Zhang et al., 2025a ) | Semantic segmentation guidance; SAR-to-optical conversion; Structural information retention |
| Representative algorithm | Core ideas | Acquisition channels |
| Wang et al. (2018) | Closed-loop feedforward/feedback mapping; Self-learning for small datasets | - |
| Hughes et al. (2020) | Sequential networks: feature extraction, multi-scale matching, error rejection | - |
| D-DenseNet ( Li et al., 2022a ) | Weight-sharing D-DenseNet; Direct learning of 4-corner displacement parameters | - |
| Li et al. (2021a) | Semantic template matching framework; Spatial location probability mapping | [Code] |
| MU-Net ( Ye et al., 2022a ) | Unsupervised MU-Net; Multi-scale parameter learning; Structural similarity loss | [Code] |
| Li et al. (2023b) | Self-attention fusion framework; Cross-modal feature regression via point displacements | [Code] |
| Datasets | Image types | Quantity | Size | Resolution | Download |
| SARptical ( Wang and Zhu, 2018 ) | Optical-SAR | 10,108 pairs | 112 112 | Opt 0.2m / SAR 1m | [Link] |
| SEN1-2 ( Schmitt et al., 2018 ) | Optical-SAR | 282,384 pairs | 256 256 | 10m | [Link] |
| University-1652 ( Zheng et al., 2020 ) | Multiple | 1,652 buildings | - | - | [Link] |
| OS-dataset ( Xiang et al., 2020b ) | Optical-SAR | 10,692 pairs | 256 256 | 1m | [Link(vriw)] |
| QXS-SAROPT ( Huang et al., 2021 ) | Optical-SAR | 20,000 pairs | 256 256 | 1m | [Link] |
| Multimodal Image Matching Datasets ( Jiang et al., 2021 ) | Multiple | 164 pairs | - | - | [Link] |