Organizations: College of Computer Science and Engineering, Dalian Minzu University, Dalian, China · School of Artificial Intelligence and Robotics, Hunan University, Hunan, China · Department of Automation, Tsinghua University, Beijing, China
Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art (SOTA) performance. Code: github.com/Wangjiayu0512/SIDFusion.
Figures & tables
Figure 1: Existing decomposition paradigms vs Ours. We introduce signal-level decomposition and image-level reconstruction pretext tasks for accurate decomposition and high-quality fusion.
Figure 2: Schematic of the proposed model. It adopts a two-stage self-supervised paradigm. Stage I: feature decomposition via 1D signal-level and 2D image-level pretext tasks; Stage II: fuses the decomposed features to reconstruct the final fused image.
Figure 3: Qualitative comparison of various fusion models.
VIF Task
M 3 FD
MSRS
TNO
Methods
Pub/Year
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
CDDFuse
CVPR 23
0.5741
0.8123
0.4554
0.4746
3.8730
0.4410
0.8698
0.7576
0.8229
0.5399
0.5669
4.9089
0.5106
0.8267
0.4654
0.8090
0.3789
0.4504
3.0691
0.4428
0.7874
LRRNet
TPAMI 23
0.6353
0.8073
0.3709
0.4346
2.8124
0.3550
0.7209
0.6608
0.8082
0.3372
0.3928
2.8662
0.2933
0.5067
0.3686
0.8062
0.2303
0.4831
2.3868
0.3685
0.7011
EMMA
CVPR 24
0.5564
0.8118
0.4602
0.4775
3.7712
0.4690
0.7979
0.6413
0.8161
0.4961
0.5432
4.1637
0.5230
0.7840
0.4356
0.8081
0.3550
0.5112
2.9089
0.4847
0.8072
TC-MoA
CVPR 24
0.5011
0.8096
0.5133
0.4934
3.3330
0.4603
0.8538
0.7446
0.8119
0.4833
0.5558
3.4925
0.4924
0.8825
0.4068
0.8071
0.4020
0.5124
2.5747
0.4312
0.8664
Text-Difuse
NeurIPS 24
0.4262
0.8057
0.4323
0.3398
2.0535
0.1574
0.2566
0.7252
0.8032
0.5378
0.3660
1.3095
0.1570
0.2122
0.2977
0.8047
0.1167
0.4296
1.8773
0.2638
0.4410
Table 1: Quantitative comparison of various fusion models on the VIF and MIF tasks. Best is bold ; second best is underlined .
VIF Task
MIF Task
Task
Case
Configurations
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
Task
Case
Configurations
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
VIF
I
w/o signal-level Dec.
0.6282
0.8181
0.5350
0.4716
3.8631
0.4391
0.8729
MIF
I
w/o signal-level Dec.
0.9012
0.8071
0.3627
0.6715
3.6129
0.3689
0.9127
w/o DWT
0.6425
0.8192
0.5551
0.4860
3.9812
0.4534
0.8847
w/o DWT
0.9185
0.8090
0.3825
0.6880
3.8310
0.3897
0.9312
w/ IDWT ← learnable Rec.
0.6531
0.8198
0.5631
0.4925
4.0200
0.4594
0.8911
w/ IDWT ← learnable Rec.
0.9253
0.8096
0.3928
0.6905
3.9025
0.3954
0.9361
w/ signal- ← image-level Dec.
0.6380
0.8185
0.5460
0.4793
3.9022
0.4468
0.8781
w/ signal- ← image-level Dec.
0.9108
0.8086
0.3755
0.6810
3.7420
0.3798
0.9275
Default setups (ours)
0.6608
0.8209
0.5735
0.4998
4.1050
0.4694
0.8988
Default setups (ours)
0.9369
0.8108
0.4034
0.7008
4.0184
0.4034
0.9439
Table 2: Quantitative results of ablation experiments on VIF and MIF tasks.
Figure 4: Visualization of intermediate 1D signals. The learned common signal cs(l) follows the shared trend of the source signals sA(l) and sB(l) .
Figure 5: Qualitative results on downstream tasks. Left: object detection performance using fused images generated by different fusion methods. Right: medical image segmentation. Our fused images provide more accurate results.
Figure 6: Broader impact of our paradigm. Refined variants (denoted by “*”) obtain fused images with clearer structures and more salient target details, as highlighted by the red and yellow circles.
Figure 7: Qualitative results on multi-focus image fusion, showing that the proposed signal-level decomposition paradigm can be naturally transferred to other image fusion tasks.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Visual Attribute
Pixel-level Manifestation
Commonly Used Constraints
Intensity / Brightness
Absolute pixel magnitude, local luminance
L1 loss, L2 loss, reconstruction loss
Contrast
Relative intensity difference between a pixel and its local neighborhood, or between foreground and background
SSIM contrast term, local contrast loss, entropy-based or information-based constraints
Edge
Sharp local intensity variation around object boundaries and structural transitions
Gradient loss, Sobel loss, Laplacian loss
Texture
fine details, and high-frequency responses
Gradient loss, frequency-domain loss
Boundary
Continuous object outlines and region boundaries
SSIM loss, gradient loss, boundary-aware loss
Structural Layout
Spatial arrangement of objects, organs, roads, targets, or anatomical regions
SSIM loss, perceptual loss
Appendix
Table 3: Representative visual attributes coupled in image pixels and their commonly used constraints in MMIF.
Figure 8: Full qualitative comparisons on six datasets. The top three groups show MIF task and the bottom three groups show VIF cases. In each row, the first two columns are the source image pairs and the remaining columns are fused results of different methods, with our method in the rightmost column, yielding sharper structures and richer complementary details.
Method
People
Car
Bus
Motorcycle
Lamp
Truck
mAP
IR
0.512
0.581
0.447
0.468
0.403
0.452
0.477
VIS
0.563
0.638
0.483
0.521
0.436
0.494
0.522
CDDFuse
0.662
0.726
0.618
0.634
0.573
0.602
0.636
LRRNet
0.641
0.707
0.597
0.612
0.551
0.582
0.615
EMMA
0.651
0.719
0.612
0.631
0.562
0.593
0.628
TC - MoA
0.642
0.715
0.606
0.623
0.559
0.601
0.624
Appendix
Table 4: AP for object detection on source images and fused images from different methods. mAP is the mean of the six category APs.
Metrics
T1
Flair
CDDFuse
LRRNet
EMMA
TC - MoA
Text - Diffuse
IoU
0.615
0.657
0.762
0.731
0.743
0.689
0.685
Dice
0.607
0.727
0.850
0.836
0.841
0.829
0.751
Metrics
CCF
SigFusion
BSAFusion
Mask - Difuser
MTG - Fusion
C2RF
Ours
IoU
0.661
0.735
0.745
0.723
0.775
0.772
0.790
Dice
0.752
0.770
0.832
0.829
0.846
0.848
0.856
Appendix
Table 5: Quantitative comparison for medical image segmentation.
Figure 9: Visualization of common and unique feature maps produced by different decomposition strategies. For each source pair, we display the estimated common features and the modality-unique features. Compared with existing methods that generate two inconsistent common features, our signal-level decomposer yields a single common map and better-separated unique maps.
Figure 10: t-SNE visualization of low-/high-frequency components and learned decomposed features. The common features are close to the low-frequency components of both modalities, while the unique features are closer to their corresponding high-frequency components, supporting the design of extracting common information from low frequencies and unique information from high frequencies.
Figure 11: Visualization of training curves. We compare the proposed signal-level objective with a conventional 2D image-domain mixed-loss strategy in terms of training loss and feature similarity, including SSIM, Pearson correlation, and cosine similarity. The results show that our signal-level objective converges faster and achieves a lower final loss, while maintaining higher and more stable similarity between the learned common features and source signals.
Method
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
DIDFuse
0.6130
0.8083
0.3570
0.4830
3.2370
0.3610
0.9020
DIDFuse*
0.7466
0.8196
0.4198
0.5299
3.9038
0.4596
0.9101
Gain
21.8% ↑
1.4% ↑
17.6% ↑
9.7% ↑
20.6% ↑
27.3% ↑
0.9% ↑
CDDFuse
0.6610
0.8107
0.4010
0.6940
3.9820
0.3990
0.9380
CDDFuse*
0.7582
0.8277
0.4523
0.7051
4.3563
0.4736
0.9446
Gain
14.7% ↑
2.1% ↑
12.8% ↑
1.6% ↑
9.4% ↑
18.7% ↑
0.7% ↑
Appendix
Table 6: The refinement effects results of our paradigm. Original vs. upgraded models (* denotes our paradigm inserted).
Method
Pub/Year
Dataset: LYTRO
Dataset: MFFW
QG↑
QM↑
QP↑
MI↑
SD↑
VIFF↑
QG↑
QM↑
QP↑
MI↑
SD↑
VIFF↑
CUNet
TPAMI 20
0.526
0.553
0.696
5.441
58.70
1.022
0.482
0.455
0.552
4.593
56.33
0.847
U2Fusion
TPAMI 20
0.580
0.480
0.742
5.677
58.37
1.086
0.537
0.405
0.611
4.876
55.26
0.825
DeFusion
ECCV 22
0.455
0.325
0.660
5.984
54.39
1.028
0.418
0.296
0.518
5.137
51.55
0.876
DIFNet
CVPR 22
0.437
0.325
0.688
5.774
49.67
1.032
0.422
0.309
0.577
4.867
46.66
0.890
FusionDiff
ESWA 23
0.629
0.821
0.783
6.554
56.13
1.188
0.545
0.602
0.659
5.334
53.27
0.993
Appendix
Table 7: Quantitative comparison on the MFIF task against 9 competing methods using 6 evaluation metrics on the LYTRO and MFFW datasets. Although MFIF may provide fused ground truth for evaluation, the decomposition of common and unique features still lacks direct supervision. Our method remains applicable in this setting and achieves strong overall performance.
I. Effect of Signal-Level Decomposer
VIF task
MIF task
Configuration settings
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
VIF: M 3 FD
MIF: MRI-CT
w/o signal-level decomposer
0.6282
0.8181
0.5350
0.4716
3.8631
0.4391
0.8729
0.9012
0.8071
0.3627
0.6715
3.6129
0.3689
0.9127
w/o DWT
0.6425
0.8192
0.5551
0.4860
3.9812
0.4534
0.8847
0.9185
0.8090
0.3825
0.6880
3.8310
0.3897
0.9312
w/ IDWT ← learnable reconstructor
0.6531
0.8198
0.5631
0.4925
4.0200
0.4594
0.8911
0.9253
0.8096
0.3928
0.6905
3.9025
0.3954
0.9361
Appendix
Table 8: Quantitative results of multiple ablation experiments for VIF and MIF tasks on six datasets.
Figure 12: Visualization of ablation results for different configurations on representative VIF and MIF cases.
Figure 13: Reproducibility verification under different random seeds. For each metric, markers denote independent runs with different initializations, and their tight clustering with small standard deviations indicates that our method is robust to random initialization.
Figure 14: Sensitivity to key hyperparameters.
VIF
MIF
Method
Time (s)
GFLOPs (G)
Params (M)
Method
Time (s)
GFLOPs (G)
Params (M)
CDDFuse
0.006
116.851
1.186
CDDFuse
0.025
116.851
1.186
LRRNet
0.001
0.001
0.049
LRRNet
0.001
0.001
0.049
EMMA
0.058
8.861
1.516
EMMA
0.009
8.861
1.516
TC-MoA
0.543
61.000
340.580
TC-MoA
0.161
61.000
340.580
Text-Difuse
23.818
18513
119.460
Text-Difuse
24.424
2742.5
119.460
Appendix
Table 9: Comparison of runtime, computational cost (GFLOPs), and model size (Params) on VIF and MIF tasks. All values are averaged per image.
Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: github.com/GMY628/RCS-Fusion.
Zeyu Wang, Mingyu Ge, Haiyu Song +1
College of Computer Science and Engineering Dalian Minzu University · Department of Automation Tsinghua University
Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity. Existing methods often struggle with effective representation learning under such conditions, limiting their practical performance. To address these challenges, we propose a mask-guided MMIF method that integrates feature restoration and interaction. We first introduce "Pseudo Ground Truth" to simplify training, promoting faster and more effective feature learning. Then, we design a mask generation mechanism based on the mapping relationship between the fused result and the source images, quantifying the relative contribution of each modality during the fusion process. By incorporating the proposed mask-guided cross-modal cross-attention mechanism, the network is encouraged to selectively attend to informative features during modality interaction, mitigating the risk of overfitting to the static distribution of the "Pseudo Ground Truth". Additionally, we propose a mask-guided learning strategy and a task-coupled degradation-aware learning strategy to balance feature restoration and interaction. Extensive experiments on synthetic and real-world datasets demonstrate that our method surpasses state-of-the-art approaches in visual quality, quantitative metrics, and downstream tasks. The source code is available at https://github.com/ixilai/AMG-Fuse.
Xilai Li, Xiaosong Li, Haishu Tan +3
Foshan University, Foshan, China · China University of Mining and Technology, Beijing, China · Kunming University of Science and Technology, Kunming, China
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros +2
National Technical University of Athens, Greece · Athena Research Center, Greece · University of Bern, Switzerland +2