Organizations: College of Computer Science and Engineering, Dalian Minzu University, Dalian, China · School of Artificial Intelligence and Robotics, Hunan University, Hunan, China · Department of Automation, Tsinghua University, Beijing, China
Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art (SOTA) performance. Code: github.com/Wangjiayu0512/SIDFusion.
Figures & tables
Figure 1: Existing decomposition paradigms vs Ours. We introduce signal-level decomposition and image-level reconstruction pretext tasks for accurate decomposition and high-quality fusion.
Figure 2: Schematic of the proposed model. It adopts a two-stage self-supervised paradigm. Stage I: feature decomposition via 1D signal-level and 2D image-level pretext tasks; Stage II: fuses the decomposed features to reconstruct the final fused image.
Figure 3: Qualitative comparison of various fusion models.
VIF Task
M 3 FD
MSRS
TNO
Methods
Pub/Year
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
CDDFuse
CVPR 23
0.5741
0.8123
0.4554
0.4746
3.8730
0.4410
0.8698
0.7576
0.8229
0.5399
0.5669
4.9089
0.5106
0.8267
0.4654
0.8090
0.3789
0.4504
3.0691
0.4428
0.7874
LRRNet
TPAMI 23
0.6353
0.8073
0.3709
0.4346
2.8124
0.3550
0.7209
0.6608
0.8082
0.3372
0.3928
2.8662
0.2933
0.5067
0.3686
0.8062
0.2303
0.4831
2.3868
0.3685
0.7011
EMMA
CVPR 24
0.5564
0.8118
0.4602
0.4775
3.7712
0.4690
0.7979
0.6413
0.8161
0.4961
0.5432
4.1637
0.5230
0.7840
0.4356
0.8081
0.3550
0.5112
2.9089
0.4847
0.8072
TC-MoA
CVPR 24
0.5011
0.8096
0.5133
0.4934
3.3330
0.4603
0.8538
0.7446
0.8119
0.4833
0.5558
3.4925
0.4924
0.8825
0.4068
0.8071
0.4020
0.5124
2.5747
0.4312
0.8664
Text-Difuse
NeurIPS 24
0.4262
0.8057
0.4323
0.3398
2.0535
0.1574
0.2566
0.7252
0.8032
0.5378
0.3660
1.3095
0.1570
0.2122
0.2977
0.8047
0.1167
0.4296
1.8773
0.2638
0.4410
Table 1: Quantitative comparison of various fusion models on the VIF and MIF tasks. Best is bold ; second best is underlined .
VIF Task
MIF Task
Task
Case
Configurations
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
Task
Case
Configurations
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
VIF
I
w/o signal-level Dec.
0.6282
0.8181
0.5350
0.4716
3.8631
0.4391
0.8729
MIF
I
w/o signal-level Dec.
0.9012
0.8071
0.3627
0.6715
3.6129
0.3689
0.9127
w/o DWT
0.6425
0.8192
0.5551
0.4860
3.9812
0.4534
0.8847
w/o DWT
0.9185
0.8090
0.3825
0.6880
3.8310
0.3897
0.9312
w/ IDWT ← learnable Rec.
0.6531
0.8198
0.5631
0.4925
4.0200
0.4594
0.8911
w/ IDWT ← learnable Rec.
0.9253
0.8096
0.3928
0.6905
3.9025
0.3954
0.9361
w/ signal- ← image-level Dec.
0.6380
0.8185
0.5460
0.4793
3.9022
0.4468
0.8781
w/ signal- ← image-level Dec.
0.9108
0.8086
0.3755
0.6810
3.7420
0.3798
0.9275
Default setups (ours)
0.6608
0.8209
0.5735
0.4998
4.1050
0.4694
0.8988
Default setups (ours)
0.9369
0.8108
0.4034
0.7008
4.0184
0.4034
0.9439
Table 2: Quantitative results of ablation experiments on VIF and MIF tasks.
Figure 4: Visualization of intermediate 1D signals. The learned common signal cs(l) follows the shared trend of the source signals sA(l) and sB(l) .
Figure 5: Qualitative results on downstream tasks. Left: object detection performance using fused images generated by different fusion methods. Right: medical image segmentation. Our fused images provide more accurate results.
Figure 6: Broader impact of our paradigm. Refined variants (denoted by “*”) obtain fused images with clearer structures and more salient target details, as highlighted by the red and yellow circles.
Figure 7: Qualitative results on multi-focus image fusion, showing that the proposed signal-level decomposition paradigm can be naturally transferred to other image fusion tasks.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Visual Attribute
Pixel-level Manifestation
Commonly Used Constraints
Intensity / Brightness
Absolute pixel magnitude, local luminance
L1 loss, L2 loss, reconstruction loss
Contrast
Relative intensity difference between a pixel and its local neighborhood, or between foreground and background
SSIM contrast term, local contrast loss, entropy-based or information-based constraints
Edge
Sharp local intensity variation around object boundaries and structural transitions
Gradient loss, Sobel loss, Laplacian loss
Texture
fine details, and high-frequency responses
Gradient loss, frequency-domain loss
Boundary
Continuous object outlines and region boundaries
SSIM loss, gradient loss, boundary-aware loss
Structural Layout
Spatial arrangement of objects, organs, roads, targets, or anatomical regions
SSIM loss, perceptual loss
Appendix
Table 3: Representative visual attributes coupled in image pixels and their commonly used constraints in MMIF.
Figure 8: Full qualitative comparisons on six datasets. The top three groups show MIF task and the bottom three groups show VIF cases. In each row, the first two columns are the source image pairs and the remaining columns are fused results of different methods, with our method in the rightmost column, yielding sharper structures and richer complementary details.
Method
People
Car
Bus
Motorcycle
Lamp
Truck
mAP
IR
0.512
0.581
0.447
0.468
0.403
0.452
0.477
VIS
0.563
0.638
0.483
0.521
0.436
0.494
0.522
CDDFuse
0.662
0.726
0.618
0.634
0.573
0.602
0.636
LRRNet
0.641
0.707
0.597
0.612
0.551
0.582
0.615
EMMA
0.651
0.719
0.612
0.631
0.562
0.593
0.628
TC - MoA
0.642
0.715
0.606
0.623
0.559
0.601
0.624
Appendix
Table 4: AP for object detection on source images and fused images from different methods. mAP is the mean of the six category APs.
Metrics
T1
Flair
CDDFuse
LRRNet
EMMA
TC - MoA
Text - Diffuse
IoU
0.615
0.657
0.762
0.731
0.743
0.689
0.685
Dice
0.607
0.727
0.850
0.836
0.841
0.829
0.751
Metrics
CCF
SigFusion
BSAFusion
Mask - Difuser
MTG - Fusion
C2RF
Ours
IoU
0.661
0.735
0.745
0.723
0.775
0.772
0.790
Dice
0.752
0.770
0.832
0.829
0.846
0.848
0.856
Appendix
Table 5: Quantitative comparison for medical image segmentation.
Figure 9: Visualization of common and unique feature maps produced by different decomposition strategies. For each source pair, we display the estimated common features and the modality-unique features. Compared with existing methods that generate two inconsistent common features, our signal-level decomposer yields a single common map and better-separated unique maps.
Figure 10: t-SNE visualization of low-/high-frequency components and learned decomposed features. The common features are close to the low-frequency components of both modalities, while the unique features are closer to their corresponding high-frequency components, supporting the design of extracting common information from low frequencies and unique information from high frequencies.
Figure 11: Visualization of training curves. We compare the proposed signal-level objective with a conventional 2D image-domain mixed-loss strategy in terms of training loss and feature similarity, including SSIM, Pearson correlation, and cosine similarity. The results show that our signal-level objective converges faster and achieves a lower final loss, while maintaining higher and more stable similarity between the learned common features and source signals.
Method
QMI↑
QNICE↑
QP↑
QCB↑
MI↑
VIFp↑
QY↑
DIDFuse
0.6130
0.8083
0.3570
0.4830
3.2370
0.3610
0.9020
DIDFuse*
0.7466
0.8196
0.4198
0.5299
3.9038
0.4596
0.9101
Gain
21.8% ↑
1.4% ↑
17.6% ↑
9.7% ↑
20.6% ↑
27.3% ↑
0.9% ↑
CDDFuse
0.6610
0.8107
0.4010
0.6940
3.9820
0.3990
0.9380
CDDFuse*
0.7582
0.8277
0.4523
0.7051
4.3563
0.4736
0.9446
Gain
14.7% ↑
2.1% ↑
12.8% ↑
1.6% ↑
9.4% ↑
18.7% ↑
0.7% ↑
Appendix
Table 6: The refinement effects results of our paradigm. Original vs. upgraded models (* denotes our paradigm inserted).
Method
Pub/Year
Dataset: LYTRO
Dataset: MFFW
QG↑
QM↑
QP↑
MI↑
SD↑
VIFF↑
QG↑
QM↑
QP↑
MI↑
SD↑
VIFF↑
CUNet
TPAMI 20
0.526
0.553
0.696
5.441
58.70
1.022
0.482
0.455
0.552
4.593
56.33
0.847
U2Fusion
TPAMI 20
0.580
0.480
0.742
5.677
58.37
1.086
0.537
0.405
0.611
4.876
55.26
0.825
DeFusion
ECCV 22
0.455
0.325
0.660
5.984
54.39
1.028
0.418
0.296
0.518
5.137
51.55
0.876
DIFNet
CVPR 22
0.437
0.325
0.688
5.774
49.67
1.032
0.422
0.309
0.577
4.867
46.66
0.890
FusionDiff
ESWA 23
0.629
0.821
0.783
6.554
56.13
1.188
0.545
0.602
0.659
5.334
53.27
0.993
Appendix
Table 7: Quantitative comparison on the MFIF task against 9 competing methods using 6 evaluation metrics on the LYTRO and MFFW datasets. Although MFIF may provide fused ground truth for evaluation, the decomposition of common and unique features still lacks direct supervision. Our method remains applicable in this setting and achieves strong overall performance.
I. Effect of Signal-Level Decomposer
VIF task
MIF task
Configuration settings
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
QMI ↑
QNICE ↑
QP ↑
QCB ↑
MI ↑
VIFp ↑
QY ↑
VIF: M 3 FD
MIF: MRI-CT
w/o signal-level decomposer
0.6282
0.8181
0.5350
0.4716
3.8631
0.4391
0.8729
0.9012
0.8071
0.3627
0.6715
3.6129
0.3689
0.9127
w/o DWT
0.6425
0.8192
0.5551
0.4860
3.9812
0.4534
0.8847
0.9185
0.8090
0.3825
0.6880
3.8310
0.3897
0.9312
w/ IDWT ← learnable reconstructor
0.6531
0.8198
0.5631
0.4925
4.0200
0.4594
0.8911
0.9253
0.8096
0.3928
0.6905
3.9025
0.3954
0.9361
Appendix
Table 8: Quantitative results of multiple ablation experiments for VIF and MIF tasks on six datasets.
Figure 12: Visualization of ablation results for different configurations on representative VIF and MIF cases.
Figure 13: Reproducibility verification under different random seeds. For each metric, markers denote independent runs with different initializations, and their tight clustering with small standard deviations indicates that our method is robust to random initialization.
Figure 14: Sensitivity to key hyperparameters.
VIF
MIF
Method
Time (s)
GFLOPs (G)
Params (M)
Method
Time (s)
GFLOPs (G)
Params (M)
CDDFuse
0.006
116.851
1.186
CDDFuse
0.025
116.851
1.186
LRRNet
0.001
0.001
0.049
LRRNet
0.001
0.001
0.049
EMMA
0.058
8.861
1.516
EMMA
0.009
8.861
1.516
TC-MoA
0.543
61.000
340.580
TC-MoA
0.161
61.000
340.580
Text-Difuse
23.818
18513
119.460
Text-Difuse
24.424
2742.5
119.460
Appendix
Table 9: Comparison of runtime, computational cost (GFLOPs), and model size (Params) on VIF and MIF tasks. All values are averaged per image.
Foshan University, Foshan, China · China University of Mining and Technology, Beijing, China · Kunming University of Science and Technology, Kunming, China