cs.CVSep 30, 2026

Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

Authors: Zeyu Wang, Mingyu Ge, Haiyu Song, Haoran Duan

Organizations: College of Computer Science and Engineering Dalian Minzu University · Department of Automation Tsinghua University

Abstract

Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: github.com/GMY628/RCS-Fusion.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion

    Sep 30, 2026Zeyu Wang, Jiayu Wang, Haiyu Song +1Multimodal FusionVisible Image Fusion

  2. Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction

    Jun 25, 2026Xilai Li, Xiaosong Li, Haishu Tan +3Visible Image FusionQuery-Specific Degradation-Aware Representation

  3. ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

    Jul 26, 2026Guo Yurong, He Yufei, Li Yonghao +3Visible Image FusionInfrared-Visible Image Fusion