cs.CVFeb 5, 2026

MambaVF: State Space Model for Efficient Video Fusion

Authors: Zixiang Zhao, Yukun Cui, Lilun Deng, Haowen Bai, Haotong Qin, Tao Feng, Konrad Schindler

Organizations: ETH Zürich · Xi’an Jiaotong University · Nanyang Technological University · Tsinghua University

Abstract

Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime. Project page: https://mambavf.github.io

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

    Apr 2, 2026Xilai Li, Weijun Jiang, Xiaosong Li +5Visible Image FusionOptical Flow

  2. TTF: Temporal Token Fusion for Efficient Video-Language Model

    May 8, 2026Simin Huo, Ning LIVideo-Language ModelsQwen3

  3. FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection

    May 23, 2025Damith Chamalke Senadeera, Muhammad Awais, Shibo Li +2Fine-Grained Video UnderstandingViolence Detection