MambaVF: State Space Model for Efficient Video Fusion
Organizations: ETH Zürich · Xi’an Jiaotong University · Nanyang Technological University · Tsinghua University
Abstract
Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime. Project page: https://mambavf.github.io
Figures & tables
| VF-Bench Multi-Exposure Fusion Branch (540p) | VF-Bench Multi-Focus Fusion Branch (480p) | ||||||||||||
| VIF | SSIM | MI | Qabf | BiSWE | MS2R | VIF | SSIM | MI | Qabf | BiSWE | MS2R | ||
| CUNet | 0.50 | 0.85 | 1.85 | 0.39 | 7.55 | 0.20 | CUNet | 0.53 | 0.86 | 3.52 | 0.68 | 10.23 | 0.42 |
| HoLoCo | 0.50 | 0.86 | 2.56 | 0.42 | 8.22 | 0.19 | RFL | 0.77 | 0.90 | 6.31 | 0.78 | 8.46 | 0.28 |
| CRMEF | 0.62 | 0.94 | 2.60 | 0.63 | 8.72 | 0.19 | EPT | 0.76 | 0.90 | 6.33 | 0.78 | 8.50 | 0.29 |
| TC-MoA | 0.74 | 0.99 | 2.93 | 0.72 | 7.82 | 0.16 | TC-MoA | 0.75 | 0.90 | 5.27 | 0.77 | 8.39 | 0.28 |
| FILM | 0.77 | 0.99 | 4.35 | 0.72 | 8.28 | 0.17 | FILM | 0.75 | 0.89 | 5.06 | 0.78 | 8.61 | 0.33 |
| Descriptions | Configurations | Metrics | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Scan type | 3D Decoder | Multi-frame | VIF | SSIM | MI | Qabf | BiSWE | MS2R | |
| Exp. I: 2D spatial scan | 2D | ✗ | ✓ | 0.74 | 0.98 | 3.83 | 0.70 | 8.02 | 0.20 |
| Exp. II: 1D temporal scan | 1D | ✗ | ✓ | 0.62 | 0.95 | 2.97 | 0.58 | 8.35 | 0.17 |
| Exp. III: w/o state interaction | 3D | ✗ | ✓ | 0.79 | 0.99 | 4.16 | 0.73 | 7.87 | 0.16 |
| Exp. IV: Single-frame input | 2D | ✗ | ✗ | 0.68 | 0.98 | 3.49 | 0.70 | 7.94 | 0.18 |
| Exp. V: 3D Decoder | 3D | ✓ | ✓ | 0.79 | 0.99 | 3.86 | 0.73 | 7.68 | 0.16 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.