cs.CVFeb 5, 2026

MambaVF: State Space Model for Efficient Video Fusion

Authors: Zixiang Zhao, Yukun Cui, Lilun Deng, Haowen Bai, Haotong Qin, Tao Feng, Konrad Schindler

Organizations: ETH Zürich · Xi’an Jiaotong University · Nanyang Technological University · Tsinghua University

Abstract

Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime. Project page: https://mambavf.github.io

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Apr 2, 2026cs.CV

MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

Infrared and visible video fusion combines the object saliency from infrared images with the texture details from visible images to produce semantically rich fusion results. However, most existing methods are designed for static image fusion and cannot effectively handle frame-to-frame motion in videos. Current video fusion methods improve temporal consistency by introducing interactions across frames, but they often require high computational cost. To mitigate these challenges, we propose MAVFusion, an end-to-end video fusion framework featuring a motion-aware sparse interaction mechanism that enhances efficiency while maintaining superior fusion quality. Specifically, we leverage optical flow to identify dynamic regions in multi-modal sequences, adaptively allocating computationally intensive cross-modal attention to these sparse areas to capture salient transitions and facilitate inter-modal information exchange. For static background regions, a lightweight weak interaction module is employed to maintain structural and appearance integrity. By decoupling the processing of dynamic and static regions, MAVFusion simultaneously preserves temporal consistency and fine-grained details while significantly accelerating inference. Extensive experiments demonstrate that MAVFusion achieves state-of-the-art performance on multiple infrared and visible video benchmarks, achieving a speed of 14.16 FPS at 640×480640 \times 480 resolution. The source code will be available at https://github.com/ixilai/MAVFusion.
May 8, 2026cs.CV

TTF: Temporal Token Fusion for Efficient Video-Language Model

Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at 448×448448{\times}448 resolution already yield >8,000 visual tokens in Qwen3-VL, making LLM prefill the dominant throughput bottleneck. Existing methods often rely on global similarity or attention-guided compression, incurring offsets to their gains. We propose \textbf{Temporal Token Fusion (TTF)}, a training-free, plug-and-play pre-LLM token compression framework that exploits structured temporal redundancy in video. TTF automatically selects an anchor frame, then for each subsequent frame, performs a local window similarity search (e.g.,3×33\times 3), fusing tokens that exceed a threshold. The compressed sequence maintains positional consistency across both prefill and decoding through coordinate realignment, enabling seamless integration with existing VLM pipelines. On Qwen3-VL-8B with threshold t=0.70, TTF removes about 67% of visual tokens while retaining 99.5% of the baseline accuracy and introducing only ≈0.16{\approx}0.16,GFLOPs of matching overhead. Overall, TTF offers a practical, efficient solution for video understanding. The code is available at https://github.com/Cominder/ttf
May 23, 2025cs.CV

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection

The rapid proliferation of surveillance cameras has increased the demand for automated violence detection. While CNNs and Transformers have shown success in extracting spatio-temporal features, they struggle with long-term dependencies and computational efficiency. We propose FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion (GCTF), an efficient architecture combining a dual-branch design and a state-space model (SSM) backbone where one branch captures spatial features, while the other focuses on temporal dynamics. The model performs continuous fusion via a gating mechanism from the spatial branch into the temporal branch to enhance detection of violent activities even in challenging surveillance scenarios. We also present a new benchmark by merging RWF-2000, RLVS, SURV and VioPeru datasets in video violence detection, ensuring strict separation between training and testing sets. Experimental results demonstrate that our model achieves state-of-the-art performance on this benchmark and also on DVD dataset which is a recently introduced dataset on video violence detection, offering an optimal balance between accuracy and computational efficiency, demonstrating the promise of SSMs for scalable, resource efficient video violence detection. The code and pre-trained models are available at https://github.com/damith92/FuseMamba-VD.