cs.CVSep 28, 2026

LVMT: Video Mask Transformer for Long-term Video Segmentation

Authors: Narges Norouzi, Niccolò Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus

Organizations: Eindhoven University of Technology · RWTH Aachen University

Abstract

Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoVISA: Multi-Token Reasoning for Video Object Segmentation

    Sep 24, 2026Ruining Zhao, Ho Kei Cheng, Alexander G SchwingVideo Object SegmentationMultimodal Large Language Models

  2. Open-World Video Segmentation

    Jun 14, 2026Qing Su, Kaiyang Li, Yuan Zhuang +2Video Object SegmentationDiscovery

  3. VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

    Sep 15, 2026Haoyu Guo, Yuan Feng, Junlin Lv +3Video Multimodal Large Language ModelsMultimodal Large Language Models