Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt
Figures & tables
Figure 1 : PMT vs. LVMT. Mean AP ± std. dev. over five runs. Across ViT-L/B/S, LVMT improves AP by at least +4.6 over the efficient PMT [ 3 ] baseline at similar FPS. Evaluated on OVIS val .
Figure 2 : LVMT architecture. The Plain Mask Decoder (PMD) takes input queries and projected patch features from a frozen ViT encoder to produce segmentation queries for mask and class prediction. At t=0 , learnable queries Qlrn are used to initialize both the decoder input and the GRU hidden state. Thereafter, a GRU cell adaptively updates the hidden state from the current segmentation queries to yield propagation queries for the next frame. Truncated Query Propagation (TQP), visualized in Fig. 3 , is applied at training time.
Figure 3 : Truncated Query Propagation (TQP). A video of T frames is partitioned into M chunks of F frames, and processed in order in a single iteration. For each chunk, we compute a loss and backpropagate it to calculate the gradients. The optimizer is updated with the average, accumulated gradients after all chunks are consumed. Queries are detached and propagated across chunk boundaries, bounding peak memory to a single chunk and preventing the model from suffering from vanishing gradients caused by long-horizon recurrence.
Method
Step
GRU
TQP
Ttrain
Mean AP ↑
Params ↓
GFLOPs ↓
FPS ↑
PMT [ 3 ]
(0)
×
×
5
51.8±0.3
358M
1014
97
(1)
✓
×
5
54.2±0.4
363M
1015
95
(2)
✓
×
10
52.3±0.3
363M
1015
95
(3)
✓
✓
10
56.1±0.3
363M
1015
95
LVMT (Ours)
(4)
✓
✓
15
56.4±0.4
363M
1015
95
Table 1 : Stepwise modifications from PMT to LVMT on OVIS val [ 31 ] . Mean AP ± standard deviation are reported over five independent runs. Ttrain is the training-clip length in frames.
Method
Backbone
Pre-training
Encoder
OVIS val [ 31 ]
YouTube-VIS 2022 val [ 42 ]
AP
AP 75
AR 10
GFLOPs
FPS
AP L
AP 75L
AR 10L
GFLOPs
FPS
DVIS++ [ 44 ]
ViT-Adapter-L [ 5 ]
DINOv2
49.6
55.0
54.6
868
17
37.5
39.4
43.5
820
18
CAVIS [ 20 ]
ViT-Adapter-L [ 5 ]
DINOv2
53.2
59.1
58.2
863
15
39.5
40.5
44.9
815
15
DVIS-DAQ [ 45 ] †
ViT-Adapter-L [ 5 ]
DINOv2
54.3
60.2
59.8
1173
8
42.0
43.0
48.4
826
10
LOMM [ 19 ]
ViT-Adapter-L [ 5 ]
DINOv2
51.7
57.5
56.2
899
12
48.2
53.2
52.6
842
12
VidEoMT [ 28 ] †
ViT-L [ 12 ]
DINOv2
52.5
57.2
57.5
934
115
42.6
46.1
48.1
557
161
Table 2 : LVMT for VIS on OVIS and YouTube-VIS 2022 [ 31 , 42 ] . † Input resolution of 544 shortest image side for OVIS.
Method
IDF1 (%) ↑
AssA (%) ↑
MT (%) ↑
ML (%) ↓
Total IDS ↓
CAVIS [ 20 ]
79.2
74.7
76.3
6.1
2683
VidEoMT [ 28 ]
78.7
73.3
75.5
6.2
2801
PMT [ 3 ]
78.9
73.3
75.7
6.3
2755
LVMT (Ours)
81.8
78.4
78.8
5.5
2001
Table 3 : Tracking quality on OVIS val [ 31 ] . We compare LVMT with existing methods using specialized tracking metrics.
Table 6 : Effect of memory mechanism on OVIS val [ 31 ] . We compare various memory mechanisms for temporal modeling.
Ttrain
GRU
TQP
AP
AP 75
AR 10
GFLOPs
FPS
5
×
×
52.0
56.0
57.7
1014
97
5
✓
×
54.7
60.6
59.7
1015
95
10
×
×
51.1
56.0
56.7
1014
97
10
✓
×
52.6
56.6
58.4
1015
95
10
×
✓
54.5
59.6
59.9
1014
97
10
✓
✓
56.5
61.5
61.5
1015
95
Table 7 : Effect of GRU and TQP on OVIS val [ 31 ] . We evaluate the effects of GRU-based query propagation and TQP across different training clip lengths Ttrain . All results are measured using 8 NVIDIA H100 (94GB) with batch size 8 (one clip per GPU) and multi-scale resolution (320–640px shortest side, capped at 768).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Backbone
Pre-training
Encoder
YouTube-VIS 2019 val [ 42 ]
YouTube-VIS 2021 val [ 42 ]
AP
AP 75
AR 10
GFLOPs
FPS
AP
AP 75
AR 10
GFLOPs
FPS
DVIS++ [ 44 ]
ViT-Adapter-L [ 5 ]
DINOv2
67.7
75.3
73.7
846
18
62.3
70.2
68.0
830
17
DVIS-DAQ [ 45 ]
ViT-Adapter-L [ 5 ]
DINOv2
68.3
76.1
73.5
851
10
62.4
70.8
68.0
836
10
CAVIS [ 20 ]
ViT-Adapter-L [ 5 ]
DINOv2
68.9
76.2
73.6
838
15
64.6
72.5
69.3
824
15
LOMM [ 19 ]
ViT-Adapter-L [ 5 ]
DINOv2
69.1
76.5
73.5
842
12
65.0
72.7
69.1
842
12
VidEoMT [ 28 ]
ViT-L [ 12 ]
DINOv2
68.6
75.6
73.9
566
160
63.1
69.3
68.1
560
160
Appendix
Table A : LVMT for VIS on YouTube-VIS 2019 and 2021 [ 42 ] .
GT Metric
High-IDS
Low-IDS
Mean objects per video
10
3
Disappear rate † (%)
63
13
Mean disappearance length (frames)
13
6
Appendix
Table B : Analysis of high- and low-IDS videos on OVIS val . High- and low-IDS groups are defined as the 20 videos with the highest and lowest number of identity switches, respectively, measured using the GRU-based model from step (1) in Tab. 1 in the main paper. IDS denotes identity switches. † Fraction of frames where ≥ 1 object is absent due to occlusion or leaving the scene. All statistics are computed from ground-truth (GT) annotations.
Model
GRU
TQP
AP
AP 75
AR 10
GFLOPs
FPS
VidEoMT [ 28 ]
×
×
50.1
53.9
55.8
934
104
LVMT (Ours)
×
×
51.1
56.0
56.7
1014
97
VidEoMT [ 28 ]
✓
×
51.8
56.5
57.3
935
102
LVMT (Ours)
✓
×
52.6
56.6
58.4
1015
95
VidEoMT [ 28 ]
×
✓
52.1
56.7
57.5
934
104
LVMT (Ours)
×
✓
54.5
59.6
59.9
1014
97
Appendix
Table C : Effect of GRU and TQP. Impact of GRU-based propagation and TQP training on VidEoMT [ 28 ] and LVMT on OVIS val with 10-frame training clips.
Method
Size
AP
Params
GFLOPs
FPS
VidEoMT [ 28 ]
L
51.9
316M
934
104
PMT [ 3 ]
52.0
358M
1014
97
LVMT (Ours)
56.7
363M
1015
95
VidEoMT [ 28 ]
B
42.7
93M
304
178
PMT [ 3 ]
42.9
116M
350
160
LVMT (Ours)
48.1
120M
351
155
Appendix
Table D : Impact of model size on OVIS val [ 31 ] . We compare VidEoMT [ 28 ] , PMT [ 3 ] , and LVMT across different model sizes.
Method
Backbone
Encoder
AP
AP 75
AR 10
DINOv2 [ 29 ]
PMT [ 3 ]
ViT-L [ 12 ]
51.8
57.7
56.0
LVMT (Ours)
ViT-L [ 12 ]
55.6
61.6
60.5
PMT [ 3 ]
ViT-L [ 12 ]
53.8
56.0
58.8
LVMT (Ours)
ViT-L [ 12 ]
56.5
61.2
61.7
DINOv3 [ 34 ]
Appendix
Table E : Effect of encoder fine-tuning on OVIS val [ 31 ] . We compare frozen and fine-tuned encoders for PMT [ 3 ] and LVMT using DINOv2 and DINOv3 pre-training.
Figure A : Temporal gradient flow with and without TQP. TQP shortens the backward path and increases within-chunk gradient retention from approximately 5% to 37%.
Figure B : GRU memory retention around object occlusion. Occluded-object queries rely more strongly on memory during the hidden interval.
Chunk size
AP
AP 75
AR 10
GFLOPs
FPS
1
50.9
54.2
55.4
1015
95
3
55.8
60.9
61.0
1015
95
5
56.7
61.8
61.6
1015
95
10
54.3
59.7
60.0
1015
95
Appendix
Table F : Effect of chunk size on OVIS val [ 31 ] . We explore different chunk sizes used by TQP during training.
Hidden state initialization
AP
AP 75
AR 10
Zero Init
55.2
59.6
60.6
Random Gaussian Init
55.8
61.5
60.8
Learnable State
55.7
61.8
60.6
Learnable Object Queries
56.7
61.8
61.6
Appendix
Table G : Effect of hidden-state initialization on OVIS val [ 31 ] . We evaluate different GRU hidden-state initialization strategies.
Figure C : Qualitative results on OVIS [ 31 ] . Comparison between PMT [ 3 ] , LVMT, and the ground-truth annotations on selected frames t={0,4,8,10,11} . PMT suffers from identity switches at t=4 and t=8 under heavy occlusion among similar motorcycles and riders, while LVMT preserves more consistent identities across the sequence.
Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.
While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) existing methods are not designed to support object discovery and identity maintenance in long videos of dynamic ego-motion, and (2) existing evaluation protocols rely on a rigid 1:1 matching that unfairly penalizes semantically valid predictions with mismatched granularity. To address both gaps, we introduce Savvy, a practical and strong system for zero-shot open-world long-horizon video segmentation. Savvy combines hierarchical mask discovery, deferred admission, and track consolidation to support persistent object discovery, safe track promotion, and stable long-range identity maintenance. We further propose OGA, a granularity-aware evaluation suite for open-world video segmentation. Built on a Granularity-Agnostic (GA) matching protocol, OGA relaxes conventional 1:1 matching to an n:1 mapping, but still enforces temporal rigor by detecting support discontinuities through sever points and scoring each reference object through its dominant coherent fragment. This prevents fragmented or flickering support from being over-rewarded while enabling GA-adapted metrics and structural diagnostics: identity persistence (IP), and identity concentration (IC). On VIPSeg, we show that standard 1:1 evaluation substantially underestimates open-world methods, whereas GA evaluation recovers much of their suppressed performance. On the more realistic long-horizon benchmarks: ScanNet and HM3D, Savvy consistently outperforms strong baselines across both classical and proposed metrics, including STQ, VPQ∞, IP and IC. Together, these results establish a practical benchmark and a strong baseline for open-world long-horizon video segmentation.
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the efficiency gains. {In this work, we identify a more fundamental inefficiency underlying this dilemma: while fine-grained visual details are essential for detailed understanding, they are largely redundant for the preliminary task of selecting semantically relevant regions. } Motivated by this, we introduce \textbf{VideoMM}, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity. Specifically, our framework {decouples selection from reasoning} by executing semantic filtering on a cost-effective \textit{Macro Proxy} (derived from downscaled frames), and projecting the selected regions onto high-fidelity \textit{Micro Tokens} for detailed understanding only when necessary. Extensive evaluations show that VideoMM significantly outperforms existing solutions. It achieves a 6.13× speedup and a 7.4% accuracy gain over full-context baselines on LongVideoBench, and further accelerates inference by 2.73× over current leading methods, establishing a highly scalable paradigm for long-video understanding. Our code is available at: https://github.com/adfh917k/VideoMM.
Haoyu Guo, Yuan Feng, Junlin Lv +3
School of Biomedical Engineering, University of Science and Technology of China · Data Darkness Lab, MIRACLE Center, Suzhou Institute for Advanced Research · China +1