cs.CVSep 24, 2026

MoVISA: Multi-Token Reasoning for Video Object Segmentation

Authors: Ruining Zhao, Ho Kei Cheng, Alexander G Schwing

Organizations: University of Illinois Urbana-Champaign

Abstract

Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.

Explore similar work

CardsList
  1. VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

    Jun 5, 2026Ming Dai, Sen Yang, Boqiang Duan +4Video Object SegmentationChain-of-Thought Reasoning

  2. RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

    May 8, 2026Junwei Wen, Deshui Miao, Guangming Lu +2Video Object SegmentationKeyframe Selection

  3. SteerSeg: Attention Steering for Reasoning Video Segmentation

    May 14, 2026Ali Cheraghian, Hamidreza Dastmalchi, Abdelwahed Khamis +3Video Object SegmentationVision-Language Model Grounding