cs.CVSep 28, 2026

Learning to Reason with Persistent Object States for Video Instance Segmentation

Authors: Yongxue Xu, Boxue Yang, Ziqian Liu, Shaoqiu Zhang, Rui Qian, Haopeng Chen

Organizations: Shanghai Jiao Tong University · Sun Yat-sen University · Fudan University

Abstract

Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, plug-and-play framework that explicitly decides when an observation should change an object's state. Each persistent state records identity, confidence, and absence history. A sparse state-observation graph supports Propose-Verify reasoning: provisional associations are revisited using object history, predicted presence, and competition among identities. The verified decisions determine whether to retain, update, reactivate, or suppress each state, while a learned gate controls the evidence written back to memory. Only verified transitions update the persistent state used in subsequent frames. POSReasoner uses standard video annotations and keeps the base model frozen, enabling integration with diverse VOS and VIS architectures. Experiments across long-term VOS and VIS benchmarks show consistent improvements over strong baselines, with the largest gains under occlusion and object reappearance.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

    Jun 5, 2026Ming Dai, Sen Yang, Boqiang Duan +4Video Object SegmentationChain-of-Thought Reasoning

  2. MoVISA: Multi-Token Reasoning for Video Object Segmentation

    Sep 24, 2026Ruining Zhao, Ho Kei Cheng, Alexander G SchwingVideo Object SegmentationMultimodal Large Language Models

  3. QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

    Jul 27, 2026Arian Kheirandish, Fardin Ayar, Ehsan Javanmardi +2Video Object SegmentationLong Videos