cs.CVSep 30, 2026

Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation

Authors: Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo

Organizations: School of Electrical Engineering, Korea Advanced Institute of Science & Technology (KAIST), Daejeon 34141, Republic of Korea

Abstract

Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of initial keyframes and then relies on propagation. In long or dynamic videos, this can cause stale grounding and persistent false positives when the object composition changes (e.g., distractors enter or the target disappears/re-appears). We propose Event-Driven Refresh + Recurrence Memory (EDRRM), an enhancement that selectively re-invokes Sa2VA only at stable change points. EDRRM triggers refresh boundaries using an EMA-smoothed event score computed from tracking-derived cues (births/deaths and coarse composition/layout changes) with temporal constraints. A recurrence memory further retrieves anchor frames via CLIP similarity to re-condition the model on re-appearance events. Experiments on Ref-DAVIS17, MeViS, and ReVOS show that EDRRM achieves a competitive accuracy-efficiency trade-off relative to fixed-window and FrameDiff-SSIM baselines, maintaining comparable or superior J&F scores at substantially lower average refresh-call budgets and reducing false-positive failures. End-to-end runtime analysis further confirms that the overhead introduced by tracking, CLIP-based recurrence matching, and the identifiability gate remains modest relative to the dominant Sa2VA inference cost, thereby validating the efficiency of the proposed pipeline.

Figures & tables

Explore similar work

CardsList
  1. ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

    Jul 27, 2026Yuanjia Li, Tianyang Xu, Tao Zhou +3Video Object SegmentationSpatial Grounding

  2. AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method

    Apr 20, 2026Deshui Miao, Chao Yang, Chao Tian +4Video Object SegmentationNeural Mask Estimation

  3. VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

    Jun 5, 2026Ming Dai, Sen Yang, Boqiang Duan +4Video Object SegmentationChain-of-Thought Reasoning