Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of initial keyframes and then relies on propagation. In long or dynamic videos, this can cause stale grounding and persistent false positives when the object composition changes (e.g., distractors enter or the target disappears/re-appears). We propose Event-Driven Refresh + Recurrence Memory (EDRRM), an enhancement that selectively re-invokes Sa2VA only at stable change points. EDRRM triggers refresh boundaries using an EMA-smoothed event score computed from tracking-derived cues (births/deaths and coarse composition/layout changes) with temporal constraints. A recurrence memory further retrieves anchor frames via CLIP similarity to re-condition the model on re-appearance events. Experiments on Ref-DAVIS17, MeViS, and ReVOS show that EDRRM achieves a competitive accuracy-efficiency trade-off relative to fixed-window and FrameDiff-SSIM baselines, maintaining comparable or superior J&F scores at substantially lower average refresh-call budgets and reducing false-positive failures. End-to-end runtime analysis further confirms that the overhead introduced by tracking, CLIP-based recurrence matching, and the identifiability gate remains modest relative to the dominant Sa2VA inference cost, thereby validating the efficiency of the proposed pipeline.
Figures & tables
Fig. 1: Process flow of the proposed pipeline.
Fig. 2: Recurrence memory flowchart.
Parameter
Default
Role
wb ( wb )
1.0
Weight for instance births term in the event score.
wd ( wd )
1.0
Weight for instance deaths term in the event score.
wc ( wc )
0.5
Weight for class-histogram change term in the event score.
wl ( wl )
0.5
Weight for spatial layout-grid change term in the event score.
wr ( wr )
2.0
Weight for recurrence-count term (strong trigger).
Table 1: Frozen default event-score weights (kept fixed across all benchmarks).
Config
thr_track
ema_alpha
event_thr
cooldown
min_chunk
sim_thr
ttl
recurrence_cooldown
pre_ctx
post_ctx
DAVIS [ 5 ]
MEVIS [ 7 ]
REVOS [ 8 ]
edrrm-cfg-1
0.5
0.3
2.5
14
14
0.34
110
25
2
2
76.78
61.50
65.34
edrrm-cfg-2
0.5
0.45
2.0
10
10
0.34
90
25
2
2
76.78
61.89
65.34
edrrm-cfg-3
0.5
0.25
2.8
16
16
0.36
240
50
2
2
76.77
61.43
65.34
edrrm-cfg-4
0.42
0.45
2.0
10
10
0.42
110
35
2
2
76.76
61.08
65.33
edrrm-cfg-5
0.25
0.65
1.4
4
4
0.22
240
8
2
2
76.58
63.13
65.75
edrrm-cfg-6
0.3
0.7
1.6
6
6
0.20
300
10
4
4
76.21
62.43
65.60
Table 2: EDRRM Configuration Ablation
Fig. 3: Event-refresh visualization and interpretability.
Config
threshold
cooldown
min_length
DAVIS [ 5 ]
MEVIS [ 7 ]
REVOS [ 8 ]
ssim-cfg-1
0.95
4
4
76.83
62.48
65.25
ssim-cfg-2
0.93
8
8
77.39
62.24
65.55
ssim-cfg-3
0.90
8
12
76.85
61.76
65.41
ssim-cfg-4
0.87
12
16
76.79
61.59
65.52
ssim-cfg-5
0.82
20
32
76.47
59.85
65.43
ssim-cfg-6
0.80
8
8
77.24
62.35
65.21
Table 3: FrameDiff-SSIM Configuration Ablation
Config
length
DAVIS [ 5 ]
MEVIS [ 7 ]
REVOS [ 8 ]
window-cfg-1
8
77.45
62.34
65.55
window-cfg-2
16
76.53
55.25
65.88
window-cfg-3
32
76.99
59.85
65.40
window-cfg-4
64
75.94
58.92
65.02
window-cfg-5
128
76.31
59.00
65.40
Table 4: Fixed Window Configuration Ablation
Variant
DAVIS [ 5 ]
MeViS [ 7 ]
ReVOS [ 8 ]
J&F
#ref
J&F
#ref
J&F
#ref
Sa2VA Original
75.20
1.00
57.00
1.00
57.60
1.00
EDRRM (Ours)
76.84
3.85
63.13
7.89
66.16
2.22
w/o recurrence memory
75.97
3.53
61.79
5.91
66.14
1.59
w/o identifiability gate
76.80
3.82
63.23
7.89
66.69
2.17
w/o anchor injection
76.24
3.82
62.77
7.22
66.14
2.22
Table 5: Component Ablation. J&F (%) and average #refresh (#ref) for each variant on Ref-DAVIS17, MeViS, and ReVOS under the oracle-tuned best configuration. Each ablated row removes exactly one module from the full EDRRM system.
Fig. 4: Component ablation Pareto frontier. J&F versus average #refresh for all ablation variants on Ref-DAVIS17 (a), MeViS (b), and ReVOS (c). The dashed curve marks the Pareto frontier. EDRRM (Ours) consistently lies on or near the frontier, confirming that the full system achieves the best accuracy-efficiency balance among all variants.
Fig. 5: J&F score versus integer refresh budget (5–30 Sa2VA calls). EDRRM, FrameDiff-SSIM, and Fixed-Window on Ref-DAVIS17, MeViS, and ReVOS. EDRRM maintains higher J&F across a broad budget range, particularly in MeViS, demonstrating that event-driven scheduling outperforms fixed-schedule and appearance-based baselines at matched inference cost.
thr_track
Tracking Quality
DAVIS [ 5 ]
MeViS [ 7 ]
ReVOS [ 8 ]
J&F
#ref
J&F
#ref
J&F
#ref
0.25
Very noisy
76.59
3.25
62.53
6.44
65.96
2.76
0.30
Noisy
76.53
3.21
61.99
6.82
65.57
2.49
0.42
Strict
76.76
1.78
61.21
4.08
65.34
1.02
0.50
Very Strict
76.78
1.25
61.61
2.46
65.34
1.05
Table 6: Tracker Sensitivity Analysis. Effect of thr_track on J&F (%) and average #refresh (#ref). Higher thr_track produces stricter, less noisy tracking with fewer but more reliable events.
Fig. 6: Tracker sensitivity analysis. J&F score (blue, left axis) and average #refresh (red dashed, right axis) versus thr_track on Ref-DAVIS17 (a), MeViS (b), and ReVOS (c). Higher thr_track reduces refresh count but also suppresses valid events, creating a graceful accuracy-cost trade-off.
Dataset
Total Segments
Gate “no”
Gate “no” Rate
DAVIS
934
192
20.56%
MeViS
6224
1383
22.22%
ReVOS
6691
1713
25.60%
Table 7: Identifiability Gate Statistics. Candidate refresh segments suppressed by the gate under the default EDRRM configuration ( edrrm-cfg-5 ).
Fig. 7: Refresh trigger analysis. Left : Proportion of refresh triggers attributed to EMA event score only (blue), recurrence matching only (red), or both simultaneously (purple) per dataset. Right : SSIM value distribution across all consecutive frame pairs in each benchmark, illustrating why SSIM-based detection is threshold-sensitive.
Fig. 8: J&F score versus average #refresh (Sa2VA segmentation calls) for all configurations.
Fig. 9: Budget-conditioned mean J&F versus refresh-rate budget. Note: Budgets where this shared intersection is empty yield undefined means; we therefore report the non-empty range (10–22%) in our benchmarks.
Method
Refresh Rate Budget ↓
Segment Length (Avg) ↑
Segment Length (Min/Max) ↑
Peak budget-conditioned J&F ↑
DAVIS
MEVIS
REVOS
DAVIS
MEVIS
REVOS
DAVIS
MEVIS
REVOS
DAVIS
MEVIS
REVOS
Window
15
14
17
7.58
7.52
6.72
4/8
4/8
3/7
78.78
64.03
66.38
SSIM
10
14
17
8.04
7.83
7.85
4/9
5/8
4/9
78.66
63.78
66.37
Ours
10
10
13
36.37
9.97
13.79
30/44
5/18
11/16
82.60
75.22
67.25
Table 8: Comparison of refresh rate budget, segment length statistics, and J&F scores.
Budget (avg #calls)
Method
DAVIS [ 5 ]
MeViS [ 7 ]
ReVOS [ 8 ]
≤ 2
Window
–
–
67.37
SSIM
93.64
–
64.71
EDRRM
77.55
66.08
67.06
≤ 4
Window
–
63.50
66.03
SSIM
78.66
–
66.55
EDRRM
78.49
58.62
66.30
Table 9: Matched Budget Comparison. J&F (%) evaluated on the shared intersection of video instances where each method uses at most the stated average number of Sa2VA calls. A dash (–) indicates an empty intersection for that dataset at that budget. Note that at budget <=2, the qualifying subset is very small and consists predominantly of static or easy videos, so per-method scores at this level are not indicative of general performance.
Component
DAVIS [ 5 ]
MeViS [ 7 ]
ReVOS [ 8 ]
MASA Tracking
18.73
18.97
8.46
Event Score
0.18
0.64
0.19
CLIP Recurrence
0.17
0.63
0.18
Identifiability Gate
2.74
6.38
1.78
Sa2VA Inference
21.66
24.93
9.26
EDRRM Total
43.48
51.55
19.87
Table 10: End-to-End Runtime / Latency Breakdown (seconds per video). Per-component latency for EDRRM and total end-to-end latency compared against Sa2VA Original, Fixed-window, and FrameDiff-SSIM baselines.
Fig. 10: Qualitative Comparison.
Method/Model
Tuning
DAVIS [ 5 ]
MEVIS [ 7 ]
REVOS [ 8 ]
Mean
VideoLISA-3.8B
-
68.80
44.40
-
56.60
VISA-13B
-
70.40
44.50
50.90
55.27
Sa2VA-1B
-
72.30
50.80
47.60
56.90
Sa2VA-4B
-
73.80
52.10
53.20
59.70
Sa2VA-8B
-
75.20
57.00
57.60
63.27
Sa2VA-26B
-
77.00
57.30
58.40
64.23
Table 11: Method/Model comparison: Global-tuned (single setting) across all datasets.
School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, China · Centre for Vision, Speech and Signal Processing, University of Surrey, Guildford, GU2 7XH, UK
School of Automation, Southeast University, Nanjing, China · Baidu Inc, Beijing, China · College of Information Science & Electronic Engineering, Zhejiang University, Hangzhou, China