Learning to Reason with Persistent Object States for Video Instance Segmentation
Organizations: Shanghai Jiao Tong University · Sun Yat-sen University · Fudan University
Abstract
Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSReasoner, a trainable, plug-and-play framework that explicitly decides when an observation should change an object's state. Each persistent state records identity, confidence, and absence history. A sparse state-observation graph supports Propose-Verify reasoning: provisional associations are revisited using object history, predicted presence, and competition among identities. The verified decisions determine whether to retain, update, reactivate, or suppress each state, while a learned gate controls the evidence written back to memory. Only verified transitions update the persistent state used in subsequent frames. POSReasoner uses standard video annotations and keeps the base model frozen, enabling integration with diverse VOS and VIS architectures. Experiments across long-term VOS and VIS benchmarks show consistent improvements over strong baselines, with the largest gains under occlusion and object reappearance.
Figures & tables
| Backbone | Method | YTVIS 2019 | YTVIS 2021 | YTVIS 2022 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AP | AP50 | AP75 | AP | AP50 | AP75 | AP | AP50 | AP75 | ||
| R50 | GenVIS ( Heo et al., 2023 ) | 50.0 | 71.5 | 54.6 | 47.1 | 67.5 | 51.5 | 37.5 | 61.6 | 41.5 |
| CTVIS ( Ying et al., 2023 ) | 51.8 | 74.1 | 55.2 | 49.5 | 72.5 | 53.6 | 45.4 | 67.5 | 48.9 | |
| + POSReasoner | 52.8 | 75.2 | 56.5 | 50.6 | 73.7 | 55.1 | 47.5 | 69.7 | 51.4 | |
| DVIS++ ( Zhang et al., 2023b ) | 55.7 | 80.8 | 59.8 | 50.4 | 71.4 | 55.2 | 46.3 | 66.6 | 50.5 | |
| + POSReasoner | 56.0 | 80.6 | 60.9 | 51.1 | 72.6 | 55.8 | 47.0 | 68.2 | 51.0 | |
| Method | AP | AP | AP50 | AP75 |
|---|---|---|---|---|
| ResNet-50 | ||||
| CTVIS ( Ying et al., 2023 ) | – | |||
| + POSReasoner | +2.7 | |||
| DVIS++ ( Zhang et al., 2023b ) | – | |||
| + POSReasoner | +2.1 | 38.9 | 39.4 | |
| DAQ ( Zhou et al., 2024 ) | – | 65.2 | ||
| Variant | State | React. | Cross | |||
| LVOS v1 cumulative components | ||||||
| SAM3 host ( Carion et al., 2025 ) | ✗ | ✗ | ✗ | 79.5 | 90.6 | 85.1 |
| + state reasoning | ✗ | ✗ | 82.1 | 92.4 | 87.2 | |
| + reactivation | ✗ | 83.3 | 93.6 | 88.4 | ||
| Full POSReasoner | 83.4 | 93.7 | 88.6 | |||
| LVOS v2 transfer | ||||||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | AP | AP | AP50 | AP75 |
|---|---|---|---|---|
| Frozen host | 34.63 | – | 59.94 | 33.80 |
| Host cues only | 34.82 | +0.18 | 60.21 | 34.09 |
| Shuffled state | 34.67 | +0.04 | 59.95 | 33.94 |
| State, no graph | 36.55 | +1.92 | 62.88 | 35.61 |
| State, one step | 36.13 | +1.50 | 62.01 | 35.27 |
| State, shared two-step | 36.97 | +2.33 | 63.00 | 36.17 |
| Candidates | Params | Latency | Memory |
|---|---|---|---|
| 10 | 0.90M | 5.80 ms | 12.71 MB |
| 20 | 0.90M | 5.66 ms | 12.99 MB |
| 40 | 0.90M | 5.64 ms | 14.03 MB |
| Dataset | Gap | Events | Videos | Host | POSR | [95% CI] | H/T/D |
|---|---|---|---|---|---|---|---|
| LVOS v1 | 5–10 | 45 | 16 | 70.6 | 74.9 | [ , ] | 22/9/14 |
| 11–30 | 34 | 15 | 51.1 | 64.2 | [ , ] | 20/6/8 | |
| 31–100 | 14 | 8 | 72.2 | 78.0 | [ , ] | 7/3/4 | |
| 6 | 4 | 23.6 | 30.6 | [ , ] | 1/2/3 | ||
| LVOS v2 | 5–10 | 114 | 42 | 66.9 | 65.8 | [ , ] | 53/17/44 |
| 11–30 | 83 | 39 | 56.6 | 61.1 | [ , ] | 38/15/30 |
| Method | Backbone | AP | AP50 | AP75 |
|---|---|---|---|---|
| GenVIS ( Heo et al., 2023 ) | ResNet-50 | 35.8 | 60.8 | 36.2 |
| CTVIS ( Ying et al., 2023 ) | ResNet-50 | 35.5 | 60.8 | 34.9 |
| DVIS++ ( Zhang et al., 2023b ) | ResNet-50 | 37.2 | 62.8 | 37.3 |
| DAQ ( Zhou et al., 2024 ) | ResNet-50 | 38.7 | 65.5 | 37.6 |
| CTVIS | ResNet-50 | 34.6 | 59.9 | 33.8 |
| CTVIS + POSReasoner | ResNet-50 | 37.3 | 63.2 | 37.0 |
| Method | Backbone | YTVIS 2019 | YTVIS 2021 | YTVIS 2022 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AP | AP50 | AP75 | AP | AP50 | AP75 | AP | AP50 | AP75 | ||
| Query-based Association | ||||||||||
| MinVIS ( Huang et al., 2022 ) | R50 | 47.4 | 69.0 | 52.1 | 44.2 | 66.0 | 48.1 | 23.3 | 47.9 | 19.3 |
| VITA ∗ ( Heo et al., 2022 ) | R50 | 49.8 | 72.6 | 54.5 | 45.7 | 67.4 | 49.5 | 32.6 | 53.9 | 39.3 |
| GenVIS ( Heo et al., 2023 ) | R50 | 50.0 | 71.5 | 54.6 | 47.1 | 67.5 | 51.5 | 37.5 | 61.6 | 41.5 |
| Memory and Decoupled Tracking | ||||||||||