Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory requirements confine them to short video clips and their single-prompt inference design makes multi-category open-vocabulary tracking computationally prohibitive. This work introduces two contributions toward fully automated tracking annotation of arbitrary video. The Generalized Presence Token (GPT) reformulates SAM3's inference pipeline to process N text prompts simultaneously via virtual prompt batching, reducing image encoding cost from O(N) to O(1) with no modifications to any learned component. The Annotation and Segmentation Handler (ASH) extends any memory-attention VIS tracker to sequences of arbitrary length through overlapping temporal chunks with IoU-based inter-chunk identity matching, requiring no dataset-specific training. Instantiated on SAM3, the resulting pipeline -- SAM3-ASH -- achieves state-of-the-art HOTA on MOTS20 under fully zero-shot conditions and remains competitive with trained specialists across seven additional benchmarks, while peak GPU memory consumption stays below 25 GB, establishing a practical baseline for scalable, training-free automated video annotation.
Figures & tables
Figure 1: SAM3-ASH output on MOT20-04 [ 1 ] at t=88 and t=158 ( 2.8 s apart). Mask colours encode track identity: the orange- and peach-outlined pedestrians at the right of each panel retain their identities across the interval. Scenes of this density, with many small and mutually occluding instances per frame, are precisely where manual annotation becomes prohibitive. MOT20 is shown for illustration only; its box ground truth is not commensurable with visible-mask output, so it is not used quantitatively.
Figure 2: Overview of the Annotation and Segmentation Handler (ASH) pipeline. ASH first attempts full-sequence Video Instance Segmentation (VIS) with chunk-level checkpointing. If an Out-of-Memory (OOM) error or length limit is encountered, the pipeline switches to chunk mode. The Chunk Division stage computes overlapping temporal segments and selects optimal chunk boundaries to maximise object density at transitions. Each chunk is then processed by the Chunk Processor, which partitions the full prompt set into batches of size B , runs an independent VIS session per batch, applies a per-batch Identity (ID) offset of (k−1)⋅Ω to prevent collision across batches, and merges the resulting segments. The Inter-Chunk Consistency stage resolves local object IDs into globally consistent track IDs via overlap-region Intersection over Union (IoU) matching; if further chunks remain, the pipeline loops back with updated chunk sequence checkpoints. Once all chunks are processed, Post-Hoc Classification re-queries the VIS model independently for each class name in the dataset vocabulary and assigns semantic labels to tracked objects via IoU matching against per-class reference masks. The final annotated tracked sequence is written to the output.
Method
Pub.
Backbone
Training
YouTube-VIS
OVIS
AP 19
AP 21
AP 22
AP
OV2Seg [ 27 ]
ICCV’23
R50
LVIS
27.2
23.6
—
—
OVFormer [ 36 ]
ECCV’24
R50
LV-VIS
34.8
29.8
—
15.1
SOV [ 37 ]
TOMM’26
R50
LVIS
35.2
31.1
—
—
BriVIS [ 38 ]
AAAI’25
R50
LV-VIS
45.3
39.5
—
14.3
GLEE-Lite [ 3 ]
CVPR’24
R50
Joint (5M+)
53.1
—
—
27.1
Table 1: Video Instance Segmentation results on YouTube-VIS 2019, 2021, and 2022 validation sets and the OVIS validation set . First place , Second place , Third place . Formatting convention used throughout all tables.
Method
Pub.
Backbone
Training
Validation
Test
AP
AP b
AP n
AP
AP b
AP n
InstFormer [ 40 ]
AAAI’25
ViT-B/32
YT-VIS
—
—
12.2
—
—
—
GLEE-Lite [ 3 ]
CVPR’24
R50
Joint (5M+)
19.6
22.1
17.7
—
—
—
OV2Seg [ 27 ]
ICCV’23
SwinB
LVIS
21.1
27.5
16.3
16.4
23.3
11.5
OVFormer [ 36 ]
ECCV’24
R50
LV-VIS
21.9
22.1
21.8
15.2
18.0
13.1
SOV [ 37 ]
TOMM’26
R50
LVIS
22.3
18.4
25.3
18.3
18.5
18.2
Table 2: Open-vocabulary Video Instance Segmentation results on LV-VIS . AP b and AP n denote Average Precision on base (641) and novel (555) categories respectively.
Method
Pub.
HOTA
sMOTSA
DetA
AssA
IDF1
TrackR-CNN [ 28 ]
CVPR’19
—
52.7
—
—
—
TrackFormer [ 9 ]
CVPR’22
—
54.9
—
—
63.6
OmniTracker-L ‡ [ 20 ]
TPAMI’25
—
67.5
—
—
69.2
Seg2Track-SAM2 † [ 43 ]
arXiv’25
61.1
55.9
57.9
66.0
72.2
ReMOTSv2 [ 44 ]
IVC’22
65.4
70.4
71.5
60.9
75.8
EMNT [ 45 ]
TIP’22
66.0
70.0
71.0
62.3
77.0
Table 3: Multi-Object Tracking and Segmentation results on the MOTS20 test set (combined across all sequences). Formatting as in Table 1 .
Method
Pub.
HOTA ↑
AssA ↑
MOTA ↑
IDF1 ↑
Closed-vocabulary specialist trackers
BPMTrack [ 46 ]
TIP’24
—
—
81.3
—
MSPNet [ 18 ]
PR’24
—
—
74.6
71.9
SambaMOTR [ 12 ]
ICLR’25
59.7
59.7
72.9
71.0
C-TWiX [ 19 ]
PR’25
63.1
62.5
78.1
76.3
DiffMOT [ 10 ]
CVPR’24
64.5
64.6
79.8
79.3
Table 4: Results on the MOT17 test set (COMBINED row, averaged over the DPM, FRCNN, and SDP detector variants). Formatting as in Table 1 .
Method
Pub.
HOTA ↑
MOTA ↑
IDF1 ↑
TETA ↑
Closed-vocabulary (8-class)
UNINEXT [ 50 ]
CVPR’23
—
67.1
69.9
—
SUSHI [ 14 ]
CVPR’23
—
68.4
75.6
—
SAM2MOT † [ 22 ]
AAAI’26
—
57.5
70.8
—
SAM2MOT ‡ [ 22 ]
AAAI’26
—
44.1
63.6
—
GHOST [ 11 ]
CVPR’23
61.7
68.1
70.9
—
Table 5: Comparison on the BDD100K val set . Proposed and SAM3/SAM2MOT rows: masks converted to axis-aligned bounding boxes, evaluated with CLEAR, HOTA, and TETA. Closed-vocabulary baselines report published metrics. Formatting as in Table 1 .
Method
Pub.
HOTA ↑
AssA ↑
MOTA ↑
IDF1 ↑
Tracking-by-detection
OC-SORT [ 51 ]
CVPR’23
54.6
40.2
89.6
54.6
AED ∗ [ 52 ]
TIP’25
55.2
57.0
91.0
—
SparseTrack [ 13 ]
CSVT’25
55.5
39.1
91.3
58.3
C-TWiX [ 19 ]
PR’25
62.1
47.2
91.4
63.6
DiffMOT [ 10 ]
CVPR’24
62.3
47.2
92.8
63.0
Table 6: Comparison on the DanceTrack test set . Methods grouped by paradigm. Formatting as in Table 1 .
Figure 3: Qualitative identity consistency across chunk boundaries ( Δ=50 , ω=10 ). Each row shows SAM3-ASH mask overlays around an inter-chunk seam of a DanceTrack test sequence: frames left of the dashed line are produced by chunk ck , frames right of it by chunk ck+1 , with colors keyed to the final merged identities (cf. the IoU merge stage of Fig. 2 ). Despite heavy inter-dancer occlusion at both seams, every identity survives the transition: overlap-window matching associates all objects one-to-one (7/7 on dancetrack0017 , 11/11 on dancetrack0003 ), and the two chunks’ masks agree at the seam frame itself with 0.91–0.995 IoU per object.
Benchmark
Vocab. ( N )
Seq.
Frames
Time (h)
FPS
DanceTrack test
1
35
38,551
4.3
2.51
MOT17 test †
1
2
4,932
0.7
2.06
BDD100K val
8
200
39,973
6.0
1.86
OVIS val
25
140
8,784
1.6
1.51
YT-VIS 2019 val
40
302
8,289
2.9
0.79
YT-VIS 2021 val
40
421
13,195
5.0
0.74
Table 7: SAM3-ASH inference throughput per benchmark on a single 40 GB MIG H100 partition ( Δ=50 , ω=10 , prompt batch size B=5 ). Time covers chunked tracking and post-hoc classification, excluding visualization and disk I/O.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1: GPU memory versus frame index on DanceTrack sequence dancetrack0017 (1,601 frames) on a 40 GB MIG H100 partition. Full-sequence SAM3 accumulates memory-bank state roughly linearly with sequence length and exhausts the partition at frame 993, before completing the sequence. SAM3-ASH ( Δ=50 , ω=10 ) bounds the per-chunk footprint to a constant ≈ 9.3 GiB plateau, independent of sequence length, completing all 1,601 frames.
Figure B.2: GPT prompt-batching sweep on a fixed 12-sequence YT-VIS 2019 val subset (341 frames, full 40-class vocabulary, 40 GB MIG H100). (a) Inference-only throughput versus prompt batch size B , with per-point speedups over the sequential B=1 baseline; the grey curve is the two-term cost model a⌈N/B⌉+c . (b) Peak reserved GPU memory versus B ; all settings, including the fully batched B=40 , fit the 40 GB partition.
Axis
Value
AP
HOTA
DetA
AssA
IDF1
Defaults ( Δ=50 , ω=10 )
46.8
70.6
61.1
82.9
79.1
ω
2
46.2
70.0
60.6
82.4
78.5
5
45.6
69.9
60.6
82.2
78.3
20
46.9
70.4
60.7
83.1
79.1
40 †
45.7
69.3
59.3
82.5
78.0
τov
0.05 / 0.10 ‡
46.8
70.6
61.1
82.9
79.1
Appendix
Table C.1: Sensitivity of SAM3-ASH to ASH hyperparameters on OVIS val (full 25-class vocabulary). Each row varies one parameter from the default configuration; the Defaults and Δ=100 rows are the two configurations reported in the main OVIS comparison.
Benchmark
Instantiation
Prompt
HOTA ↑
MOTA ↑
IDF1 ↑
DanceTrack
FLASH [ 6 ]
box
62.0
64.1
72.5
SAM3-ASH (Ours)
text
66.6
82.4
68.6
MOT17 †
FLASH [ 6 ]
box
43.4
28.7
56.5
SAM3-ASH (Ours)
text
45.2
31.2
55.3
BDD100K
FLASH [ 6 ]
box
58.8
11.1
55.8
SAM3-ASH (Ours)
text
56.9
53.8
69.2
Appendix
Table D.1: ASH instantiated on two backbones with two prompt modalities: box-prompted FLASH (ASH on SAM2 [ 6 ] ) vs. text-prompted SAM3-ASH . Bold marks the better score per benchmark.
Method
Pub.
HOTA ↑
DetA ↑
AssA ↑
MOTA ↑
IDF1 ↑
Closed-vocabulary, Train only
ByteTrack [ 54 ]
ECCV’22
62.8
69.8
51.2
94.1
77.1
OC-SORT [ 51 ]
CVPR’23
71.9
72.2
59.8
94.5
86.4
DiffMOT [ 10 ]
CVPR’24
72.1
72.8
60.5
94.5
86.0
AED ∗ [ 52 ]
TIP’25
72.8
76.8
61.4
95.0
86.3
Deep-EIoU [ 55 ]
WACV’24
74.1
75.0
63.1
95.1
87.2
Appendix
Table E.1: Comparison with state-of-the-art trackers on the SportsMOT test set . Methods are grouped by training data split. Formatting as in Table 1 .