Many visual tracking methods use rejection mechanisms to suppress unreliable predictions. However, these mechanisms can also reject correctly localized candidates, leaving useful information unused. We investigate how to identify and recover these candidates while preserving native accepted outputs and candidate coordinates. To this end, we propose P-SRM (Post-rejection Selective Recovery Method), which combines spatial responses, past accepted states, and native decision margins to reassess candidates and selectively restore reliable predictions. We evaluate P-SRM on six trackers and four datasets spanning category-specific, point, and generic object tracking. Across all nine configurations, P-SRM improves rejected-candidate ranking and overall tracking performance. These results show that post-rejection verification can identify and recover useful predictions discarded by native rejection, demonstrating the value of reusing rejected information. Project repository: https://github.com/PalestyHR/P-SRM.
Figures & tables
Figure 1: Overview of P-SRM. Rejected candidates are selectively recovered using spatial, historical, and native decision evidence. Dashed arrows indicate training-only paths.
Method / Source
Dataset
Setting
Acc.
Pre.
Rec.
F1
AP r
AJ
OA
NC
nˉC (recovery %)
FPS
Extra latency (ms/frame)
Category-specific tracking
TrackNetV1 AVSS 2019 [ 2 ]
Shuttlecock
Native
52.06
88.99
46.48
61.06
62.83
–
–
125
25.3
7.14
+P-SRM
52.89
88.14
47.92
62.09
73.72
–
–
124.3 (99.47)
21.5
RacketVision
Native
55.40
79.57
59.18
67.88
49.21
–
–
23
25.3
+P-SRM
55.61
78.79
60.06
68.16
54.02
–
–
22.0 (95.65)
21.5
TrackNetV2 ICPAI 2020 [ 3 ]
Shuttlecock
Native
83.62
97.60
82.35
89.33
53.92
–
–
771
39.8
4.27
Table 1: Tracking results (%). RTX 4090 inference (KCF: CPU). NC : correct rejected candidates; nˉC : three-seed mean correct recoveries (recovery %). ‡ : TAP-Net FPS uses clip-amortized latency. Overhead includes clip-to-frame evidence preparation; separate profiling measures 3.94 ms/frame for preparation and 1.08 for core recovery.
Figure 2: Recovery examples across three tracking categories. Gray dashed markers show native-rejected candidates; blue markers show the same candidates after recovery. Examples are from RacketVision, Shuttlecock, Kinetics, and OTB2013.
Comparison
Δt
Δ AP r
Δ F1 pool
Δ F1 video [95% CI]
TAP-Net
Q→Q+H
1.00
3.92
0.020
0.017[−0.046,0.076]
Q+H→Full
0.30
0.11
-0.011
−0.016[−0.024,0.007]
Q+M→Full
1.36
4.04
-0.001
−0.010[−0.070,0.044]
M→Full
6.26
32.89
0.588
0.917[0.620,1.258]
M+H→Full
5.51
20.16
0.508
0.715[0.460,1.008]
Table 2: Cue ablations (pp); CIs: 10,000 paired video bootstraps with fixed models/thresholds. Δt : five-run mean end-to-end change (ms/frame, RTX 4090).