Organizations: School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology, China. · Mine Digitization Engineering Research Center of the Ministry of Education, China. · University of Auckland, Auckland, New Zealand. · Australian Institute for Machine Learning, Adelaide University, Australia.
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.
Figures & tables
Figure 1: (a) Conventional Two-stage Self-Supervised Paradigm. (b) Our End-to-End Box Free Framework where forward tracking and backward tracking adopt the same tracking pipeline
Figure 2: The pipeline of ESMTrack, which learns forward-backward tracking consistency through two complementary triplet constraints: a cross-frame temporal triplet loss for unlabelled search frames with backward IoU filtering to ensure temporal semantic consistency, and a grounding triplet loss for the annotated initial frame to enhance spatial discriminability
Figure 3: Forward tracking framework, where the fusion branch computes modality importance with each unimodal branch, and the Adaptive Modality Decoupling module suppresses the relatively more dominant modality to enhance model robustness
Figure 4: Visualization results of ESMTrack alongside other state-of-the-art methods on the RGBT234 dataset. The top/bottom two rows depict RGB and infrared frames of elecbikechange2 , flower2 and kettle , respectively
Figure 5: Attribute-based evaluation on RGBT234 dataset. The radial axis for PR ranges from 0.4 to 0.9, while SR is scaled from 0.2 to 0.7
Tracker
USOT+RGBT
AFter ∗
ViPT ∗
TBSI ∗
S2OTFormer
GDSTrack
ESMTrack
NO
52.7/38.2
56.2/45.0
72.4 / 56.5
60.3/46.6
67.0/48.8
67.8/52.1
78.0/57.3
PO
20.4/12.6
27.6/24.4
33.5/29.4
25.6/23.2
35.3/26.4
42.6 / 33.0
42.8 / 32.8
TO
15.4/9.6
30.2/25.5
29.4/26.3
23.6/21.2
35.0/25.4
42.2/32.1
41.2 / 31.0
HO
6.2/4.8
18.9/25.7
15.4/21.2
16.4/24.9
20.3/19.1
34.6/36.1
31.6 / 31.7
OV
20.5/16.9
16.9/22.4
55.0 / 51.4
25.2/28.3
30.6/28.7
50.1/41.8
52.4 / 46.7
LI
15.3/10.2
30.8/26.7
30.7/26.9
25.7/23.5
33.3/25.1
36.7/29.2
36.2 / 28.0
Table 2: PR/SR evaluation results on Attribute Challenges comparing to state-of-the-art methods on the LasHeR dataset
Figure 6: Precision Rate and Success Rate on RGBT234 dataset compared against other RGB-T trackers
Figure 7: Precision Rate, Normalized Precision Rate and Success Rate on LasHeR dataset compared against other RGB-T trackers
AMD
Ground. Tri.
Temporal Tri.
GTOT
RGBT234
RGBT210
LasHeR
72.1/59.3
64.1/43.9
63.6/43.4
40.4/36.1/30.2
✓
76.6/63.4
68.2/47.4
65.4/45.1
44.1/39.5/33.9
✓
✓
78.0/63.5
68.7/47.4
67.4 /45.9
46.3/42.1/ 36.0
✓
✓
✓
79.3/65.1
69.1/48.5
66.5/ 46.3
47.0/42.5 /35.1
Table 3: PR, NPR, and SR of the model with different module
Trainable Parameters
GTOT
RGBT234
RGBT210
LasHeR
Task-specific modules only
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
+ LayerNorm & CLS token †
76.8/64.0
70.6/49.6
68.2/47.4
47.3/42.7 / 36.0
Table 4: Ablation on trainable parameter configurations. “ † ” denotes the final adopted setting
Modality Suppression
GTOT
RGBT234
RGBT210
LasHeR
w/o Suppression
72.2/60.0
62.3/43.7
60.8/41.9
42.0/38.6/32.4
Random Suppression
75.7/63.3
63.8/44.9
61.7/42.6
44.0/39.9/33.1
AMD
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 5: PR, NPR, and SR of the model with different modality suppression methods
Reliability Metric
GTOT
RGBT234
RGBT210
LasHeR
Negative Entropy
77.5/64.1
64.8/45.4
63.9/44.6
44.5/40.1/34.1
Max Response
68.3/56.9
67.3/46.7
65.4/45.0
44.6/40.1/34.9
APCE
76.8/64.0
70.6/49.6
68.2/47.4
47.3/42.7 / 36.0
Table 6: Ablation on response-map reliability metrics for the AMD module
Grounding Tri.
Temporal Tri.
GTOT
RGBT234
RGBT210
LasHeR
Cross Sequence
Shifted Box
75.2/62.3
63.2/44.1
62.0/42.8
42.8/37.9/32.8
Cross Sequence
Cross Sequence
80.3/66.2
67.5/47.0
65.3/44.8
44.6/39.8/33.6
Shifted Box
Shifted Box
74.5/61.4
62.5/43.6
58.1/40.3
43.6/39.2/32.6
Shifted Box
Cross Sequence
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 7: PR, NPR, and SR of the model with different negative sampling strategies for triplet loss
Temporal Triplet Loss
GTOT
RGBT234
RGBT210
LasHeR
w/o Filter
74.0/61.2
65.2/46.0
64.4/44.8
44.7/40.5/33.8
Forward APCE Filter
75.8/62.6
67.8/47.5
66.3/45.8
47.2/42.6/35.7
Backward IOU Filter
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 8: PR, NPR, and SR of the model with different pseudo-label filtering strategy
Cross Modal.
Cross View
GTOT
RGBT234
RGBT210
LasHeR
74.4/61.8
61.3/42.8
60.3/41.6
43.0/38.9/32.4
✓
71.2/59.7
63.9/44.6
63.2/43.4
43.2/39.1/32.5
✓
✓
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 9: PR, NPR, and SR of the model with different Lrgb-tir settings
Figure 8: Visualization of the Adaptive Modality Decoupling (AMD) module on five representative sequences from the LasHeR dataset.
Backward IOU
0.1
0.2
0.3
0.4
GTOT
77.6/63.2
73.6/59.6
79.3/65.1
77.2/62.9
RGBT234
66.5/46.5
64.5/45.3
69.1/48.5
67.9/47.4
LasHeR
44.4/39.3/33.4
42.8/38.3/32.2
47.0/42.5/35.1
44.6/40.0/33.7
Table 10: PR, NPR, and SR on the backward IoU threshold for temporal triplet loss sample filtering
Figure 9: Failure cases on sequence crouch and diamond in RGBT234 dataset. The first row shows RGB images, and the second row shows the corresponding thermal images