Organizations: School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology, China. · Mine Digitization Engineering Research Center of the Ministry of Education, China. · University of Auckland, Auckland, New Zealand. · Australian Institute for Machine Learning, Adelaide University, Australia.
RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bounding box annotations. Given only the standard initial-frame annotation used in visual tracking, ESMTrack learns discriminative and temporally consistent representations through two complementary objectives: a grounding triplet loss on annotated initial frames and a cross-frame temporal triplet loss on unlabeled search frames, with reliable samples selected by forward-backward consistency. To address modality dominance bias, ESMTrack employs a three-branch architecture consisting of a fusion branch and two unimodal branches for RGB and thermal inputs. We quantify modality contributions using the Average Peak-to-Correlation Energy by measuring response discrepancies between the fusion and unimodal branches. The resulting reliability estimates guide a training-time modality decoupling mechanism that suppresses dominant-modality shortcuts and adaptively weights cross-modal contrastive learning for task-level alignment. Extensive experiments on five RGB-T tracking benchmarks show that ESMTrack achieves competitive state-of-the-art performance, strong cross-dataset generalization, and real-time inference speed. The source code is available at https://github.com/LiShenglana/ESMTrack.
Figures & tables
Figure 1: (a) Conventional Two-stage Self-Supervised Paradigm. (b) Our End-to-End Box Free Framework where forward tracking and backward tracking adopt the same tracking pipeline
Figure 2: The pipeline of ESMTrack, which learns forward-backward tracking consistency through two complementary triplet constraints: a cross-frame temporal triplet loss for unlabelled search frames with backward IoU filtering to ensure temporal semantic consistency, and a grounding triplet loss for the annotated initial frame to enhance spatial discriminability
Figure 3: Forward tracking framework, where the fusion branch computes modality importance with each unimodal branch, and the Adaptive Modality Decoupling module suppresses the relatively more dominant modality to enhance model robustness
Figure 4: Visualization results of ESMTrack alongside other state-of-the-art methods on the RGBT234 dataset. The top/bottom two rows depict RGB and infrared frames of elecbikechange2 , flower2 and kettle , respectively
Figure 5: Attribute-based evaluation on RGBT234 dataset. The radial axis for PR ranges from 0.4 to 0.9, while SR is scaled from 0.2 to 0.7
Tracker
USOT+RGBT
AFter ∗
ViPT ∗
TBSI ∗
S2OTFormer
GDSTrack
ESMTrack
NO
52.7/38.2
56.2/45.0
72.4 / 56.5
60.3/46.6
67.0/48.8
67.8/52.1
78.0/57.3
PO
20.4/12.6
27.6/24.4
33.5/29.4
25.6/23.2
35.3/26.4
42.6 / 33.0
42.8 / 32.8
TO
15.4/9.6
30.2/25.5
29.4/26.3
23.6/21.2
35.0/25.4
42.2/32.1
41.2 / 31.0
HO
6.2/4.8
18.9/25.7
15.4/21.2
16.4/24.9
20.3/19.1
34.6/36.1
31.6 / 31.7
OV
20.5/16.9
16.9/22.4
55.0 / 51.4
25.2/28.3
30.6/28.7
50.1/41.8
52.4 / 46.7
LI
15.3/10.2
30.8/26.7
30.7/26.9
25.7/23.5
33.3/25.1
36.7/29.2
36.2 / 28.0
Table 2: PR/SR evaluation results on Attribute Challenges comparing to state-of-the-art methods on the LasHeR dataset
Figure 6: Precision Rate and Success Rate on RGBT234 dataset compared against other RGB-T trackers
Figure 7: Precision Rate, Normalized Precision Rate and Success Rate on LasHeR dataset compared against other RGB-T trackers
AMD
Ground. Tri.
Temporal Tri.
GTOT
RGBT234
RGBT210
LasHeR
72.1/59.3
64.1/43.9
63.6/43.4
40.4/36.1/30.2
✓
76.6/63.4
68.2/47.4
65.4/45.1
44.1/39.5/33.9
✓
✓
78.0/63.5
68.7/47.4
67.4 /45.9
46.3/42.1/ 36.0
✓
✓
✓
79.3/65.1
69.1/48.5
66.5/ 46.3
47.0/42.5 /35.1
Table 3: PR, NPR, and SR of the model with different module
Trainable Parameters
GTOT
RGBT234
RGBT210
LasHeR
Task-specific modules only
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
+ LayerNorm & CLS token †
76.8/64.0
70.6/49.6
68.2/47.4
47.3/42.7 / 36.0
Table 4: Ablation on trainable parameter configurations. “ † ” denotes the final adopted setting
Modality Suppression
GTOT
RGBT234
RGBT210
LasHeR
w/o Suppression
72.2/60.0
62.3/43.7
60.8/41.9
42.0/38.6/32.4
Random Suppression
75.7/63.3
63.8/44.9
61.7/42.6
44.0/39.9/33.1
AMD
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 5: PR, NPR, and SR of the model with different modality suppression methods
Reliability Metric
GTOT
RGBT234
RGBT210
LasHeR
Negative Entropy
77.5/64.1
64.8/45.4
63.9/44.6
44.5/40.1/34.1
Max Response
68.3/56.9
67.3/46.7
65.4/45.0
44.6/40.1/34.9
APCE
76.8/64.0
70.6/49.6
68.2/47.4
47.3/42.7 / 36.0
Table 6: Ablation on response-map reliability metrics for the AMD module
Grounding Tri.
Temporal Tri.
GTOT
RGBT234
RGBT210
LasHeR
Cross Sequence
Shifted Box
75.2/62.3
63.2/44.1
62.0/42.8
42.8/37.9/32.8
Cross Sequence
Cross Sequence
80.3/66.2
67.5/47.0
65.3/44.8
44.6/39.8/33.6
Shifted Box
Shifted Box
74.5/61.4
62.5/43.6
58.1/40.3
43.6/39.2/32.6
Shifted Box
Cross Sequence
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 7: PR, NPR, and SR of the model with different negative sampling strategies for triplet loss
Temporal Triplet Loss
GTOT
RGBT234
RGBT210
LasHeR
w/o Filter
74.0/61.2
65.2/46.0
64.4/44.8
44.7/40.5/33.8
Forward APCE Filter
75.8/62.6
67.8/47.5
66.3/45.8
47.2/42.6/35.7
Backward IOU Filter
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 8: PR, NPR, and SR of the model with different pseudo-label filtering strategy
Cross Modal.
Cross View
GTOT
RGBT234
RGBT210
LasHeR
74.4/61.8
61.3/42.8
60.3/41.6
43.0/38.9/32.4
✓
71.2/59.7
63.9/44.6
63.2/43.4
43.2/39.1/32.5
✓
✓
79.3/65.1
69.1/48.5
66.5/46.3
47.0/42.5/35.1
Table 9: PR, NPR, and SR of the model with different Lrgb-tir settings
Figure 8: Visualization of the Adaptive Modality Decoupling (AMD) module on five representative sequences from the LasHeR dataset.
Backward IOU
0.1
0.2
0.3
0.4
GTOT
77.6/63.2
73.6/59.6
79.3/65.1
77.2/62.9
RGBT234
66.5/46.5
64.5/45.3
69.1/48.5
67.9/47.4
LasHeR
44.4/39.3/33.4
42.8/38.3/32.2
47.0/42.5/35.1
44.6/40.0/33.7
Table 10: PR, NPR, and SR on the backward IoU threshold for temporal triplet loss sample filtering
Figure 9: Failure cases on sequence crouch and diamond in RGBT234 dataset. The first row shows RGB images, and the second row shows the corresponding thermal images
Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.
Andong Lu, Ziyi Zha, Jiandong Jin +4
School of Computer Science and Technology, Anhui University, China · School of Artificial Intelligence, Anhui University, China
Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, and usability. Thus, we introduce OneTrackerV2, a unified multi-modal tracking framework that enables end-to-end training for any modality. We propose Meta Merger to embed multi-modal information into a unified space, allowing flexible modality fusion and robustness. We further introduce Dual Mixture-of-Experts (DMoE): T-MoE models spatio-temporal relations for tracking, while M-MoE embeds multi-modal knowledge, disentangling cross-modal dependencies and reducing feature conflicts. With a shared architecture, unified parameters, and a single end-to-end training, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks and 12 benchmarks, while maintaining high inference efficiency. Notably, even after model compression, OneTrackerV2 retains strong performance. Moreover, OneTrackerV2 demonstrates remarkable robustness under modality-missing scenarios.
Lingyi Hong, Jinglun Li, Xinyu Zhou +6
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University
RGB-T detectors leverage the complementary strengths of visible and thermal infrared modalities, achieving robust performance under challenging conditions. Many of them resort to heavy dual backbones and exhaustive cross-modality fusion across the entire image, leading to impractically high computational costs. We observe that most image regions are smooth backgrounds (e.g., sky, ground) that can be easily handled by lightweight single-modality models. In light of this observation, we propose a sparse fusion mechanism for efficient RGB-T detection: first rapidly scanning the image to identify the proposals and then carefully examining the remaining sparse proposals via feature fusion. We propose a two-stage framework to instantiate this mechanism, which performs detection in two stages: 1) a lightweight and modality-specific detection stage that produces high-recall RoIs, and 2) a fusion-driven examination and refinement stage that filters out the false positives and refines the bounding boxes. This design enables the detector to adaptively allocate more computational resources to the potential foregrounds, improving the efficiency while ensuring detection accuracy. Extensive experiments show that our method achieves competitive performance with substantially fewer parameters and lower cost, while maintaining strong scalability to high-resolution images.
Chao Tian, Zikun Zhou, Chao Yang +2
Harbin Institute of Technology, Shenzhen, Guangdong, China