Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking
Authors: Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran
Organizations: Department of Electrical and Computer Engineering University of Victoria 3800 Finnerty Road Victoria, BC, Canada · Department of Mechanical Engineering University of Victoria 3800 Finnerty Road Victoria, BC, Canada
Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking where paired modalities are simultaneously available, cross-modal tracking requires localization when template and search frames originate from different active modalities. Accordingly, we introduce TSDA-Track, a Template-Search Domain Adaptation framework to reduce modality discrepancy during training. We investigate two feature alignment strategies. Pre-AFA TSDA-Track applies adversarial alignment before transformer's template-search interaction to suppress modality-specific bias. Enc-CFA TSDA-Track applies contrastive alignment to encoder representations after interaction to strengthen target-level cross-modal correspondence. Both variants retain a shared inference pipeline without modality-specific branches. Experiments on LasHeR, and zero-shot evaluations on RGBT234 and GTOT under multiple cross-modal protocols demonstrate improvements over representative state-of-the-art trackers. For instance, under the modality-switch protocol on RGBT234, Pre-AFA TSDA-Track achieves an SR/PR of 43.2/56.0, compared with 36.8/50.0 for ToMP-101 baseline. In addition, a study on Anti-UAV-024 further verifies the applicability of TSDA-Track to aerial tracking. Our study highlights the effectiveness of feature alignment domain adaptation for cross-modal tracking.
Figures & tables
Figure 1: Comparison of (a) single-modal, (b) multi-modal RGB-T, and (c) proposed cross-modal tracking. Switches denote adaptive alignment strategy and interaction stage.
Figure 2: Overview of the proposed TSDA-Track for cross-modal tracking. AFA and CFA denote selectable offline-training alignment branches. Pre-AFA aligns pooled template and search features before the encoder via adversarial learning, whereas Enc-CFA aligns pooled encoder representations through contrastive learning.
Figure 3: PCA visualization of normalized pre-encoder representations on RGBT234 Li et al. (2019) .
Figure 4: Qualitative comparison on RGBT234 Li et al. (2019) under RGB→T, T→RGB, and Switch protocols.
Figure 5: Attribute-based (SR 0.5 , PR) comparison on RGBT234 Li et al. (2019) under the RGB → T, T → RGB, and Switch protocols.
Figure 6: Tracking performance on RGBT234 Li et al. (2019) under the RGB → T, T → RGB, and Switch evaluation protocols.
Figure 7: Visual comparison on Anti-UAV-024 Jiang et al. (2021) test split under RGB → T, T → RGB, and Switch mode.
Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. However, existing multi-modal trackers are typically designed for fixed modality combinations, requiring separate models for different inputs. This leads to a poor adaptability to missing or imperfect modalities, and limited generalization. To address these issues, we propose a novel unified framework called AnyTrack for object tracking with any modalities. Specifically, we design a Modality-aware Interaction Module (MIM) to facilitate dynamic interaction across diverse modalities. This module bridges modality discrepancies and aggregates temporal cues to maintain spatio-temporal consistency during cross-modal interaction. Furthermore, we introduce a Context Understanding Module (CUM) to establish spatial correspondence between visual features and target locations via global-local prompts. This module employs target-aware context modeling to enhance foreground-background discrimination for precise localization. Finally, to support the training and evaluation under diverse modalities, we extend existing multi-modal object tracking benchmarks by incorporating grayscale images, language descriptions, and audio clips. Extensive experiments with both complete and missing modality settings demonstrate that our AnyTrack achieves state-of-the-art performance, validating its effectiveness and flexibility. The source code is available at https://github.com/IdolLab/AnyTrack.
Hao Li, Yunzhi Zhuge, Wenning Hao +4
Army Engineering University of PLA Nanjing, China · Dalian University of Technology Dalian, China · National University of Defense Technology Nanjing, China
Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, and usability. Thus, we introduce OneTrackerV2, a unified multi-modal tracking framework that enables end-to-end training for any modality. We propose Meta Merger to embed multi-modal information into a unified space, allowing flexible modality fusion and robustness. We further introduce Dual Mixture-of-Experts (DMoE): T-MoE models spatio-temporal relations for tracking, while M-MoE embeds multi-modal knowledge, disentangling cross-modal dependencies and reducing feature conflicts. With a shared architecture, unified parameters, and a single end-to-end training, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks and 12 benchmarks, while maintaining high inference efficiency. Notably, even after model compression, OneTrackerV2 retains strong performance. Moreover, OneTrackerV2 demonstrates remarkable robustness under modality-missing scenarios.
Lingyi Hong, Jinglun Li, Xinyu Zhou +6
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University
Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context, and have now reached a bottleneck. While high-performance tracking increasingly relies on foundation models, existing methods use them monolithically, adapting a foundation model into a tracker or modify a segmentation foundation model into a tracking pipeline, which fails to exploit complementary strengths. Matching-based trackers excel at instance-level correspondence but lack semantic discrimination and fine-grained foreground perception, whereas segmentation foundation models produce precise masks yet struggle with instance discrimination and multimodal extension. Both paradigms also lack error-correction capabilities for long-term tracking. To address these issues, we propose ACTrack, an agentic coordination framework that treats heterogeneous models as invocable tools under an event-triggered mechanism. ACTrack coordinates a Tracker-based Instance Matching Tool for target discrimination, a SAM3 Motion Tool for mask-derived motion priors, a SAM3 Perception Tool for detecting distractors and instance-conflict cues, and a VLM Reprompt Tool activated only under persistent conflict to mitigate error accumulation. We design a complete tool-invocation trigger mechanism and an inter-tool coordination mechanism, enabling the complementary strengths of different model tools to be fully integrated. Experiments show that ACTrack substantially surpasses the strongest and the largest trackers on eight RGB benchmarks. Furthermore, a parameter-efficient adaptation strategy enables parameter sharing and reuse across tools, achieving unified multimodal tracking with only 30% trainable parameters while substantially outperforming prior methods on multimodal benchmarks such as LasHeR, VisEvent, TNL2K, and DepthTrack.
Wenrui Cai, Yuzhe Li, Qingjie Liu +1
State Key Laboratory of Virtual Reality Technology and System, Beihang University. · Beihang University. · School of Computer Science and Engineering, Beihang University. +1