cs.CVSep 29, 2026

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

Authors: Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran

Organizations: Department of Electrical and Computer Engineering University of Victoria 3800 Finnerty Road Victoria, BC, Canada · Department of Mechanical Engineering University of Victoria 3800 Finnerty Road Victoria, BC, Canada

Abstract

Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi-modal tracking where paired modalities are simultaneously available, cross-modal tracking requires localization when template and search frames originate from different active modalities. Accordingly, we introduce TSDA-Track, a Template-Search Domain Adaptation framework to reduce modality discrepancy during training. We investigate two feature alignment strategies. Pre-AFA TSDA-Track applies adversarial alignment before transformer's template-search interaction to suppress modality-specific bias. Enc-CFA TSDA-Track applies contrastive alignment to encoder representations after interaction to strengthen target-level cross-modal correspondence. Both variants retain a shared inference pipeline without modality-specific branches. Experiments on LasHeR, and zero-shot evaluations on RGBT234 and GTOT under multiple cross-modal protocols demonstrate improvements over representative state-of-the-art trackers. For instance, under the modality-switch protocol on RGBT234, Pre-AFA TSDA-Track achieves an SR/PR of 43.2/56.0, compared with 36.8/50.0 for ToMP-101 baseline. In addition, a study on Anti-UAV-024 further verifies the applicability of TSDA-Track to aerial tracking. Our study highlights the effectiveness of feature alignment domain adaptation for cross-modal tracking.

Figures & tables

Explore similar work

Aug 7, 2026cs.CV

AnyTrack: Unifying Visual Object Tracking with Any Modalities

Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. However, existing multi-modal trackers are typically designed for fixed modality combinations, requiring separate models for different inputs. This leads to a poor adaptability to missing or imperfect modalities, and limited generalization. To address these issues, we propose a novel unified framework called AnyTrack for object tracking with any modalities. Specifically, we design a Modality-aware Interaction Module (MIM) to facilitate dynamic interaction across diverse modalities. This module bridges modality discrepancies and aggregates temporal cues to maintain spatio-temporal consistency during cross-modal interaction. Furthermore, we introduce a Context Understanding Module (CUM) to establish spatial correspondence between visual features and target locations via global-local prompts. This module employs target-aware context modeling to enhance foreground-background discrimination for precise localization. Finally, to support the training and evaluation under diverse modalities, we extend existing multi-modal object tracking benchmarks by incorporating grayscale images, language descriptions, and audio clips. Extensive experiments with both complete and missing modality settings demonstrate that our AnyTrack achieves state-of-the-art performance, validating its effectiveness and flexibility. The source code is available at https://github.com/IdolLab/AnyTrack.
May 5, 2026cs.CV

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, and usability. Thus, we introduce OneTrackerV2, a unified multi-modal tracking framework that enables end-to-end training for any modality. We propose Meta Merger to embed multi-modal information into a unified space, allowing flexible modality fusion and robustness. We further introduce Dual Mixture-of-Experts (DMoE): T-MoE models spatio-temporal relations for tracking, while M-MoE embeds multi-modal knowledge, disentangling cross-modal dependencies and reducing feature conflicts. With a shared architecture, unified parameters, and a single end-to-end training, OneTrackerV2 achieves state-of-the-art performance across five RGB and RGB+X tracking tasks and 12 benchmarks, while maintaining high inference efficiency. Notably, even after model compression, OneTrackerV2 retains strong performance. Moreover, OneTrackerV2 demonstrates remarkable robustness under modality-missing scenarios.
Aug 1, 2026cs.CV

Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking

Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context, and have now reached a bottleneck. While high-performance tracking increasingly relies on foundation models, existing methods use them monolithically, adapting a foundation model into a tracker or modify a segmentation foundation model into a tracking pipeline, which fails to exploit complementary strengths. Matching-based trackers excel at instance-level correspondence but lack semantic discrimination and fine-grained foreground perception, whereas segmentation foundation models produce precise masks yet struggle with instance discrimination and multimodal extension. Both paradigms also lack error-correction capabilities for long-term tracking. To address these issues, we propose ACTrack, an agentic coordination framework that treats heterogeneous models as invocable tools under an event-triggered mechanism. ACTrack coordinates a Tracker-based Instance Matching Tool for target discrimination, a SAM3 Motion Tool for mask-derived motion priors, a SAM3 Perception Tool for detecting distractors and instance-conflict cues, and a VLM Reprompt Tool activated only under persistent conflict to mitigate error accumulation. We design a complete tool-invocation trigger mechanism and an inter-tool coordination mechanism, enabling the complementary strengths of different model tools to be fully integrated. Experiments show that ACTrack substantially surpasses the strongest and the largest trackers on eight RGB benchmarks. Furthermore, a parameter-efficient adaptation strategy enables parameter sharing and reuse across tools, achieving unified multimodal tracking with only 30% trainable parameters while substantially outperforming prior methods on multimodal benchmarks such as LasHeR, VisEvent, TNL2K, and DepthTrack.