Organizations: Department of Computer Science, Bowling Green State University, OH 43402, USA · Department of Computer Science, University of Alabama at Birmingham, AL 35294, USA
In autonomous driving perception, visual object tracking systems must satisfy stringent latency and power constraints while remaining robust in complex and dynamic environments. Although transformer-based trackers achieve state-of-the-art accuracy, their substantial computational and memory overheads hinder deployment on real-time, resource-constrained platforms. To move toward this goal, we propose Difference Feature Map Knowledge Distillation (DFM-KD), a novel relational distillation framework tailored for transformer-based visual object tracking. Unlike conventional feature distillation methods that minimize point-wise discrepancies (e.g., mean squared error) between teacher and student feature representations, DFM-KD transfers knowledge through inter-sample feature differences, explicitly aligning the relational structure of the feature space. By distilling how the teacher models appearance variation and consistency across samples, rather than enforcing similarity in absolute activations, DFM-KD enables the student to better capture the structural dynamics of visual changes within a batch. As a result, the distilled model exhibits enhanced feature robustness and improved tracking performance. Extensive experiments demonstrate that DFM-KD consistently outperforms conventional feature-level distillation methods in both tracking precision and success rates.
Figures & tables
Fig. 1 : General idea of Difference Feature Map Distillation (DFM-KD). Unlike classic instance-wise KD, which aligns teacher and student features for each sample independently, DFM-KD distills the differences between samples to preserve the teacher’s batch-wise topology. In the left panel, single-headed dotted arrows depict conventional per-sample alignment. In the right panel, DFM-KD introduces relational alignment: double-headed arrows represent the alignment of difference vectors between neighboring samples, enabling the student to capture how the teacher’s representations vary across the batch rather than matching isolated feature points.
Fig. 2 : Illustration of the Difference Feature Map based distillation. Note: Grey cubes are selected for knowledge transfer. Dashed yellow cubes that are eliminated do not participate in distillation. The right panel provides a detailed view of how difference feature maps are constructed along the mini-batch dimension.
Method
Params(M)
FLOPs(G)
AUC
P
NP
Teacher
86.60
42.10
75.33
81.10
86.66
MixFormerV2
58.25
28.16
73.21
77.90
84.27
TDSC-KD
73.78
78.39
84.74
DFM-KD
74.03
78.85
85.25
TABLE I : Comparison of methods on MixFormerV2.
Method
Params(M)
FLOPs(G)
AUC
P
NP
Teacher
27.87
5.80
68.63
72.91
79.95
Student-baseline
19.19
3.29
64.20
62.77
72.07
Conventional KD
65.94
67.98
74.59
DFM-KD
68.76
72.17
78.65
TABLE II : Comparison of methods on STARK. FLOPs are computed for the transformer component only.
Method
AUC
Precision
Norm Precision
Baseline
64.20
62.77
72.07
Feat-only
65.94
67.98
74.59
DFM-only
67.81
70.70
77.75
DFM-KD (combined)
68.76
72.17
78.65
TABLE III : Ablation of DFM-KD components on STARK.
Fig. 3 : Qualitative tracking results. The dark blue bounding box marks the target in the initial frame. Across the following three frames, the green bounding box shows the output of our DFM-KD model, whereas the red bounding box corresponds to the student baseline (STARK architecture).
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for robust real-world performance. While recent works incorporate powerful visual foundation model (VFM) features into tracking pipelines, no prior work has systematically analyzed which VFM provides the most robust representations for point tracking. We present the first such analysis, evaluating diverse VFMs in a zero-shot setting on both standard and robustness benchmarks for point tracking. Our study reveals that video diffusion transformers (DiTs) consistently yield the most temporally coherent and discriminative features, even surpassing ResNet backbones explicitly supervised on tracking data. We hypothesize this advantage stem from large-scale video pretraining, full 3D spatio-temporal attention, and a diffusion training objective. Motivated by this finding, we propose DiTracker, which integrates video DiT features into existing tracking frameworks through query-key matching cost computation, cost-level fusion with a lightweight ResNet branch, and LoRA adaptation. Under the same tracking head, DiTracker is trained solely on synthetic data with far fewer iterations, yet outperforms CoTracker3 trained with additional real-world videos, with the largest gains under challenging and corrupted scenarios. It further generalizes across tracking heads and scales with backbone size, confirming that generative video pretraining provides real-world priors that reduce the dependence on large-scale real-data supervision.
Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in complex scenarios. To alleviate this issue, we propose EATrack, an efficient and asymmetric UAV tracking framework centered around a teacher-guided dual-branch distillation strategy that enhances the feature expressiveness of the lightweight student model. Specifically, EATrack investigates two complementary perspectives of knowledge transfer: spatially focused feature-level distillation that compensates for weakened representations by guiding the student to learn strong target representations, and prediction-level distillation that enhances spatial localization by learning the teacher's capability for accurate target localization. Furthermore, to enhance robustness against appearance variations, we introduce a fine-grained target-aware distillation strategy that selectively transfers the teacher's target modeling capacity to the student. A temporal adaptation module is incorporated at inference to enhance robustness over time. Experiments on five UAV benchmarks demonstrate that EATrack achieves a favorable balance between accuracy and speed. Code: https://github.com/GXNU-ZhongLab/EATrack
Hongtao Yang, Bineng Zhong, Qihua Liang +4
Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University, Guilin, 541004, China · Guangxi Key Lab of Multi-Source Information Mining and Security, Guangxi Normal University, Guilin, 541004, China · Nanjing University of Science and Technology +1
Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction. Existing temporal mechanisms typically update templates, prompts, queries, or prediction states, while intermediate representations are rarely reused to modulate corresponding processing stages. We propose \textbf{FeedbackTrack}, a visual-cortex-inspired framework that introduces sparse, group-level layer-aligned cross-frame feedback into pretrained Transformer trackers. Previous-frame intermediate states are detached, cached, and returned to corresponding Transformer groups in the current frame through two lightweight pathways: Query Feedback for token-level query modulation and Gate Feedback for context-dependent feature modulation. FeedbackTrack preserves the original tracking pipeline with only a fixed-size one-frame cache. Across SPMTrack and ARTrackV2, FeedbackTrack consistently improves five backbone configurations on LaSOT and GOT-10k, achieving 83.4 AO and 79.1 AUC with SPMTrack-G while adding less than 1% parameters. Controlled comparisons show that cross-frame feedback outperforms same-frame modulation by 1.8--3.2 AO points, demonstrating that the gains mainly come from recurrent historical information. Further analysis reveals a non-uniform depth-dependent organization of learned feedback strengths, highlighting the effectiveness of recurrent feedback for Transformer tracking.