SBMVTrack: Spike-Budgeted Multi-View Learning for Power-Efficient UAV Tracking
Authors: Pengzhi Zhong, Jiwei Mo, Haolun Li, Ge Zheng, Jingqi Wang, Xinyi Bo, Shuiwang Li
Organizations: College of Computer Science and Engineering, Guilin University of Technology, Guilin 541004, China · School of Information Engineering, Wuhan University of Technology, Wuhan 430070, China
With sparse and event-driven computation, spiking neural networks show great potential for achieving accurate and power-efficient UAV visual tracking. However, existing SNN-based trackers typically use spike firing rates only for power consumption and lack explicit optimization of actual spike activity. Moreover, regulating spike activity alone does not explicitly encourage stable target representations under partial observations and temporal appearance changes. We propose SBMVTrack, a fully spiking tracking framework that combines spike activity regulation with complementary multi-view representation learning. Specifically, SBMVTrack introduces Energy-Weighted Spike Budgeting (EWSB), which incorporates layer-wise computational costs when regulating spike firing rates and penalizing saturated activations, thereby reducing redundant spike computation. To further improve target representations under the spike budget constraint, we introduce Masked Multi-View Target Modeling (MVTM), which treats the initial template, online template, and search region as temporal views of the same target. By aligning target embeddings between masked and corresponding unmasked views and enforcing cross-view identity consistency, MVTM encourages robustness to missing local cues and temporal appearance changes. Experiments on four UAV benchmarks demonstrate competitive tracking performance with a 24.1% reduction in estimated power consumption relative to the baseline. On VisDrone2018, SBMVTrack achieves a success rate of 70.0%, exceeding SpikeTrack by 9.7 percentage points while reducing estimated power consumption by 45.7%. The source code will be released upon acceptance.
Figures & tables
Figure 1: Power–accuracy trade-off on VisDrone2018. SBMVTrack attains 70.0% AUC at 4.4 mJ inference energy, showing a superior balance between accuracy and efficiency.
Figure 2: Overview of the SBMVTrack architecture. The network comprises an original branch and a masked branch, both processed by a weight-shared spiking backbone. EWSB regulates energy-weighted spike activity, while MVTM performs masked feature reconstruction and cross-view identity-consistency learning through Target-Aware Pooling. During inference, only the original branch, spiking backbone, and SNN tracking head are retained.
Tracker
Source
UAVTrack112
UAVDT
VisDrone2018
UAV123
Power
Param.
Prec.
Succ.
Prec.
Succ.
Prec.
Succ.
Prec.
Succ.
(mJ)
(M)
fDSST ( Danelljan et al. 2016 )
TPAMI 17
56.8
39.1
66.6
38.3
69.8
51.0
58.3
40.5
-
-
MCCT_H ( Wang et al. 2018 )
CVPR 18
63.4
43.6
66.8
40.2
80.3
56.7
65.9
45.7
-
-
ARCF ( Huang et al. 2019b )
ICCV 19
67.3
45.6
72.0
45.8
79.7
58.4
67.1
46.8
-
-
AutoTrack ( Li et al. 2020 )
CVPR 20
69.4
46.5
71.8
45.0
78.8
57.3
68.9
47.2
-
-
HiFT ( Cao et al. 2021 )
ICCV 21
74.2
57.0
65.2
47.5
71.9
52.6
78.7
59.0
33.1
9.9
Table 1: Comparison of SBMVTrack with representative trackers in terms of Precision (Prec.), Success rate (Succ.), Power (mJ), and Parameters (Params.) on UAVTrack112, UAVDT, VisDrone2018, and UAV123. The best and second-best results for each metric are highlighted in bold and underlined, respectively. The compared trackers are categorized into DCF-based, CNN-based, ViT-based, and SNN-based methods.
EWSB
MVTM
UAV123
VisDrone2018
Power
Prec.
Succ.
Prec.
Succ.
(mJ)
86.1
67.0
82.5
65.0
5.8
✓
87.1
67.9
84.2
65.8
4.4
✓
87.4
68.0
84.8
66.7
5.8
✓
✓
87.6
68.0
90.0
70.0
4.4
Table 2: Ablation study of EWSB and MVTM on UAV123 and VisDrone2018.
Figure 3: Comparison of search feature activation maps. The second and third rows show the activation maps produced by SBMVTrack without and with EWSB, respectively.
Method
Prec.
Succ.
Power (mJ)
Baseline
86.1
67.0
5.8
Unweighted Budget
85.5
66.9
4.8
Energy-Weighted Budget
86.5
67.4
5.1
EWSB
87.1
67.9
4.4
Table 3: Ablation study of the key designs in EWSB on UAV123.
rtar
UAV123
UAVDT
Avg. SFR ↓
Power (mJ) ↓
Succ.
Prec.
Succ.
Prec.
Baseline
67.0
86.1
61.2
78.5
0.188
5.80
0.08
68.8
88.7
61.9
79.9
0.131
4.12
0.10
67.4
86.4
61.8
80.2
0.133
4.21
0.12
67.9
87.1
64.1
82.9
0.140
4.42
0.14
66.6
85.3
63.2
80.8
0.146
4.61
Table 4: Impact of rtar on tracking accuracy, average spike firing rate (SFR), and power consumption.
Recon.
Cons.
UAV123
VisDrone2018
Prec.
Succ.
Prec.
Succ.
86.1
67.0
82.5
65.0
✓
84.9
66.4
80.8
63.3
✓
85.3
66.6
84.4
65.5
✓
✓
87.4
68.0
84.8
66.7
Table 5: Ablation study of the key objectives in MVTM on UAV123 and VisDrone2018.
Figure 4: Qualitative evaluation on three video sequences from UAVDT, UAV123, and VisDrone2018 (i.e., S1101, car15, and uav0000074_04320_s).
λcon
UAV123
VisDrone2018
UAVDT
Prec.
Succ.
Prec.
Succ.
Prec.
Succ.
0.1
87.3
67.8
83.7
65.9
77.3
60.4
0.5
87.4
68.0
84.8
66.7
79.7
61.6
1.0
86.9
67.3
84.6
66.3
80.9
62.8
1.5
86.9
67.3
82.7
65.1
78.2
60.5
Table 6: Impact of different values of λcon .
Method
GPU
CPU
Prec.
Succ.
Power (mJ)
OSTrack
65.4
5.9
84.2
64.8
98.9
ORTrack
226.4
55.4
88.6
66.8
11.0
SpikeTrack
60.4
12.1
80.2
60.3
8.1
SBMVTrack
118.8
24.4
90.0
70.0
4.4
SBMVTrack-S
157.0
33.3
85.3
66.5
2.5
Table 7: Inference speed, tracking accuracy, and theoretical power consumption on VisDrone2018.
Multi-object tracking (MOT) plays a fundamental role in visual perception, where accurate trajectory prediction is essential for reliable target association under complex motion patterns. Recent trackers have improved motion modeling with densely activated artificial neural networks, yet they largely overlook whether such dense responses are necessary for trajectory prediction. In this paper, we formulate activation sparsity preference (ASP) by tackling two key questions: 1. How can we identify a model architecture that appropriately and formally explains ASP, and 2. How can we translate this explanation into competitive tracking performance. Theoretical analysis shows that sparse gating is no worse than state-independent dropout under the same activation rate. Based on this insight, SpikingMOT is proposed as a spike-driven tracker that adaptively models sparse trajectory dynamics with spiking neural networks (SNNs). Specifically, SpikingMOT decomposes each trajectory state into pseudo-trajectory bases and uses the current prediction error to calibrate the posterior for next-frame prediction. With this brain-inspired loop, SpikingMOT achieves state-of-the-art performance in extensive experiments, 74.9 HOTA on SportsMOT and 56.5 HOTA on DanceTrack, while reducing the parameters and energy by 72% and 86.7%, respectively. These results bring SNNs into MOT, opening a promising direction for efficient tracking.
Yiding Sun, Xiangyang Yang, Dongxu Zhang +7
1Xi’an Jiaotong University · 2Peking University · 3Guilin University of Technology
Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in complex scenarios. To alleviate this issue, we propose EATrack, an efficient and asymmetric UAV tracking framework centered around a teacher-guided dual-branch distillation strategy that enhances the feature expressiveness of the lightweight student model. Specifically, EATrack investigates two complementary perspectives of knowledge transfer: spatially focused feature-level distillation that compensates for weakened representations by guiding the student to learn strong target representations, and prediction-level distillation that enhances spatial localization by learning the teacher's capability for accurate target localization. Furthermore, to enhance robustness against appearance variations, we introduce a fine-grained target-aware distillation strategy that selectively transfers the teacher's target modeling capacity to the student. A temporal adaptation module is incorporated at inference to enhance robustness over time. Experiments on five UAV benchmarks demonstrate that EATrack achieves a favorable balance between accuracy and speed. Code: https://github.com/GXNU-ZhongLab/EATrack
Hongtao Yang, Bineng Zhong, Qihua Liang +4
Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University, Guilin, 541004, China · Guangxi Key Lab of Multi-Source Information Mining and Security, Guangxi Normal University, Guilin, 541004, China · Nanjing University of Science and Technology +1
Multi-object tracking (MOT) onboard agile unmanned aerial vehicles (UAVs) remains challenging due to severe viewpoint jitter induced by camera ego-motion. Rapid attitude changes during flight often lead to significant target displacement across frames, causing inaccurate target association and degraded tracking performance. Existing UAV MOT methods are primarily evaluated on offline benchmarks and seldom address the practical requirements of real-world onboard deployment, including robustness to camera motion and active target following. To address these challenges, we propose JitTrack, an active onboard multi-object tracking framework that accommodates drone dynamics and camera ego-motion. Built upon a query-based transformer tracker, JitTrack introduces semantic refinement to improve the detection of emerging targets, motion-aware query rectification to compensate for target misalignment caused by viewpoint jitter, and a motion-inspired denoising training strategy that simulates camera motion patterns for robust supervision. Furthermore, we develop a perception-planning-control closed-loop tracking pipeline for real-world deployment, enabling collision-free and physically feasible target following on agile UAVs. Extensive experiments on public UAV MOT benchmarks demonstrate consistent improvements over the baseline method, while real-world flight experiments validate the effectiveness and practicality of JitTrack for robust onboard visual tracking under viewpoint jitter.
Yachun Shan, Feitian Zhang
The Department of Advanced Manufacturing and Robotics, Peking University, Beijing, China