Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance
Authors: Khac Duc Giang Nguyen, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
Organizations: University of Amsterdam Amsterdam, The Netherlands · Department of Computer Science and Technology, SUNY Empire State University Saratoga Springs, NY, USA
Counter-UAV systems based on thermal infrared detection must stay accurate as operational datasets evolve, yet sequential fine-tuning causes catastrophic forgetting of prior tasks, a problem that remains insufficiently characterized in this domain. This continual-learning study measures the stability-plasticity trade-off in YOLOMG, a YOLOv5-based detector run as a single thermal-infrared stream with the motion channel disabled, trained sequentially across three anti-UAV benchmarks of rising scale difficulty: Anti-UAV-RGBT, Anti-UAV410, and CST Anti-UAV. Naive fine-tuning on CST yields a Forgetting Measure of -0.605 against the Stage 1 ceiling, corresponding to a 90% capability loss, with -0.572 occurring in Stage 3 alone. In contrast, knowledge distillation from a frozen teacher is associated with FM = -0.033 +/- 0.004 across three seeds, corresponding to 95% retention. Because no Stage 2 no-KD control is included, this result establishes retention under KD training rather than a causal KD effect. Per-stratum analysis shows large-target detection collapsing to near zero within the first epoch, despite an inter-stage cosine similarity of 0.987 over the gradient-updated weights, pointing to scale-conditioned gradient imbalance, rather than weight drift, as a candidate mechanism. Scale-Stratified Herding (SSH), a 300-exemplar buffer balanced across four UAV size strata, roughly halves the forgetting (FM = -0.605 to -0.311) and keeps large-target detection non-zero. An ablation attributes the gain primarily to scale stratification rather than herding: random-stratified replay performs at least as well (FM = -0.221 versus -0.311 for SSH). These replay results are single-seed and should therefore be treated as preliminary.
Figures & tables
Figure 1 . Overview of the three-stage continual-learning curriculum. A single YOLOMG detector, run as a single thermal-infrared stream with the motion channel zeroed, is trained sequentially: Stage 1 supervised on Anti-UAV-RGBT, Stage 2 with knowledge distillation from the frozen Stage 1 teacher on Anti-UAV410, and Stage 3 (naive fine-tuning, Scale-Stratified Herding replay, or the random-stratified control) on CST Anti-UAV. The replay buffer is drawn from the Anti-UAV-RGBT training split (300 exemplars, 75 per size stratum). After each stage, Task-1 retention is evaluated on the Anti-UAV-RGBT validation split to compute the Forgetting Measure, overall and per size stratum.
Figure 2 . Scale distribution across the three training datasets (train + val splits, bbox area: tiny < 256 px 2 , small 256–1024 px 2 , normal 1024–4096 px 2 , large ≥ 4096 px 2 ). Anti-UAV-RGBT is dominated by normal targets (67.8%; 95.8% normal or small). CST Anti-UAV shows a near-complete inversion: 87.9% tiny; 99.8% tiny or small. All three datasets are thermal infrared, so the modality is held constant; the shift studied here is this scale/size distribution, not an arbitrary domain shift.
Stratum
After S1
After S2
After S3
Δ
Tiny
0.009
0.018
0.001
− 0.008
Small
0.579
0.559
0.060
− 0.519
Normal
0.719
0.684
0.079
− 0.640
Large
0.461
0.494
0.000
− 0.461
Overall
0.673
0.640
0.068
− 0.605
Table 1 . Per-stratum T1 mAP@0.5 on Anti-UAV-RGBT across training stages. Stage 3 values are from the best naive checkpoint (epoch 3). After S k is T1 mAP after Stage k ; Δ is the net Stage 1 to Stage 3 change. Knowledge distillation retains every stratum through Stage 2—the large stratum is even slightly strengthened (0.461 to 0.494)—and the collapse occurs only at Stage 3. After S2 is the mean over three Stage 2 seeds (per-stratum std ≤0.006 ); After S1 (the Stage 1 ceiling) and After S3 (the naive run) are single checkpoints. Strata are defined by bounding-box area (thresholds 256/1024/4096 px 2 ).
Figure 3 . Stage 3 naive fine-tuning on CST Anti-UAV (19 epochs). (A) T3 detection mAP@0.5 on CST val; star marks the best checkpoint at epoch 3 (0.083). (B) T1 retention on Anti-UAV-RGBT val; shaded region shows the gap to the Stage 1 ceiling (0.6725). (C) Per-stratum T1 mAP@0.5: the large stratum (green) collapses to 0.000 at epoch 1, indicating scale-specific erasure. (D) Forgetting Measure: FM abs (vs Stage 1 ceiling, dashed) and FM stage3 (vs Stage 2 baseline, solid); worst FM abs=−0.656 at epoch 18.
Metric
Value
T2 mAP@0.5 on Anti-UAV410 (best epoch)
0.428 ± 0.002
T1 mAP@0.5 retained (Anti-UAV-RGBT)
0.640
Forgetting Measure FM
− 0.033 ± 0.004
T1 retention
95%
Best T2 epoch (seed 42)
16 / 32
Table 2 . Stage 2 aggregate results: KD fine-tuning on Anti-UAV410 ( λkd=1.0 , mean ± std across three seeds).
Figure 4 . Stage 2 training dynamics across three seeds (n=3, mean with shaded ± 1 std). (A) T2 mAP@0.5 on Anti-UAV410 val. (B) T1 mAP@0.5 on Anti-UAV-RGBT val; green dashed line is the Stage 1 ceiling (0.6725). (C) F1 score for T1 and T2 tasks. (D) Forgetting Measure per epoch; dashed horizontal marks the − 0.05 boundary.
Stratum
Area range
mAP@0.5 (mean)
Tiny
< 256 px 2
0.001
Small
256–1024 px 2
0.278
Normal
1024–4096 px 2
0.681
Large
≥ 4096 px 2
0.276
Table 3 . Stage 2 scale-stratified mAP@0.5 on Anti-UAV410 validation set (mean across seeds 42, 123, 999). Normal-stratum performance (0.681) is the model’s primary strength; tiny-stratum performance (0.001) is essentially zero throughout all stages.
Transition
Cosine (all)
Cosine (learn.)
FM
Stage 1 to Stage 2
0.911
0.935
− 0.033
Stage 2 to Stage 3
0.967
0.987
− 0.605
Table 4 . Mean cosine similarity between successive stage checkpoints, computed over all 335 named tensors and, separately, over the 204 learnable (gradient-updated) tensors only, with BatchNorm running-statistic buffers excluded. Learnable-only values are for seed 42. Higher value means smaller weight change. FM is the Forgetting Measure relative to the Stage 1 ceiling.
Figure 5 . Relative L2 parameter drift (Stage 2 → Stage 3, naive) per layer group, restricted to the learnable (gradient-updated) tensors. The neck adapts most (median ≈ 0.20) and the backbone least (median ≈ 0.03); the three detection-head tensors show only modest drift (median ≈ 0.05). Excluding the BatchNorm running-statistic buffers removes the spurious head-layer drift seen when all parameters are pooled: the gradient-updated head weights barely move, consistent with starvation of the absent large-target stratum rather than active overwriting of the output layers.
Quantity
CST (0% large)
RGBT (all)
RGBT (large only)
∥g∥ P3 (fine)
0.510
0.000
0.000
∥g∥ P4 (mid)
0.278
0.426
0.000
∥g∥ P5 (coarse)
0.388
0.615
0.227
Lbox
0.096
0.020
0.001
Table 5 . Direct gradient probe at the Stage 2 → 3 boundary. Per-head gradient L2 norm and box-regression loss Lbox (mean over three Stage 2 seeds, 60 batches each, no optimiser step), under the CST distribution, the full T1 distribution, and T1 with only large ( ≥ 4096 px 2 ) ground truth retained. Heads are ordered by anchor size (P3: ∼ 3 px; P5: up to 26×14 px).
Stratum
Eligible frames
Selected
Sampling rate
Tiny
1,091
75
6.9%
Small
43,560
75
0.2%
Normal
98,929
75
0.1%
Large
4,788
75
1.6%
Total
148,368
300
—
Table 6 . SSH buffer composition (Anti-UAV-RGBT training split). Each stratum contributes exactly 75 exemplars regardless of size; the tiny stratum is over-represented (6.9% sampling rate) to protect rare tiny-target knowledge. Strata are by bounding-box area (thresholds 256/1024/4096 px 2 ).
Condition
T3 mAP
T1 mAP
FM
Large T1
Naive (ep. 3)
0.083
0.068
− 0.605
0.000
Random-strat. (ep. 2)
0.064
0.451
− 0.221
0.129
SSH (ep. 2)
0.064
0.362
− 0.311
0.079
Table 7 . Stage 3 forgetting: naive fine-tuning versus random-stratified replay and Scale-Stratified Herding (SSH), each at its best-T3 checkpoint (single seed, seed 42). Both replay variants roughly halve the Forgetting Measure and keep the large stratum non-zero; herding does not improve on random selection.
Figure 6 . Detection output on the same Anti-UAV-RGBT validation frame (sequence 20190925_101846, clip 1_4, frame 0) under three successive checkpoints. Left: Stage 1 best checkpoint (T1 ceiling, mAP@0.5 = 0.6725). Centre: Stage 2 best checkpoint after knowledge distillation (mAP@0.5 = 0.640, FM = − 0.033). Right: Stage 3 naive fine-tuning at epoch 3 (T1 mAP@0.5 = 0.068, FM = − 0.605); the absence of large-target detections is in line with the gradient-starvation interpretation discussed above.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 . Per-split counts of UAV-present frames—frames in which a UAV is present (the exist =1 flag), labelled ‘annotated’ on the figure axis. Anti-UAV410 is the largest (428,703 across train/val/test), followed by Anti-UAV-RGBT (293,209). CST Anti-UAV (208,221) and ARD100 (199,291) are comparable in size; ARD100 is used only for motion-mask pre-computation and is not part of the continual learning curriculum.
Figure 8 . Target visibility per dataset split. A frame is marked visible if it contains an actively flying UAV (the exist =1 flag); occluded or out-of-frame instances are marked absent. Anti-UAV-RGBT has the highest visibility rate, CST Anti-UAV the lowest, consistent with the difficulty of detecting sub-16-pixel targets in cluttered thermal scenes.
Stage 1
Stage 2
Stage 3
Dataset
Anti-UAV-RGBT (IR)
Anti-UAV410
CST Anti-UAV
Objective
Supervised det.
KD + det.
Naive / replay det. †
Epochs
100 (early stop)
50 (early stop)
Naive: 19 epochs, best ep. 3; SSH: 5 epochs, best ep. 2; random: 3 epochs, best ep. 2
Best checkpoint
Epoch 49
Ep. 16/15/27 (seeds 42/123/999)
See Epochs row
Optimiser
SGD (Nesterov)
SGD (Nesterov)
SGD (Nesterov)
Initial LR
10−2
10−3
5×10−4
Appendix
Table 8 . Training configuration per stage. Stage 2 is run with three seeds; Stages 1 and 3 use seed 42 only. Epochs are 0-indexed (epoch 0 is the first).
Epoch
Ldet
Lkd
kd/det
T2 mAP
FM
0
0.0580
0.0791
1.36
0.403
− 0.007
5
0.0355
0.0768
2.16
0.419
− 0.035
10
0.0317
0.0739
2.33
0.406
− 0.046
15
0.0293
0.0712
2.43
0.417
− 0.042
16
0.0289
0.0705
2.44
0.426
− 0.037
20
0.0273
0.0677
2.48
0.419
− 0.041
Appendix
Table 9 . Epoch-level Stage 2 results (seed 42, 32 epochs). kd/det is the ratio of Lkd to Ldet , averaged over batches. Epoch 16 is the peak T2 checkpoint. FM values here are seed 42’s Stage 2 training-time evaluation (e.g. −0.037 at epoch 16). The post-Stage-2 baseline used for the headline Forgetting Measure and as the Stage 3 starting point is a clean re-evaluation of this checkpoint (T1 mAP 0.640 , FM −0.033 ); the ∼ 0.004 gap reflects training-time versus re-evaluation differences. The three-seed mean Stage 2 FM is −0.033±0.004 .
Ep.
T3 mAP
T1 mAP
FM abs
FM stage3
Tiny
Small
Normal
Large
pre
—
0.640
− 0.033
0.000
—
—
—
—
0
0.042
0.206
− 0.466
− 0.434
0.041
0.130
0.251
0.010
1
0.066
0.127
− 0.545
− 0.513
0.005
0.090
0.151
0.000
2
0.074
0.082
− 0.591
− 0.559
0.001
0.075
0.092
0.000
3
0.083
0.068
− 0.605
− 0.572
0.001
0.060
0.079
0.000
4
0.066
0.057
− 0.615
− 0.583
0.000
0.056
0.065
0.000
Appendix
Table 10 . Stage 3 naive condition — per-epoch results. Pre-training T1 mAP (before any Stage 3 gradient steps) = 0.640. T1 ceiling = 0.6725. Bold marks the best T3 checkpoint (epoch 3). Large-stratum T1 mAP reaches 0.000 at epoch 1 and never recovers; the small and normal strata fall steeply in parallel (from 0.559 and 0.684 post-Stage-2, Table 1 , to 0.060 and 0.079 at the best checkpoint).
Dataset
Frame source
Annotation source
Anti-UAV-RGBT
MP4 video (OpenCV)
infrared.json
Anti-UAV410
Pre-extracted JPEG
IR_label.json
CST Anti-UAV
Pre-extracted JPEG
gt.txt + IR_label.json
ARD100
MP4 video (OpenCV)
Pascal VOC XML per frame
Appendix
Table 11 . Native-format loading strategy per dataset. No intermediate converted files are written; format-specific logic is handled inside the custom Dataset implementation.
Figure 9 . Precision plot on CST Anti-UAV validation set (Stage 3 best checkpoint, epoch 3). The representative score at threshold 20 px (PR@20) is reported in the main tracking evaluation table.
Figure 10 . Success plot on CST Anti-UAV validation set (Stage 3 best checkpoint, epoch 3). The area under this curve (AUC) provides an IoU-threshold-independent summary of tracking quality.
Figure 11 . Per-sequence success rate at IoU 0.5 (SR@0.5) on CST Anti-UAV validation set. Sequences with predominantly tiny targets ( < 256 px 2 ) show the lowest SR@0.5, consistent with the near-zero tiny-stratum mAP reported in Table 1 .
Run
Seed(s)
Notes
Stage 1 supervised
42
Best mAP@0.5 = 0.6725 at epoch 49
Stage 2 KD
42, 123, 999
Three-seed retention analysis
Stage 3 naive
42
FM = − 0.605; best T3 checkpoint at epoch 3
Stage 3 SSH replay
42
FM = − 0.311; best T3 checkpoint at epoch 2
Stage 3 random replay
42
FM = − 0.221; best T3 checkpoint at epoch 2
Gradient probe
42, 123, 999
Stage 2 → Stage 3 boundary analysis
Appendix
Table 12 . Experimental runs underlying the reported results.
The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.
Anti-UAV systems increasingly fuse multiple sensors, yet their detection heads provide no per-modality reliability signal. This study evaluates whether evidential deep learning (EDL) heads, Dempster-Shafer (DS) evidence fusion, and uncertainty-driven temporal sensor gating improve anti-UAV detection through a controlled ablation on three benchmarks: thermal tracking (AntiUAV600), RGB-audio-RF classification (TRIDENT), and RGB-IR tracking (MM-UAV). The EDL training objective improves accuracy over retrained sigmoid baselines (+5.9 percentage points in accuracy and a tripled tracker-on-absent rate in E1; +4.8 percentage points in classification accuracy in E2, surviving a clip-clustered bootstrap, p = 0.011) and ranks classification errors substantially better (entropy UAUC approximately 0.94 vs. 0.51). The remaining components do not support their respective hypotheses. DS fusion does not outperform simple probability averaging. Dirichlet vacuity adds no ranking power beyond predictive entropy and inverts at the detection level, where extreme background imbalance causes it to encode class membership rather than error likelihood, a failure also observed for entropy and sigmoid confidence. Temporal gating preserves accuracy only when nearly inactive and yields no realised latency saving on shared-backbone hardware. The benefit of evidential learning therefore arises primarily from its training objective rather than its uncertainty estimate; a crop-level control further localises the detection-level breakdown to anchor-level evaluation rather than the learned representation.
Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
University of Amsterdam Amsterdam, The Netherlands
Unmanned aerial vehicle (UAV) infrared image super-resolution aims to recover weak thermal structures for deployment on resource-constrained platforms; lightweight models are therefore preferred, but multi-loss training can be unstable. A common strategy combines pixel-domain and frequency-domain objectives; however, low contrast, limited high-frequency content, and sensor-specific noise often make their gradients weakly aligned or conflicting. To address this optimization ambiguity, we propose Orthogonal Gradient Gaming and Frequency Rectification (OGG-FR), a plug-and-play optimization framework that decomposes the frequency gradient into a redundant parallel component and an orthogonal innovation component relative to the pixel gradient. In the conflict regime, OGG-FR computes a safe base gradient using the Multiple Gradient Descent Algorithm (MGDA) and adds a variance-rectified orthogonal innovation; in the compatible regime, it discards redundant parallel information and injects the orthogonal innovation according to a confidence score estimated from the high-frequency residual. Experimental results on the UAV thermal benchmark show broad gains under BI and BD degradations at ×4 and ×8 scales, while gradient analyses support the effectiveness of the proposed conflict-aware update rule.
Yongsong Huang, Qingzhong Wang, Xiaofeng Liu +3
Tohoku University · Amazon Web Services · Yale University