What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video
Abstract
The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.
Figures & tables
| Stage | Information used | Spatial annotations from the target dataset |
|---|---|---|
| Classifier and readout training | Frames, presence/absence labels, and weights pretrained for ImageNet classification | Not used |
| Patch harvesting and synthesis | Training frames, presence/absence labels, model responses, and pixel statistics | Not used |
| Checkpoint selection | Presence/absence labels for separate selection frames | Not used |
| Threshold calibration | Absence labels and map maxima for separate calibration frames | Not used |
| Data audits, localization/size/SCR evaluation, and diagnostics during development | Presence/absence labels and ground-truth boxes | Used |
| Baselines with bounding-box supervision | Bounding boxes from the dataset specified for each comparison and labels for checkpoint selection | Used as specified for each comparison |
| Count | Train | Select | Calib. | Test |
|---|---|---|---|---|
| CST | ||||
| Sequences | 65 | 34 | 17 | 51 |
| Positive | 15,596 | 7,183 | — | 10,778 |
| Negative | 15,596 | 7,183 | 2,364 | 10,778 |
| A410 | ||||
| Sequences | 27 | 27 | 15 | 26 |
| Method | Training annotations (UAV) | hit16 | hit8 | Pd@FA0.1 | Pd@FA0.5 |
|---|---|---|---|---|---|
| Frame labels + synthesis (main) | CST frame labels | 0.566 / 0.548 | 0.519 / 0.490 | 0.340 / 0.315 | 0.482 / 0.467 |
| Frame labels, no synthesis | CST frame labels | 0.520 / 0.506 | 0.425 / 0.390 | 0.350 / 0.316 | 0.483 / 0.466 |
| A410 frame-label model CST | A410 frame labels | 0.354 / 0.340 | 0.324 / 0.300 | 0.071 / 0.075 | 0.185 / 0.185 |
| Frame differencing | No training | — / 0.551 | — / 0.492 | — / 0.002 | — / 0.006 |
| Frozen main backbone + A410 bbox head | CST frame labels + A410 boxes | 0.559 / 0.539 | 0.553 / 0.533 | 0.443 / 0.429 | 0.553 / 0.537 |
| Frozen main backbone + CST bbox head | CST frame labels + CST boxes | 0.576 / 0.558 | 0.570 / 0.549 | 0.427 / 0.389 | 0.581 / 0.556 |
| Method | hit16r | hit8r | hit16 | hit8 | Pd@FA0.5 r |
|---|---|---|---|---|---|
| A410 frame labels | 0.641 / 0.616 | 0.637 / 0.610 | 0.610 / 0.580 | 0.460 / 0.410 | 0.639 / 0.595 |
| Frozen A410 backbone + A410 bbox head | 0.660 / 0.651 | 0.658 / 0.648 | 0.609 / 0.591 | 0.511 / 0.485 | 0.680 / 0.661 |
| A410 backbone + bbox head, fine-tuned on A410 | 0.695 / 0.677 | 0.689 / 0.672 | 0.645 / 0.625 | 0.511 / 0.485 | 0.690 / 0.674 |
| Frozen A410 backbone + CST bbox head | 0.586 / 0.551 | 0.584 / 0.549 | 0.535 / 0.501 | 0.460 / 0.411 | 0.463 / 0.430 |
| CST frame-label model A410 | 0.209 / 0.213 | 0.201 / 0.203 | 0.188 / 0.190 | 0.167 / 0.154 | 0.159 / 0.159 |
| YOLO26n: A410 boxes | — / 0.687 | — / 0.687 | — / 0.686 | — / 0.668 | — / 0.674 |
| Backbone seed | True labels (3 readouts) | Shuffled labels (9 readouts) |
|---|---|---|
| 0 | 0.6139–0.6145 | 0.0000–0.6066 |
| 1 | 0.6030–0.6060 | 0.0000–0.5861 |
| 2 | 0.6108–0.6120 | 0.0000–0.5747 |
| Metric | No synthesis | With synthesis |
|---|---|---|
| hit16r | 0.639 | 0.640 |
| hit8r | 0.630 | 0.635 |
| hit16 | 0.585 | 0.549 |
| hit8 | 0.400 | 0.388 |
| Pd@FA0.5 r | 0.629 | 0.490 |
| Size (px) | Count | hit16r diff. | 90% CI |
|---|---|---|---|
| 2 to <10 | 54 | +0.000 | [ 0.006, +0.006] |
| 10 to <30 | 918 | +0.061 | [+0.050, +0.072] |
| 30 to <50 | 164 | 0.274 | [ 0.311, 0.232] |
| 782 | 0.014 | [ 0.029, 0.001] |
| Configuration | CST hit16 | CST hit8 | A410 hit16r | A410 hit8r | A410 hit16 | A410 hit8 |
|---|---|---|---|---|---|---|
| ResNet-18, stride 8 | 0.592 | 0.516 | 0.591 | 0.584 | 0.514 | 0.402 |
| ResNet-18, stride 16 | 0.617 | 0.474 | 0.560 | 0.530 | 0.484 | 0.347 |
| ResNet-34, stride 8 | 0.605 | 0.578 | 0.671 | 0.664 | 0.624 | 0.509 |
| ResNet-34, stride 16 | 0.624 | 0.471 | 0.665 | 0.620 | 0.634 | 0.414 |
| MobileNetV3-Small, stride 16 | 0.078 | 0.055 | 0.410 | 0.383 | 0.304 | 0.180 |
| ShuffleNetV2 1.0, stride 8 | 0.136 | 0.044 | 0.413 | 0.318 | 0.286 | 0.133 |
| Model | Params. (M) | GFLOPs | hit16 | hit8 |
|---|---|---|---|---|
| Frame, single | 2.783 | 34.4 | 0.548 | 0.490 |
| Frame, ensemble | 8.349 | 103.3 | 0.566 | 0.519 |
| YOLO26n | 2.375 | 4.2 | 0.687 | 0.685 |
| YOLO26s | 9.466 | 16.6 | 0.715 | 0.713 |
| SCR interval | 2 diagonal < 10 px | 10 diagonal < 30 px | 30 diagonal < 50 px | Diagonal 50 px |
|---|---|---|---|---|
| Q1 (low) | 0.201 (378) | 0.295 (1,784) | 0.475 (282) | 0.278 (18) |
| Q2 | 0.142 (339) | 0.346 (1,939) | 0.449 (243) | 0.006 (177) |
| Q3 | 0.300 (353) | 0.738 (2,635) | 0.784 (134) | 0.000 (4) |
| Q4 (high) | 0.989 (1,132) | 0.926 (1,327) | 0.875 (32) | — |
| All SCR intervals | 0.613 (2,202) | 0.569 (7,685) | 0.544 (691) | 0.030 (199) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Condition | Synthetic positives | Null pastes | |
|---|---|---|---|
| CST, 31,192 | 0.466 | 13,361 | 15,596 |
| CST, 2,756 | 0.611 | 1,100 | 1,378 |
| CST, 930 | 0.792 | 332 | 465 |
| CST, 322 | 0.849 | 86 | 157 |
| CST, 130 | 0.457 | 41 | 64 |
| A410 with A410 patches | 0.974 | 3,950 | 4,611 |
| Scene pool | Train scenes | Train frames + / | Cal. |
|---|---|---|---|
| 8 | 6 | 133 / 167 | 71 |
| 18 | 16 | 285 / 515 | 45 |
| 35 | 32 | 626 / 974 | 89 |
| 65 | 59 | 1,120 / 1,830 | 174 |
| 125 | 113 | 2,174 / 3,476 | 405 |
| 250 | 225 | 4,468 / 6,782 | 857 |