Bee Detection and Tracking at Hive Entrance using YOLO11 and ByteTrack
Authors: Thi Thu Thao Nguyen, Johannes Reschke
Organizations: IoT Engineering Savonia University of Applied Sciences Finland · Electrical and Information Engineering Ostbayerische Technische Hochschule Regensburg Germany
This work presents an automatic bee entrance monitoring system based on YOLO11 transfer learning and the ByteTrack tracking algorithm. The study investigates the influence of data augmentation, backbone freezing, and tracker parameter optimization on the detection and counting of small, fast-moving bees. The detector with progressive backbone unfreezing strategy achieved about 97.0% precision and 98.7% mAP50, while providing more stable convergence than full fine-tuning. Experiments also showed that light augmentation outperformed heavy augmentation. For tracking, ByteTrack parameters were optimized to improve trajectory continuity under low-confidence detections. On an independent 25 FPS side-view video, the optimized YOLO11-ByteTrack system correctly counted 43 of 47 incoming bees (91.5%) and 7 of 30 outgoing bees (23.3%). Error analysis showed that most counting errors were caused by missed detections due to rapid bee motion and motion blur, while tracking failures became less frequent after parameter optimization. Overall, the results indicate that moderate augmentation, progressive backbone unfreezing, and ByteTrack tuning improve the reliability of automatic bee entrance monitoring under realistic recording conditions.
Figures & tables
Fig. 1: Example image of bees at a hive entrance, adapted from the Mendeley dataset [ 5 ] .
Fig. 2: Example image of bees at a hive entrance, adapted from the Mendeley dataset [ 5 ] .
Fig. 3: Example image of bees at a hive entrance, adapted from Dataset Ninja [ 6 ] .
Fig. 4: A frame extracted from the evaluation video by Ella Gronewold on Pexels [ 7 ] .
Fig. 5: ByteTrack pipeline.
Fig. 6: Counting logic graph.
Fig. 7: Box loss and class loss between light and heavy augmented models.
Fig. 8: mAP50 and mAP50–95 between light and heavy augmented models.
Fig. 9: Recall between light and heavy augmented models.
Model
Precision
Recall
Unfreeze all
∼ 0.969–0.973
∼ 0.960–0.964
Freeze backbone 30 epochs
∼ 0.971–0.974
∼ 0.958–0.963
Freeze backbone all epochs
∼ 0.964–0.968
∼ 0.958–0.961
TABLE I: Precision and Recall Comparison
Model
mAP50
mAP50–95
Unfreeze all
∼ 0.984
∼ 0.752–0.754
Freeze backbone 30 epochs
∼ 0.987
∼ 0.783–0.784
Freeze backbone all epochs
∼ 0.986
∼ 0.783–0.784
TABLE II: mAP50 and mAP50–95 Comparison
Model
box loss
cls loss
dfl loss
Unfreeze all
∼ 0.875
∼ 0.345
∼ 1.018
Freeze backbone 30 epochs
∼ 0.862
∼ 0.360
∼ 0.955
Freeze backbone all epochs
∼ 0.933
∼ 0.380
∼ 0.960
TABLE III: Validation Losses Comparison
Model
Epochs
Unfreeze all
81
Freeze backbone 30 epochs
93 (highest)
Freeze backbone all epochs
75
TABLE IV: Training Duration
Fig. 10: Training and validation DFL loss between fine-tune-all and progressive backbone unfreezing.
Parameter
Selected value
Rationale
track_high_thresh
0.15, lower than default 0.25
Detection confidence fluctuates because of small size and fast motion. Lowering the threshold allows more valid detections to participate in the first association stage.
track_low_thresh
0.03, lower than default 0.10
A very low threshold enables recovery of weak detections during the second association stage instead of terminating the track.
new_track_thresh
0.15, lower than default 0.25
Allows newly appearing bees to be initialized more quickly, reducing missed bee entries.
track_buffer
50, higher than default 30
Keeps lost tracks alive longer, enabling successful re-association after temporary occlusion and reducing fragmented trajectories.
match_thresh
0.85, higher than default 0.80
Requires stronger spatial consistency before assigning an existing ID, reducing incorrect ID switches between nearby bees.
fuse_score
True (default)
Combines detection confidence with motion similarity during association, improving matching robustness.
TABLE V: Optimized ByteTrack Tracking Parameters and Their Rationale
Fig. 11: Bee counting IN versus ground truth.
Fig. 12: Bee counting OUT versus ground truth.
Fig. 13: Example of a fast-moving bee exiting the hive, demonstrating motion blur and an oblique viewing angle.
The CVPPA@ECCV 2026 BuzzSpot Challenge asks us to detect bees, bumblebees, hoverflies, and moths in 1920x1080 field keyframes. Its annotations carry 2 difficulties: the median box occupies 0.16% of a frame, and bees account for 80% of the labels. To cope with the small boxes, we compare 10 recorded detector configurations on held-out keyframes; plain Co-DINO with a Swin-L backbone has the highest mAP in this comparison, so we select it. Training then addresses the bee dominance in 2 ways: fine-tuning on a crop-mosaic pool in which the combined annotation share of the 3 rare classes rises from 19.9% to 55.1%, and a class-weighted simplex equiangular tight frame (ETF) loss that pulls the projected states of matched decoder queries toward fixed class directions. The full schedule spans 12+3+2 epochs. Without inference-time ensembling or test-time augmentation, we rank first on FinalTest at 0.5062 mAP@[.5:.95].
Detecting pollinators in field video is challenging: targets are small, visually similar, and observed against cluttered vegetation under blur and occlusion. We present a systematic empirical study of small-pollinator detection under a practical single-GPU compute budget. Using the BuzzSpot challenge dataset, we compare YOLO and RF-DETR models across input resolutions and evaluate sliced inference, class-gated fusion, size-routed ensembling, and post-hoc temporal processing. RF-DETR Large at 1344-pixel resolution achieved our best hidden-test result, reaching 0.405 mAP50:95 and outperforming the 1120-pixel model (0.379) and the best single-model YOLO26m baseline (0.366). The strongest gains came from adopting RF-DETR and increasing its input resolution, indicating that detector choice and input resolution were more effective levers than added inference-time complexity; the resolution gain was strongest for small objects and the rarer bumblebee and moth classes. Sliced-inference fusion, size-routed ensembling, and warm-started 1536-pixel continuation did not surpass this result, while post-hoc temporal processing did not improve the leaked diagnostic evaluation. Error analysis identified bee-hoverfly discrimination as the clearest remaining bottleneck: neighboring frames rarely supplied correctly classified hoverfly evidence for post-hoc correction. These findings motivate learned feature-level temporal aggregation before the final classification decision.
Onur Onal, Chen Chen
Iowa State University · Iowa State University, Ames, IA, USA · Institute of AI, University of Central Florida +1
The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusions. However, this strategy is nevertheless vulnerable to erroneous track reclassification and identity switching, as detection confidence scores dictate association priority. To address this, I present a simple enhancement of the ByteTrack architecture--named ByteTraX--that optimises track continuity via a single unified matching threshold, while penalising identity switches through stringent track initiation criteria. This approach achieves consistently improved performance across a range of diverse benchmarks including GMOT-40, LC-MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea-MOT, while simultaneously increasing processing speed by >10%. Specifically, results demonstrate a >40% reduction in identity switches, accompanied by mean increases in HOTA of 3.6, IDF1 of 5.6, and FPS of 6.3. As such, adoption of the ByteTraX algorithm has the potential to substantially enhance tracking performance over the ByteTrack baseline, while retaining the efficiency needed for real-time deployment. To facilitate usage, I provide the source code, integration functionality for the YOLO family of object detection models, and deployment instructions via an open source repository.