Organizations: Jilin University, China · Shenzhen University, China · Taiyuan University of Technology, China · Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, China
Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweight detectors struggle to balance accuracy and efficiency while often neglecting quality-aware feature modeling and consistency between classification and localization, leading to unstable performance under crowded conditions. To address these issues, we propose DensePed-Lite, a unified framework built on a single principle: under occlusion the network should adapt its behavior to the quality of what it observes rather than assume complete information. This principle is realized at three points where occlusion does the most damage: unreliable confidence scoring (UQE), fragmented spatial coverage (MPSC), and incoherent multi-scale fusion (CTDM). The three mechanisms reinforce one another instead of acting in isolation, all without significantly increasing complexity. Experiments on CityPersons and CrowdHuman validate that DensePed-Lite achieves a superior accuracy-efficiency trade-off compared with recent state-of-the-art lightweight methods, making it suitable for real-time deployment in dense pedestrian scenarios.
Figures & tables
Figure 1: The architecture of the network.
Figure 2: The architecture of the UQE module.
Figure 3: The architecture of the CTDM module.
CrowdHuman
CityPersons
GFLOPs
Params
Method
Recall
AP50(%)
AP(%)
Precision
Recall
AP50(%)
AP(%)
Precision
Baseline
YOLOv11n [ 6 ]
0.678
79.2
48.4
0.838
0.513
59.6
36.1
0.760
6.3
2.58M
Traditional Lightweight Detectors
YOLOv5n [ 8 ]
0.658
77.6
47.4
0.833
0.502
59.7
36.6
0.789
7.1
2.5M
YOLOv8n [ 22 ]
0.666
79.0
48.7
0.842
0.518
60.8
37.3
0.783
8.1
3.0M
Table 1: Comparison Results on CrowdHuman and CityPersons Datasets. Boldface indicates the best result and underlining indicates the second-best result among ultra-lightweight methods (Params ≤ 4M), respectively.
Figure 4: Efficiency-accuracy trade-off for ultra-lightweight detectors on CrowdHuman. Bubble size represents parameter count. DensePed-Lite achieves the best AP50 with near-minimal GFLOPs.