Organizations: Jilin University, China · Shenzhen University, China · Taiyuan University of Technology, China · Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, China
Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweight detectors struggle to balance accuracy and efficiency while often neglecting quality-aware feature modeling and consistency between classification and localization, leading to unstable performance under crowded conditions. To address these issues, we propose DensePed-Lite, a unified framework built on a single principle: under occlusion the network should adapt its behavior to the quality of what it observes rather than assume complete information. This principle is realized at three points where occlusion does the most damage: unreliable confidence scoring (UQE), fragmented spatial coverage (MPSC), and incoherent multi-scale fusion (CTDM). The three mechanisms reinforce one another instead of acting in isolation, all without significantly increasing complexity. Experiments on CityPersons and CrowdHuman validate that DensePed-Lite achieves a superior accuracy-efficiency trade-off compared with recent state-of-the-art lightweight methods, making it suitable for real-time deployment in dense pedestrian scenarios.
Figures & tables
Figure 1: The architecture of the network.
Figure 2: The architecture of the UQE module.
Figure 3: The architecture of the CTDM module.
CrowdHuman
CityPersons
GFLOPs
Params
Method
Recall
AP50(%)
AP(%)
Precision
Recall
AP50(%)
AP(%)
Precision
Baseline
YOLOv11n [ 6 ]
0.678
79.2
48.4
0.838
0.513
59.6
36.1
0.760
6.3
2.58M
Traditional Lightweight Detectors
YOLOv5n [ 8 ]
0.658
77.6
47.4
0.833
0.502
59.7
36.6
0.789
7.1
2.5M
YOLOv8n [ 22 ]
0.666
79.0
48.7
0.842
0.518
60.8
37.3
0.783
8.1
3.0M
Table 1: Comparison Results on CrowdHuman and CityPersons Datasets. Boldface indicates the best result and underlining indicates the second-best result among ultra-lightweight methods (Params ≤ 4M), respectively.
Figure 4: Efficiency-accuracy trade-off for ultra-lightweight detectors on CrowdHuman. Bubble size represents parameter count. DensePed-Lite achieves the best AP50 with near-minimal GFLOPs.
Pedestrian head orientation recognition plays an important role in autonomous driving by providing valuable cues for understanding pedestrian attention and anticipating potential crossing behavior. However, reliable recognition in real-world traffic scenes remains challenging because pedestrian head regions are often captured at low resolution. To address this challenge, we propose a lightweight Low-Resolution Head Orientation Convolutional Neural Network (LRHO-CNN) for pedestrian head orientation recognition. We construct a new dataset by extracting pedestrian head images from multiple public datasets and manually annotating them into eight orientation categories. The collected images are systematically preprocessed and augmented to increase data diversity and better represent variations in illumination and image quality. The experimental analysis compares LRHO-CNN with three fine-tuned baseline models, namely ResNet-18, ResNet-34, and VGG-16. The results demonstrate that LRHO-CNN achieves the highest classification accuracy among the evaluated models. LRHO-CNN is further evaluated on the JAAD and PIE datasets, demonstrating its effectiveness in recognizing pedestrian head orientation in real-world traffic scenes and providing informative head-orientation cues that can support downstream pedestrian behavior and intention prediction.
Yuanzhe Li, Yidi Huang, Xiaotong Chang +1
Chair of Automotive Engineering, Technische Universität Berlin, Berlin, 13355, Germany.
In real-world applications, pedestrian trajectory prediction models rely on inputs from detection and tracking systems. Prior trajectory prediction benchmarks either contain relatively sparse pedestrian interactions, assume perfect tracking inputs, or rely on overhead viewpoints that minimize occlusion and perspective distortion, limiting evaluation in realistic dense-crowd scenarios. We present CrowdTraj, a benchmark for pedestrian trajectory prediction in natural dense crowd scenes. Unlike previous datasets, CrowdTraj supports end-to-end evaluation from detection through tracking to trajectory prediction under severe occlusion in CCTV views. It also captures diverse, natural pedestrian behaviours, including abrupt directional changes rarely observed in existing benchmarks. CrowdTraj includes five diverse scenes, with an average of 1,146 unique pedestrians per scene, maximum frame-level densities ranging from 114 to 372 pedestrians, and over 3.2 million annotated head bounding boxes. CrowdTraj provides pixel and real-world coordinates via per-scene homography matrices for physically meaningful analysis. Our experimental results show that tracking accuracy (IDF1) drops to 0.68 to 0.70 in the densest scenes, compared with approximately 0.90 in less crowded scenes. Trajectory prediction training also becomes substantially more computationally expensive in dense scenes, with training times increasing by up to 8 times. These findings show that CrowdTraj exposes limitations in current trajectory prediction pipelines that remain hidden on existing sparse-crowd benchmarks, particularly in robustness to tracking noise and computational scalability.
Antonius Bima Murti Wijaya, Paul Henderson, Marwa Mahmoud
School of Computing Science University of Glasgow University Avenue G12 8QQ, Glasgow, United Kingdom
Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and creating severe safety risks. Despite their effectiveness in controlled settings, existing physical attacks face two major limitations in practice: they lack systematic disruption of the multi-stage decision pipeline, enabling residual modules to offset perturbations, and they fail to model complex physical variations, leading to poor robustness. To overcome these limitations, we propose a novel pedestrian adversarial patch generation method that combines multi-stage collaborative attacks with robustness enhancement under physical diversity, called TriPatch. Specifically, we design a triplet loss consisting of detection confidence suppression, bounding-box offset amplification, and non-maximum suppression (NMS) disruption, which jointly act across different stages of the detection pipeline. In addition, we introduce an appearance consistency loss to constrain the color distribution of the patch, thereby improving its adaptability under diverse imaging conditions, and incorporate data augmentation to further enhance robustness against complex physical perturbations. Extensive experiments demonstrate that TriPatch achieves a higher attack success rate across multiple detector models compared to existing approaches.
Shihui Yan, Ziqi Zhou, Yufei Song +3
Huazhong University of Science and Technology China