Long-tailed 3D object detection is treated as a class-frequency problem, but LiDAR supervision quality depends on object observability: similar frequencies can hide different geometric evidence. We introduce Geometry-Augmented Exponentially Weighted Instance-Aware Repeat Factor Sampling (GA-EIRFS), a detector-agnostic method that modulates a frequency-based repeat factor with a fixed geometry score combining point count, surface-normal entropy, and surface coverage. GA-EIRFS changes only frame-sampling probabilities, leaving the detector and inference unchanged. On nuScenes it improves mean average precision (mAP) and the nuScenes detection score (NDS) in four converged experiments with CenterPoint and PointPillars over two seeds; for CenterPoint at seed 666, mAP rises from 0.552 to 0.563 and bicycle AP from 0.306 to 0.359. Per-class gains correlate with the class sampling-weight increase (Spearman rho=0.70, p=0.025) but not with geometry score alone (rho=0.32, p=0.37), so geometry amplifies frequency-driven need. KITTI results vary across seeds, most for the rarest class. Code: https://github.com/Multimodal-Sensing-Lab/GA-EIRFS.
Figures & tables
Figure 1: Rarity and geometric difficulty are close to independent across the ten nuScenes classes. Markers are classes, placed by the E-IRFS frequency term qc and the geometry score Gc of Eq. ( 5 ). Grey curves are contours of rc ( α=2.0 , β=1.0 ): vertical distance at fixed rarity is the sampling weight that geometry adds.
Figure 2: GA-EIRFS pipeline. The priors qc and Gc of Eqs. ( 1 ) and ( 5 ) are estimated once from the training split and enter the class repeat factor of Eq. ( 6 ), which becomes a frame sampling probability.
Points/
Repeat factor rc
Class
Instances
obj.
Gc
β=0
β=1
Car
339,949
98.1
0.601
1.29
1.49
Pedestrian
161,928
11.7
0.807
1.37
1.78
Barrier
107,507
62.6
0.252
1.56
1.74
Truck
65,262
209.8
0.431
1.51
1.80
Traffic cone
62,964
9.2
0.588
1.61
2.14
Table 1: nuScenes class statistics and repeat factors, ordered by instance count. Eq. ( 6 ) uses α=2 , t=0.01 : β=0 gives E-IRFS and β=1 our default.
Figure 3: Quantitative evidence on nuScenes. (a) Per-class AP change at 20 epochs, in percentage points (pp). Bars are the mean of the two seeds, vertical markers are the individual seeds, and classes are ordered by Gc , lowest at the bottom. (b) Geometry-strength sweep at 12 epochs, CenterPoint, seed 666. (c) AP change against the exposure gain Δrc=rc(β=1)−rc(β=0) , with the Spearman rank correlation over the ten classes.
Figure 4: Top-down view of nuScenes validation sample 3896, in which the two annotated bicycles are far apart and sparsely sampled. Green boxes are ground truth, dashed red boxes are predictions scoring at least 0.3. Vanilla and E-IRFS return no bicycle above the threshold, whereas GA-EIRFS recovers both, at 0.33 and 0.52. The panels differ only in the sampler.
mAP ↑
NDS ↑
Detector
Seed
V
GA
Δ
V
GA
Δ
CenterPoint
666
0.552
0.563
+1.1
0.635
0.642
+0.7
CenterPoint
1337
0.554
0.563
+0.9
0.638
0.641
+0.3
PointPillars
666
0.385
0.388
+0.4
0.536
0.538
+0.3
PointPillars
1337
0.382
0.395
+1.3
0.534
0.541
+0.7
Mean, 4 runs
+0.9
+0.5
Table 2: nuScenes validation results after 20 epochs. V is training without rebalancing and GA is GA-EIRFS. Higher is better, Δ is in percentage points computed before rounding, and the better value of each metric within a run is in bold. The last row is the mean of the four paired differences, with 95% confidence intervals [+0.27,+1.55] for mAP and [+0.12,+0.86] for NDS.
Car
Pedestrian
Cyclist
Seed
V
GA
V
GA
V
GA
666
75.88
75.87
44.19
45.34
61.90
62.29
1337
76.28
75.81
41.93
43.74
62.69
59.99
42
75.34
75.61
44.89
44.00
61.05
62.88
Mean
75.83
75.76
43.67
44.36
61.88
61.72
SD
0.47
0.14
1.55
0.86
0.82
1.53
Table 3: KITTI 3D AP (%) at moderate difficulty, PointPillars, three seeds. Better value of each pair in bold.
Post-processing is a critical stage in LiDAR-based 3D object detection, where dense and overlapping proposals must be filtered for compact and reliable perception. This work introduces two learned filtering modules that replace heuristic non-maximum suppression (NMS) by leveraging relations among detections. D2D-Rescore employs transformer-based detection-to-detection (D2D) attention, while GossipNet3D adapts the 2D GossipNet concept to 3D through localized message passing in bird's-eye view. A metric-aware matching strategy aligned with the nuScenes evaluation protocol ensures consistent training and validation behavior, improving overall detection performance. Both approaches improve mean average precision (mAP), nuScenes detection score (NDS), and true positive quality compared to CircleNMS, particularly for small and infrequent classes, while adding minimal computational overhead. These results demonstrate that learned, detection-level filtering can enhance 3D detector reliability without modifying the base network, offering a principled alternative to heuristic suppression. Code is available at https://github.com/rst-tu-dortmund/learned-3d-nms .
Timo Osterburg, Stefan Schütte, Torsten Bertram
Institute of Control Theory and Systems Engineering, TU Dortmund University, Germany.
The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primarily on isolated 3D object models, while recent LiDAR spoofing attacks target richer and more realistic driving scenes but focus mainly on physical realizability rather than understanding detector behavior or attack efficiency. In this work, we investigate how LiDAR-based detectors rely on spatial evidence in complex scenes and whether these reliance patterns can be exploited to induce failures more efficiently. To this end, we propose an explainability-guided adversarial analysis methodology. We introduce the Saliency-LiDAR (SALL) method, which aggregates Integrated Gradient attributions across scenes to produce universal saliency maps for LiDAR-based 3D object detectors. Guided by these maps, we design the Explainability-aware Frustum Attack (EFA), which selectively perturbs only the most influential frustums rather than uniformly attacking entire object regions. Experiments on KITTI and nuScenes, across detectors such as PointPillars and SECOND, show that EFA reduces detection recall by more than 15 percentage points while requiring 25-50% fewer perturbed frustums than the state-of-the-art non-saliency-aware baseline. These findings reveal that modern 3D detectors concentrate discriminative evidence in a small subset of spatial regions, exposing a structural robustness vulnerability in current LiDAR perception systems. Our code is released at https://github.com/SecMindLab/Saliency_LiDAR.
Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.
Ömer Faruk Deniz, Mustafa Taha Koçyiğit
Institute of Data Science and Artificial Intelligence, Boğaziçi University, South Campus, Bebek, 34342, Istanbul, Türkiye