Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained prediction may provide a useful target location while enclosing sparse foreground returns, background clutter, or points inconsistent with the predicted box. We refer to this mismatch as box-point inconsistency. We introduce SimFuse3D, which preserves the target placement and repairs the associated pseudo object using measured geometry from labeled source scans. Object Memory retrieves a similar labeled source instance. Target Simulation places the retrieved source geometry at the target location, aligns its points with the target viewing geometry, and filters the aligned crop to approximate the target observation. Confidence-Guided Multi-Stage Localization Reweighting (CMLR) maps each target pseudo-object confidence score to a bounded weight shared by RPN localization and R-CNN box regression. All components operate only during adaptation, leaving the detector architecture and inference graph unchanged. Across six cross-platform transfers, SimFuse3D consistently outperforms Pi3DET-Net and achieves the best performance among the compared adaptation methods on nearly all metrics. On nuScenes-to-KITTI, it ranks first among the compared adaptation methods with both evaluated detectors.
Figures & tables
Fig. 1: Motivation and qualitative comparison. (a) Pi3DET-Net [ 3 ] retains the target point observations enclosed by accepted pseudo boxes, allowing unreliable geometry to enter localization supervision. (b) Our method, SimFuse3D, retrieves measured source geometry, aligns it with the target placement and viewpoint, and reweights target localization supervision with CMLR. (c) Vehicle → Drone example. The upper dashed green circle marks one ground-truth vehicle, and the lower circle marks two adjacent vehicles. All three boxes contain only a few LiDAR returns. Pi3DET-Net misses the three vehicles in the two marked regions, whereas SimFuse3D detects them. Red boxes show ground truth, blue boxes show predictions, and purple points are returns inside the ground-truth boxes.
Fig. 2: Overview of our SimFuse3D framework . (a) The detector is initialized by source-domain pre-training, and (b) target inference generates pseudo labels for adaptation. During adaptation, (c) Object Memory retrieves labeled source instances for target pseudo objects, (d) Target Simulation repairs target pseudo objects with retrieved source geometry, and (e) CMLR reweights localization supervision according to pseudo-object confidence. The modules in (c)–(e) are used only during adaptation and are absent from the deployed inference graph.
Fig. 3: Object Memory and the proposed Target Simulation . Object Memory stores source points, ground-truth boxes, metadata, and fixed Point-NN [ 21 ] descriptors for source-instance retrieval. A target pseudo object queries the memory for a labeled source instance. Target Simulation uses the target pseudo object to guide viewpoint alignment and point-distribution simulation, producing the simulated target observation.
Source Domain
Method
Pi3DET (Quadruped)
Pi3DET (Drone)
PV-RCNN
Voxel R-CNN
PV-RCNN
Voxel R-CNN
AP@0.7
AP@0.5
AP@0.7
AP@0.5
AP@0.7
AP@0.5
AP@0.7
AP@0.5
nuScenes
Source Only
40.11 / 31.23
43.95 / 42.00
41.84 / 33.39
45.27 / 43.36
35.56 / 24.96
38.36 / 35.76
37.84 / 24.21
44.53 / 39.60
ST3D [ 4 ]
54.83 / 42.66
58.94 / 56.80
50.91 / 41.38
53.25 / 51.69
51.56 / 31.54
58.35 / 51.54
52.20 / 33.53
56.85 / 51.29
ST3D ‡ [ 4 ]
54.68 / 43.34
58.64 / 57.78
50.46 / 40.95
53.02 / 51.20
50.94 / 31.84
57.68 / 51.07
51.95 / 33.25
56.65 / 51.14
ST3D++ [ 5 ]
56.08 / 41.90
58.84 / 58.03
51.05 / 41.45
53.76 / 51.70
51.26 / 33.35
56.63 / 51.34
54.32 / 34.70
59.41 / 54.03
TABLE I: Cross-platform adaptation to Pi3DET Quadruped and Drone. Entries are APBEV/AP3D (%). Entries within each detector column share the detector and evaluation configuration; Pi3DET-Net is reproduced under this setup. ST3D/ST3D++ use random object scaling unless ‡ marks no ROS. “Target Platform” is fully supervised; the best scores among the listed adaptation methods are bolded.
Task
Method
PV-RCNN
Voxel R-CNN
AP@0.7
AP@0.5
AP@0.7
AP@0.5
Q → D
Source Only
27.52 / 11.76
32.60 / 27.62
28.02 / 12.14
33.93 / 28.54
ST3D [ 4 ]
31.14 / 13.93
38.27 / 31.90
30.31 / 20.06
36.03 / 33.05
ST3D++ [ 5 ]
34.35 / 16.46
41.97 / 36.98
33.88 / 20.89
39.29 / 36.31
ReDB [ 12 ]
27.71 / 16.79
33.57 / 29.66
32.21 / 22.57
36.80 / 34.13
MS3D++ [ 6 ]
31.77 / 16.47
40.00 / 35.27
26.68 / 19.43
29.41 / 29.20
TABLE II: Bidirectional adaptation between Pi3DET Quadruped (Q) and Drone (D). Entries are APBEV/AP3D (%). Entries within each detector column share the detector and evaluation configuration; Pi3DET-Net is reproduced under this setup. “Target Platform” is fully supervised; the best scores among the listed adaptation methods are bolded.
Method
APBEV
AP3D
Source Only
64.98
38.85
SN [ 7 ]
55.00
43.95
ST3D [ 4 ]
74.47
50.24
ST3D [ 4 ] w/ SN
78.40
70.90
ST3D++ [ 5 ]
81.70
45.35
ST3D++ [ 5 ] w/ SN
84.98
75.50
TABLE III: nuScenes → KITTI adaptation with PV-RCNN at IoU 0.7 . Entries are AP (%); the best adaptation scores are bolded. “w/ SN” denotes statistical normalization, and “Target Platform” is fully supervised.
Method
APBEV
AP3D
Source Only
41.58
17.62
SN [ 7 ]
33.95
23.26
ST3D [ 4 ]
77.29
35.58
ST3D [ 4 ] w/ SN
87.11
66.02
ST3D++ [ 5 ]
82.08
38.10
ST3D++ [ 5 ] w/ SN
85.84
68.64
TABLE IV: nuScenes → KITTI adaptation with Voxel R-CNN at IoU 0.7 . Entries are AP (%); the best adaptation scores are bolded. Unreported methods are omitted, and “Target Platform” is fully supervised.
TS
CMLR
Pi3DET (Quadruped)
Pi3DET (Drone)
AP@0.7
AP@0.5
AP@0.7
AP@0.5
58.76 / 47.44
63.68 / 61.22
65.82 / 49.59
71.22 / 65.57
✓
60.94 / 48.22
65.97 / 63.58
67.53 / 49.66
70.94 / 67.36
✓
✓
62.54 / 49.34
67.71 / 65.40
67.58 / 50.77
73.05 / 69.35
TABLE V: Component ablation with Voxel R-CNN on the Vehicle-source transfers. TS and CMLR denote Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting. Entries are APBEV/AP3D (%); best scores are bolded.
Task
GT pts.
#GT
Dist. (m)
R@0.5
R@0.7
V → Q
0–15
304
28.24
9.54/ 11.51 (+1.97)
4.93/ 5.26 (+0.33)
16–30
180
25.16
45.56/ 51.11 (+5.56)
27.22/ 30.00 (+2.78)
≥31
1,581
18.46
80.83/ 86.08 (+5.25)
66.86/ 70.46 (+3.61)
All
2,065
20.41
67.26/ 72.06 (+4.79)
54.29/ 57.34 (+3.05)
V → D
0–15
1,573
36.10
29.37/ 36.62 (+7.25)
16.59/ 19.64 (+3.05)
16–30
871
28.77
63.15/ 65.21 (+2.07)
42.37/ 44.32 (+1.95)
TABLE VI: Voxel R-CNN recall by target ground-truth point support. Entries are Pi3DET-Net/Ours ( Δ ), in percent; V, Q, and D denote Vehicle, Quadruped, and Drone. #GT is the object count, Dist. is the median sensor distance, and Δ is computed before rounding.
Fig. 4: Pseudo-label reliability on Vehicle → Drone. Each cell reports Q@IoU≥0.5 (%) and the sample count n for a teacher-confidence and in-box point-support bin.
Fig. 5: Vehicle → Drone comparison. Orange outlines mark three sparse vehicles recovered by SimFuse3D at IoU 0.5 but missed by Pi3DET-Net.