Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense annotation. This motivates a simple question of whether high-quality LiDAR training data can be produced automatically, without any human labeling. To this end, we introduce LiDAR-SAM2, a framework that turns a 2D video foundation model, SAM2, into a scalable source of supervision for the 4D LiDAR domain. On the data side, it automatically generates temporally coherent LiDAR-level labels from SAM2 video masks through multi-view projection and spatio-temporal aggregation. On the modeling side, a tailored modality interface and a two-stage learning objective adapt SAM2's video segmentation kernel to spatio-temporal LiDAR structure, so that a single click per object yields a consistent mask track across the sequence. Trained with no human LiDAR annotation, LiDAR-SAM2 produces semantic and panoptic labels on SemanticKITTI that approach the quality of full human annotation from only a few points, and models trained on these labels approach the performance of full ground-truth supervision. This positions LiDAR-SAM2 as a scalable labeling tool that substantially reduces the annotation burden for 3D and 4D scene understanding.
Figures & tables
Figure 2: Overview of the LiDAR-SAM2 framework. The training pipeline consists of two stages, both supervised by pseudo-labels.
Methods
Label type
car
bicycle
motorcycle
truck
other-vehicle
person
bicyclist
motorcyclist
road
parking
sidewalk
other-ground
building
fence
vegetation
trunk
terrain
pole
traffic-sign
mIoU
MinkU [ 6 ]
Point labels
9.7
0.5
2.4
5.6
4.0
0.8
1.6
3.4
0.8
2.5
4.1
0.2
5.3
4.0
1.8
10.4
10.4
23.1
7.4
5.2
SAM2 labels
61.0
0.6
6.5
50.3
11.9
25.1
44.7
0.1
57.8
6.0
7.4
0.6
67.6
21.8
53.5
21.3
47.8
45.7
37.2
29.8
Our labels
84.3
17.8
60.5
75.8
46.8
58.3
73.0
2.5
85.2
36.8
69.6
0.3
84.6
52.2
84.7
63.6
69.9
49.6
33.2
55.2
Full GT
96.1
35.0
64.8
83.4
59.9
73.4
88.3
0.0
93.7
53.1
81.0
7.0
91.0
61.7
88.0
67.8
75.5
62.5
48.6
64.8
PTv2 [ 36 ]
Point labels
27.6
1.9
5.8
3.2
6.6
0.9
2.0
0.0
1.2
4.8
0.2
0.1
6.2
3.7
1.8
15.0
2.7
41.0
16.6
7.4
SAM2 labels
78.7
0.1
0.0
0.0
15.3
14.6
22.3
1.9
53.9
1.4
0.7
0.0
72.1
23.1
64.5
22.8
51.5
43.2
35.9
26.4
Table 1: Semantic segmentation results (mIoU) trained with point labels, SAM2 labels, our labels, and full GT. Experiments are on the SemanticKITTI [ 2 ] validation set.
Methods
Label type
LSTQ
Sassoc
Scls
IoU st
IoU th
4D-PLS [ 1 ]
Our labels
46.9
46.3
47.6
54.3
44.3
4D-PLS [ 1 ]
Full GT
62.7
65.1
60.5
65.4
61.3
4D-StOP [ 13 ]
Our labels
57.7
64.2
51.7
56.5
51.7
4D-StOP [ 13 ]
Full GT
67.0
74.4
60.3
65.3
60.9
Mask4Former [ 41 ]
Our labels
64.8
70.7
59.4
56.9
62.8
Mask4Former [ 41 ]
Full GT
70.5
74.3
66.9
67.1
66.6
Table 2: Panoptic segmentation results trained with our labels and full GT. Experiments are evaluated on SemanticKITTI [ 2 ] validation set.
Figure 3: Qualitative results of labels annotated with LiDAR-SAM2. The top shows instance labels, while the bottom shows semantic labels.
Constructing faithful 4D worlds from LiDAR-acquired sequences is crucial for embodied AI, yet current generative frameworks apply uniform modeling capacity across all spatial regions. This ignores that perceptual difficulty varies dramatically within a single scan: distant surfaces, occluded boundaries, and small-scale objects carry far higher uncertainty than well-observed structures. We present U4D, a new framework that explicitly leverages spatial uncertainty to guide LiDAR scene generation in a "hard-to-easy" schedule. U4D derives per-point uncertainty maps via Shannon Entropy from a pretrained segmentor, then applies an unconditional diffusion stage to synthesize high-entropy areas with precise geometry, followed by a conditional completion stage that fills in the remaining regions using these structures as priors. A MoST (Mixture of Spatio-Temporal) block further maintains cross-frame coherence by dynamically balancing spatial detail and temporal continuity. Extensive experiments on nuScenes and SemanticKITTI demonstrate state-of-the-art scene fidelity, temporal consistency, and downstream performance.
Frame-wise semantic segmentation of indoor lidar scans is a fundamental step toward higher-level 3D scene understanding and mapping applications. However, acquiring frame-wise ground truth for training deep learning models is costly and time-consuming. This challenge is largely addressed, for imagery, by Visual Foundation Models (VFMs) which segment image frames. The same VFMs may be used to train a lidar scan frame segmentation model via a 2D-to-3D distillation pipeline. The success of such distillation has been shown for autonomous driving scenes, but not yet for indoor scenes. Here, we study the feasibility of repeating this success for indoor scenes, in a frame-wise distillation manner by coupling each lidar scan with a VFM-processed camera image. The evaluation is done using indoor SLAM datasets, where pseudo-labels are used for downstream evaluation. Also, a small manually annotated lidar dataset is provided for validation, as there are no other lidar frame-wise indoor datasets with semantics. Results show that the distilled model achieves up to 56% mIoU under pseudo-label evaluation and around 36% mIoU with real-label, demonstrating the feasibility of cross-modal distillation for indoor lidar semantic segmentation without manual annotations.
Haiyang Wu, Juan J. Gonzales Torres, George Vosselman +1
Department of Earth Observation Science, Faculty of Geo-Information Science and Earth Observation (ITC), University of Twente, 7522 NB Enschede, The Netherlands
Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do not fully exploit the teacher's layer-wise semantic structure and global context, as well as the rich spatiotemporal information inherent in LiDAR sequences. We propose HilDA, a self-supervised pretraining framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. HilDA combines hierarchical distillation comprising multi-layer distillation for progressive semantic alignment and global context distillation for scene-level semantics, with a temporal occupancy diffusion objective promoting spatiotemporal consistency. Models pre-trained with HilDA achieve state-of-the-art results on cross-modal distillation benchmarks and outperform models trained via prior distillation approaches on 3D object detection, scene flow, and semantic occupancy prediction. Code available at: https://maxiuw.github.io/hilda.
Maciej Wozniak, Jesper Ericsson, Hariprasath Govindarajan +4
KTH Royal Institute of Technology, Sweden · TRATON AB, Sweden · Linköping University, Sweden +1