Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers
Authors: Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li
Organizations: Department of Automotive Engineering, Clemson University, Greenville, SC, USA. · School of Computing, Clemson University, Clemson, SC, USA.
Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.
Figures & tables
Fig. 1: Lang3DSeg segments a single LiDAR sweep with no human annotation and no vision-language model at inference (top, predictions projected onto the front camera). A PointTransformerV3 backbone trained from scratch reports the highest mIoU among annotation-free methods on both nuScenes and SemanticKITTI (bottom).
Fig. 2: The Lang3DSeg framework. (1) Targets are curated offline. SAM 3 is prompted with a descriptive ensemble, masks are rasterized under an explicit class priority so that large background regions cannot overwrite small foreground objects, and each projected instance is cut at the first gap in its depth distribution di , discarding points that lie behind the mask that claimed them. A per-sweep RANSAC ground model then separates flat terrain from vertical vegetation, and unlabeled points are filled only where neighboring labels agree. (2) A PointTransformerV3 backbone is trained from scratch on these targets; numbers give the channel width of each encoder and decoder stage. (3) Training combines weighted cross-entropy, Lovász-Softmax, and a cosine alignment term that draws the point embeddings toward frozen CLIP text anchors. The anchors supervise training only and never enter the forward path, so inference consumes one LiDAR sweep and no image data.
Class
Prompts
Priority
traffic cone
traffic cone; orange construction cone
160
pedestrian
pedestrian; person
150
bicycle
bicycle
145
motorcycle
motorcycle
140
barrier
barrier; barricade; guard rail
130
constr. vehicle
excavator; bulldozer; crane;
125
TABLE I: Prompt ensembles issued to SAM 3 and the projection priority of each class. Higher priority wins during rasterization; ties within a priority level are broken by mask score. The ordering follows one rule: a category that occupies few pixels and is easily overwritten sits above one that covers large contiguous regions, and the road surface is the base layer.
Fig. 3: Qualitative results on nuScenes validation (top two rows) and SemanticKITTI sequence 08 (bottom two rows). Left to right: camera context; ground-truth annotations; the Lang3DSeg prediction in bird’s-eye view; the same prediction in an ego-centric perspective view; and an error map with misclassified points in red, correctly classified points in grey and unannotated points in white. Lang3DSeg is trained without any human annotation and consumes no image data at inference.
Method
Venue
VLM-free
Pre-train
3D backbone
Backbone type
nuScenes
SemanticKITTI
Fully supervised upper bound (trained from scratch)
PTv3 [ 6 ]
CVPR’24
–
–
PTv3
point transformer
80.4
70.8
Annotation-free methods
CLIP2Scene [ 1 ]
CVPR’23
✓
–
SPVCNN
sparse voxel
20.8
–
HICL [ 31 ]
CVPR’24
✓
–
SPVCNN
sparse voxel
23.0
–
CNS [ 21 ]
NeurIPS’23
✓
–
MinkowskiNet
sparse voxel
26.8
–
TABLE II: Annotation-free 3D semantic segmentation on nuScenes validation and SemanticKITTI sequence 08. VLM-free indicates that no vision-language model is evaluated at test time. Pre-train indicates a self-supervised initialization of the 3D backbone. Every prior method uses a voxel-based or projection-based backbone.
Class
LOSC [ 7 ]
Lang3DSeg
Δ
barrier
5.2
50.2
+45.0
bicycle
21.3
8.8
−12.5
bus
80.5
69.0
−11.5
car
67.5
73.9
+6.4
construction vehicle
10.6
27.6
+17.0
motorcycle
64.1
46.4
−17.7
TABLE III: Per-class IoU (%) on nuScenes validation. LOSC is the strongest published annotation-free baseline; Δ is our margin over it.
Curation stage
mIoU, labeled
mIoU, all
Coverage
Projection with class priority
56.4
40.6
46.1
+ occlusion depth test
58.7
41.1
45.5
+ ground-plane refinement
58.6
41.1
45.5
+ label completion
57.1
44.4
49.2
TABLE IV: Quality of the curated targets, measured against ground truth on 1,586 validation frames without any training. Labeled scores only points that carry a pseudo-label; all charges unlabeled points as errors and therefore reflects coverage as well as correctness.