Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers
Authors: Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li
Organizations: Department of Automotive Engineering, Clemson University, Greenville, SC, USA. · School of Computing, Clemson University, Clemson, SC, USA.
Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.
Figures & tables
Fig. 1: Lang3DSeg segments a single LiDAR sweep with no human annotation and no vision-language model at inference (top, predictions projected onto the front camera). A PointTransformerV3 backbone trained from scratch reports the highest mIoU among annotation-free methods on both nuScenes and SemanticKITTI (bottom).
Fig. 2: The Lang3DSeg framework. (1) Targets are curated offline. SAM 3 is prompted with a descriptive ensemble, masks are rasterized under an explicit class priority so that large background regions cannot overwrite small foreground objects, and each projected instance is cut at the first gap in its depth distribution di , discarding points that lie behind the mask that claimed them. A per-sweep RANSAC ground model then separates flat terrain from vertical vegetation, and unlabeled points are filled only where neighboring labels agree. (2) A PointTransformerV3 backbone is trained from scratch on these targets; numbers give the channel width of each encoder and decoder stage. (3) Training combines weighted cross-entropy, Lovász-Softmax, and a cosine alignment term that draws the point embeddings toward frozen CLIP text anchors. The anchors supervise training only and never enter the forward path, so inference consumes one LiDAR sweep and no image data.
Class
Prompts
Priority
traffic cone
traffic cone; orange construction cone
160
pedestrian
pedestrian; person
150
bicycle
bicycle
145
motorcycle
motorcycle
140
barrier
barrier; barricade; guard rail
130
constr. vehicle
excavator; bulldozer; crane;
125
TABLE I: Prompt ensembles issued to SAM 3 and the projection priority of each class. Higher priority wins during rasterization; ties within a priority level are broken by mask score. The ordering follows one rule: a category that occupies few pixels and is easily overwritten sits above one that covers large contiguous regions, and the road surface is the base layer.
Fig. 3: Qualitative results on nuScenes validation (top two rows) and SemanticKITTI sequence 08 (bottom two rows). Left to right: camera context; ground-truth annotations; the Lang3DSeg prediction in bird’s-eye view; the same prediction in an ego-centric perspective view; and an error map with misclassified points in red, correctly classified points in grey and unannotated points in white. Lang3DSeg is trained without any human annotation and consumes no image data at inference.
Method
Venue
VLM-free
Pre-train
3D backbone
Backbone type
nuScenes
SemanticKITTI
Fully supervised upper bound (trained from scratch)
PTv3 [ 6 ]
CVPR’24
–
–
PTv3
point transformer
80.4
70.8
Annotation-free methods
CLIP2Scene [ 1 ]
CVPR’23
✓
–
SPVCNN
sparse voxel
20.8
–
HICL [ 31 ]
CVPR’24
✓
–
SPVCNN
sparse voxel
23.0
–
CNS [ 21 ]
NeurIPS’23
✓
–
MinkowskiNet
sparse voxel
26.8
–
TABLE II: Annotation-free 3D semantic segmentation on nuScenes validation and SemanticKITTI sequence 08. VLM-free indicates that no vision-language model is evaluated at test time. Pre-train indicates a self-supervised initialization of the 3D backbone. Every prior method uses a voxel-based or projection-based backbone.
Class
LOSC [ 7 ]
Lang3DSeg
Δ
barrier
5.2
50.2
+45.0
bicycle
21.3
8.8
−12.5
bus
80.5
69.0
−11.5
car
67.5
73.9
+6.4
construction vehicle
10.6
27.6
+17.0
motorcycle
64.1
46.4
−17.7
TABLE III: Per-class IoU (%) on nuScenes validation. LOSC is the strongest published annotation-free baseline; Δ is our margin over it.
Curation stage
mIoU, labeled
mIoU, all
Coverage
Projection with class priority
56.4
40.6
46.1
+ occlusion depth test
58.7
41.1
45.5
+ ground-plane refinement
58.6
41.1
45.5
+ label completion
57.1
44.4
49.2
TABLE IV: Quality of the curated targets, measured against ground truth on 1,586 validation frames without any training. Labeled scores only points that carry a pseudo-label; all charges unlabeled points as errors and therefore reflects coverage as well as correctness.
3D vision-language segmentation aims to segment target objects in 3D scenarios according to the linguistic instructions and visual observations. Prior art heavily relies on the coarse superpoint representation to reduce the computation complexity, which suffers from poor segmentation quality and messy object boundaries. In this paper, we propose the SEGment-And-select (SEGA3D) paradigm for 3D visionlanguage segmentation that directly operates on the fine-grained visual information and is free from the superpoint dependency. Specifically, we first leverage a mask candidate generator to provide fine-grained categorical mask candidates, substantially improving the quality of candidate masks over the superpoint counterparts. Then, a Large Language Model (LLM) is utilized to generate the semantic and spatial information based on the linguistic description and visual features. The LLM output and visual features are fed to the Semantic-Spatial Selector (SSS) to produce the top-ranking mask candidates. Eventually, the Loopback Verification Module (LVM) is designed to yield the segmentation mask from the selected candidate masks. Our SEGA3D attains competitive performance on ScanRefer, ScanNet and Matterport3D benchmarks. Notably, our SEGA3D surpasses the top-performing counterpart by 8.3 mIoU and 5.3 mIoU on ScanNet and Matterport3D, respectively. Codes will be available upon publication.
Yulin Chen, Zhihang Zhong, Yuenan Hou
Shanghai AI Laboratory · University of Science and Technology of China · Shanghai Jiaotong University
Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.
Ömer Faruk Deniz, Mustafa Taha Koçyiğit
Institute of Data Science and Artificial Intelligence, Boğaziçi University, South Campus, Bebek, 34342, Istanbul, Türkiye
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a joint vision-language space known to behave like a bag-of-words on compositional tasks. Furthermore, even annotation free variants often require a large 3D training corpus and a dedicated 3D encoder per domain. Instead we use a vision-language model purely as a translator. It produces structured, entity-level descriptions of each posed image. These descriptions are grounded, projected, and aggregated directly in a general-purpose, language-only embedding space, with no 3D training corpus or encoder required. On ScanNet++, our pipeline is competitive with strong annotation free baselines trained on ScanNet. On a 5-building cultural heritage benchmark, raw scores initially favor a CLIP-based variant, but a single systematic vocabulary correction reverses this ranking. An effect confirmed by a second, independent correction on a different class, indicating that language-space embeddings track physical content more faithfully. This fidelity extends to genuinely out-of-vocabulary (OOV) objects on ScanNet++ proving that language-space embeddings separate presence from absence objects far more sharply than CLIP-based embeddings do. GoDeep also localize these OOV objects within the scene, all without any 2D-3D annotation. Because every representation remains discrete text, predictions are also explainable at the point level. Finally, exploiting both a heuristic weighting, that favors precise over merely frequent observations and GoDeep's explainability property, we propose an aggregation strategy, as a proof of concept, that favors finer elements localization.
Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos
Laboratory of Photogrammetry, School of Rural, Surveying and Geoinformatics Engineering, NTUA, 15772 Athens, Greece