cs.CVOct 1, 2026

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

Authors: Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li

Organizations: Department of Automotive Engineering, Clemson University, Greenville, SC, USA. · School of Computing, Clemson University, Clemson, SC, USA.

Abstract

Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.

Figures & tables

Explore similar work

CardsList
  1. Segment and Select: Vision-Language Segmentation in 3D Scenarios

    Jun 9, 2026Yulin Chen, Zhihang Zhong, Yuenan Hou3D Generation

  2. Open-vocabulary 3D object detection with promptable segmentation

    Sep 16, 2026Ömer Faruk Deniz, Mustafa Taha Koçyiğit3D Object DetectionObject Detection

  3. GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

    Sep 8, 2026Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos3D Scene UnderstandingVisual Embeddings