cs.CVSep 24, 2026

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

Authors: Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang

Organizations: Department of Automotive Engineering (Automotive-Computer Convergence), Hanyang University, Seoul, South Korea. · Department of Automotive Engineering, Hanyang University, Seoul, South Korea.

Abstract

LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.

Figures & tables

Explore similar work

CardsList
  1. Do LiDAR Language Models Really Understand Spatio-temporal Relationships?

    Sep 21, 2026Runyi Yang, Murat Akkoyun, Di Wen +8Spatio-Temporal ReasoningLidar

  2. OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

    Jun 21, 2026Hao Vo, Phu Loc Nguyen, Khoa Vo +7Multimodal Large Language ModelsEfficient Planning

  3. SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs

    May 27, 2026Jiawei Li, Ziyi Liu, Weijie Shi +33D Visual Grounding3D Spatial Reasoning