cs.CVOct 5, 2026

Lens3D: Target-Conditioned Visual Foveation for Fine-Grained 3D Understanding

Authors: Junming Huang, Zini Chen, Shuaiying Hou, Chi Wang, Qiang Dai, Weiwei Xu

Organizations: Zhejiang University · LIGHTSPEED

Abstract

Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.

Figures & tables

Explore similar work

CardsList
  1. VLM3: Vision Language Models Are Native 3D Learners

    May 28, 2026Zhipeng Cai, Zhuang Liu, Yunyang Xiong +33D Scene Understanding3D Generation

  2. Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

    Sep 29, 2026Jaewoo Jung, Hyeonseo Yu, Honggyu An +10Spatial ReasoningMultimodal Large Language Models

  3. Grounded 3D-Aware Spatial Vision-Language Modeling

    May 28, 2026An-Chieh Cheng, Yang Fu, Yatai Ji +123D Visual GroundingSpatial Grounding