Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.
Figures & tables
Figure 1 : Lens3D brings fine-grained object understanding to 3D LLMs. While existing models capture broad scene structure and salient objects (left), Lens3D further capture instance-level attributes and easily overlooked small objects (right), with the performance radar showing these gains in fine-grained understanding do not come at the cost of the models’ original grounding and question-answering abilities (center).
Figure 2 : Overview of Lens3D. (a) LensUnd follows Ground–Focus–Understand: a 3D grounder localizes an anchor box b from a language query q , bbox-voxel selects complementary views by coverage gain and projects the target wireframe onto them, and an external 2D VLM produces a detailed understanding y⋆ of the target. (b) LensDistill trains a 3D LLM student with the sequence-level distillation loss −logpθ(y⋆∣X,c(b),p) , using teacher-generated y⋆ as privileged supervision while the student receives uniformly sampled native views, the target center c(b) encoded as <coord> , and a task prompt. (c) After distillation, the external LensUnd path is removed; the student produces the same detailed understanding from native RGB-D, the target-center, and a task prompt.
Figure 3 : Fine-grained interactions with LensUnd. Complementary views selected by LensUnd support (a) object QA on presence, appearance, and manipulation, and (b) small-object grounding by combining per-view 2D predictions into a 3D box using enclosed depth and camera parameters.
Method
Summary F1
Summary Sem.
Attr. Mean
Human Sat.
Qwen3-VL-2B [ 3 ]
0.380
0.686
0.356
–
Qwen3-VL-2B (LensUnd)
0.445
0.747
0.397
–
Qwen3-VL-8B (LensUnd)
0.485
0.783
0.528
–
Chat-3D v2
0.222
0.522
–
–
ChatScene
0.227
0.546
–
–
Video-3D LLM
0.228
0.542
–
1.88
Table 1 : Core LensBench results. “Attr. Mean” averages Parts, Materials, and State F1 and is reported only for models that produce the required structured fields. Human satisfaction (Human Sat.dcfr) is rated on a four-point scale. Best results in each column are highlighted in red , second-best in orange , and third-best in yellow .
ScanRefer
Multi3DRefer
ScanQA
SQA3D
Method
Acc@0.25
Acc@0.5
F1@0.25
F1@0.5
CIDEr
EM
Task-specific models
ScanRefer [ 5 ]
37.3
24.3
–
–
–
–
MVT [ 22 ]
40.8
33.3
–
–
–
–
3DVG-Transformer [ 42 ]
45.9
34.5
–
–
–
–
ViL3DRel [ 7 ]
47.9
37.7
–
–
–
–
Table 2 : Comparison on standard 3D scene-understanding benchmarks. Asterisks indicate results reproduced in our local environment and evaluation stack. Lens3D uses only the corresponding base model’s native input at inference. Best results in each column are highlighted in red , second-best in orange , and third-best in yellow .
Method
Unique
Multiple
@0.25
@0.5
@0.25
@0.5
Video-3D LLM
87.80
78.16
51.06
45.40
Lens3D (Video-3D LLM)
87.26
78.16
53.84
47.87
3DRS
86.72
77.40
56.87
50.79
Lens3D (3DRS)
86.45
77.56
57.44
51.25
Table 3 : ScanRefer [ 5 ] grounding on unique and multiple subsets.
Figure 4 : More qualitative results with LensUnd. Examples across multiple scenes demonstrate LensUnd’s ability to handle varied fine-grained 3D understanding scenarios.
Metric
LensUnd
ExCap3D
Unique covered voxels
2,057.7
1,564.6
Union-normalized coverage (%)
90.3
63.6
Inter-view redundancy (%) ↓
38.0
47.8
Table 4 : Target-region view coverage under the same four-view budget. Union-normalized coverage uses, for each target, the union of voxels observed by either method as the reference set; it is not absolute full-surface coverage.
Input
Summary F1
Overall F1
Summary Sem.
Unmasked
0.489
0.470
0.753
Target masked
0.369
0.337
0.560
Table 5 : Target-region masking for Lens3D (3DRS).
Type
Model
Box source
View selection
Views
Wire.
Sum. F1
Sum. Sem.
Parts
Materials
State
Mean
LensUnd
Qwen3-VL-2B [ 3 ]
Ground truth
Temporal uniform
4
No
0.263
0.619
0.192
0.607
0.272
0.357
LensUnd
Qwen3-VL-2B
Ground truth
Temporal uniform
4
Yes
0.370
0.694
0.218
0.591
0.275
0.362
LensUnd
Qwen3-VL-2B
Ground truth
Temporal uniform
32
Yes
0.380
0.686
0.201
0.589
0.278
0.356
LensUnd
Qwen3-VL-2B
Ground truth
Quality-pool uniform
4
Yes
0.426
0.732
0.245
0.628
0.300
0.391
LensUnd
Qwen3-VL-2B
Ground truth
Video-3D LLM greedy
4
Yes
0.353
0.690
0.215
0.594
0.277
0.362
LensUnd
Qwen3-VL-2B
Ground truth
ExCap3D visibility
4
Yes
0.436
0.742
0.250
0.634
0.299
0.394
Table 6 : Full LensBench evaluation. “Mean” averages Parts, Materials, and State F1. Temporal and quality-pool denote uniform sampling from the full sequence and from a quality-filtered pool, respectively. Best and second-best results within each block are bolded and underlined, respectively.