Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.
Figures & tables
Figure 1 : Lens3D brings fine-grained object understanding to 3D LLMs. While existing models capture broad scene structure and salient objects (left), Lens3D further capture instance-level attributes and easily overlooked small objects (right), with the performance radar showing these gains in fine-grained understanding do not come at the cost of the models’ original grounding and question-answering abilities (center).
Figure 2 : Overview of Lens3D. (a) LensUnd follows Ground–Focus–Understand: a 3D grounder localizes an anchor box b from a language query q , bbox-voxel selects complementary views by coverage gain and projects the target wireframe onto them, and an external 2D VLM produces a detailed understanding y⋆ of the target. (b) LensDistill trains a 3D LLM student with the sequence-level distillation loss −logpθ(y⋆∣X,c(b),p) , using teacher-generated y⋆ as privileged supervision while the student receives uniformly sampled native views, the target center c(b) encoded as <coord> , and a task prompt. (c) After distillation, the external LensUnd path is removed; the student produces the same detailed understanding from native RGB-D, the target-center, and a task prompt.
Figure 3 : Fine-grained interactions with LensUnd. Complementary views selected by LensUnd support (a) object QA on presence, appearance, and manipulation, and (b) small-object grounding by combining per-view 2D predictions into a 3D box using enclosed depth and camera parameters.
Method
Summary F1
Summary Sem.
Attr. Mean
Human Sat.
Qwen3-VL-2B [ 3 ]
0.380
0.686
0.356
–
Qwen3-VL-2B (LensUnd)
0.445
0.747
0.397
–
Qwen3-VL-8B (LensUnd)
0.485
0.783
0.528
–
Chat-3D v2
0.222
0.522
–
–
ChatScene
0.227
0.546
–
–
Video-3D LLM
0.228
0.542
–
1.88
Table 1 : Core LensBench results. “Attr. Mean” averages Parts, Materials, and State F1 and is reported only for models that produce the required structured fields. Human satisfaction (Human Sat.dcfr) is rated on a four-point scale. Best results in each column are highlighted in red , second-best in orange , and third-best in yellow .
ScanRefer
Multi3DRefer
ScanQA
SQA3D
Method
Acc@0.25
Acc@0.5
F1@0.25
F1@0.5
CIDEr
EM
Task-specific models
ScanRefer [ 5 ]
37.3
24.3
–
–
–
–
MVT [ 22 ]
40.8
33.3
–
–
–
–
3DVG-Transformer [ 42 ]
45.9
34.5
–
–
–
–
ViL3DRel [ 7 ]
47.9
37.7
–
–
–
–
Table 2 : Comparison on standard 3D scene-understanding benchmarks. Asterisks indicate results reproduced in our local environment and evaluation stack. Lens3D uses only the corresponding base model’s native input at inference. Best results in each column are highlighted in red , second-best in orange , and third-best in yellow .
Method
Unique
Multiple
@0.25
@0.5
@0.25
@0.5
Video-3D LLM
87.80
78.16
51.06
45.40
Lens3D (Video-3D LLM)
87.26
78.16
53.84
47.87
3DRS
86.72
77.40
56.87
50.79
Lens3D (3DRS)
86.45
77.56
57.44
51.25
Table 3 : ScanRefer [ 5 ] grounding on unique and multiple subsets.
Figure 4 : More qualitative results with LensUnd. Examples across multiple scenes demonstrate LensUnd’s ability to handle varied fine-grained 3D understanding scenarios.
Metric
LensUnd
ExCap3D
Unique covered voxels
2,057.7
1,564.6
Union-normalized coverage (%)
90.3
63.6
Inter-view redundancy (%) ↓
38.0
47.8
Table 4 : Target-region view coverage under the same four-view budget. Union-normalized coverage uses, for each target, the union of voxels observed by either method as the reference set; it is not absolute full-surface coverage.
Input
Summary F1
Overall F1
Summary Sem.
Unmasked
0.489
0.470
0.753
Target masked
0.369
0.337
0.560
Table 5 : Target-region masking for Lens3D (3DRS).
Type
Model
Box source
View selection
Views
Wire.
Sum. F1
Sum. Sem.
Parts
Materials
State
Mean
LensUnd
Qwen3-VL-2B [ 3 ]
Ground truth
Temporal uniform
4
No
0.263
0.619
0.192
0.607
0.272
0.357
LensUnd
Qwen3-VL-2B
Ground truth
Temporal uniform
4
Yes
0.370
0.694
0.218
0.591
0.275
0.362
LensUnd
Qwen3-VL-2B
Ground truth
Temporal uniform
32
Yes
0.380
0.686
0.201
0.589
0.278
0.356
LensUnd
Qwen3-VL-2B
Ground truth
Quality-pool uniform
4
Yes
0.426
0.732
0.245
0.628
0.300
0.391
LensUnd
Qwen3-VL-2B
Ground truth
Video-3D LLM greedy
4
Yes
0.353
0.690
0.215
0.594
0.277
0.362
LensUnd
Qwen3-VL-2B
Ground truth
ExCap3D visibility
4
Yes
0.436
0.742
0.250
0.634
0.299
0.394
Table 6 : Full LensBench evaluation. “Mean” averages Parts, Materials, and State F1. Temporal and quality-pool denote uniform sampling from the full sequence and from a quality-filtered pool, respectively. Best and second-best results within each block are bolded and underlined, respectively.
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on expert vision models with complex task-specific designs. The key argument this work wants to make is that VLMs are native 3D learners. Our in-depth large scale study shows that 1) focal length unification, 2) text-based pixel reference and 3) data mixture and scaling, are all you need for effective 3D learning. Model architecture changes, large models, heavy data augmentations, and complex losses including the regression formulation, many of which form the foundation of expert vision models, are actually not necessary conditions. As a result, we propose VLM3, a scalable method with the simplest design that enables standard VLMs to master diverse 3D tasks. VLM3 not only advances the VLM depth estimation accuracy by a large margin (0.84 -> 0.9), but also enables diverse 3D tasks such as pixel correspondence, camera pose estimation and object-level 3D understanding, matching expert vision model accuracy while maintaining standard architectures and text-based training. We believe VLM3 opens up a new paradigm for simple and scalable 3D learning.
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introduces an implicit grounding mechanism that identifies entity mentions during generation and inserts the corresponding region tokens into the text stream, allowing the model to reference visual evidence on the fly when producing spatial chain-of-thought responses. In parallel, a region-prompted monocular 3D grounding design predicts 3D bounding boxes in the camera view from grounded region queries, supported by intrinsic-aware normalization and dense geometric supervision. Together, these grounding capabilities enable GR3D to decompose complex spatial understanding problems into grounded 2D perception followed by 3D inference. GR3D achieves consistent improvements across grounded and non-grounded spatial benchmarks, demonstrating grounding as an effective inductive bias for strengthening spatial understanding in VLMs. These grounding capabilities collectively enhance general spatial understanding beyond the grounding task itself.