Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
Figures & tables
Figure 1: Overall architecture of Imagine3D-LLM. Given multi-view images of a scene and a text instruction, our model appends a compact set of learnable Gaussian summary tokens between the image and text tokens. At an intermediate layer ℓ , the hidden states at the Gaussian summary token positions are decoded into a compact 3D Gaussian Splatting representation that abstractly reconstructs the scene. The model is trained jointly with a standard language modeling loss on the text output, a photometric reconstruction loss obtained by rendering the predicted Gaussians at the input viewpoints, and a distillation loss from a pretrained compact Gaussian teacher [ 1 ] .
Figure 2: Effects of joint reconstruction training on the LLM’s image features. (a) Image-to-image attention maps between a query point ( red dot) and other input views. Our model exhibits substantially stronger cross-view correspondences than the baseline, indicating that the Gaussian summary tokens implicitly drive the image features to align across views. (b) PCA visualization of the LLM’s image features (first three components shown as RGB). Our features are noticeably cleaner and more semantically structured, with corresponding objects (e.g., chairs, desks) encoded with consistent colors across views, reflecting the emergence of object-level abstractions.
Figure 3: Visualization of clustered Gaussian from summary tokens. We visualize the Gaussians clustered from a set of Gaussian summary tokens via K-means over the summary tokens ( cyan indicates the Gaussians belonging to one cluster). Similar summary tokens are mapped to the same object without any explicit clustering supervision.
Method
SQA3D test
Real-3DQA test
ScanQA val
EM
EM-R
EM
EM-R
CIDEr
BLEU-4
METEOR
ROUGE
EM
Specialists
SQA3D [ 52 ]
46.6
–
–
–
–
–
–
–
–
ScanQA [ 3 ]
–
–
–
–
64.9
10.1
13.1
33.3
21.1
2D MLLMs
InternVL2-8B [ 67 ]
33.0
45.3
–
–
62.5
3.3
14.5
34.3
–
Table 1: Quantitative results on 3D question-answering over SQA3D [ 52 ] , Real-3DQA [ 51 ] , and ScanQA [ 3 ] . “Specialists” refer to models tailored to individual tasks via task-specific decoders. General-purpose 2D MLLMs [ 67 , 75 , 89 ] are reported under a zero-shot protocol. We report values available from prior works; “–” marks entries unavailable to us.
Method
Scan2Cap val (IoU@0.5)
ScanRefer val
Multi3DRefer val
ROUGE
BLEU-4
METEOR
CIDEr
Acc@0.25
Acc@0.5
F1@0.25
F1@0.5
Specialists
Scan2Cap [ 11 ]
44.5
23.3
22.0
35.2
–
–
–
–
3DJCG [ 8 ]
50.8
31.0
24.2
49.5
49.6
37.3
–
26.6
ScanRefer [ 10 ]
–
–
–
–
37.3
24.3
–
–
M3DRef-CLIP [ 88 ]
–
–
–
–
51.9
44.7
42.8
38.4
Table 2: Evaluation of 3D dense captioning and visual grounding on Scan2Cap [ 11 ] , ScanRefer [ 10 ] , and Multi3DRefer [ 88 ] .
Model
SPAR-Bench
Avg.
Low
Med.
High
Proprietary
GPT-4o [ 34 ]
36.4
29.3
24.9
45.1
Claude-3.7-Sonnet [ 2 ]
21.8
25.4
7.3
23.3
General-purpose 2D MLLMs
LLaVA-Video-7B [ 89 ]
32.3
23.6
24.8
42.6
Table 3: Evaluation on SPAR-Bench [ 87 ] .
Method
SQA3D
ScanQA
Scan2Cap
Baseline
56.5
26.2
63.1
Imagine3D-LLM (Ours)
63.8
29.9
67.6
Method
ScanRefer
Multi3DRefer
SPAR
Baseline
58.3
57.4
60.9
Imagine3D-LLM (Ours)
62.8
60.2
68.5
Table 4: Controlled comparison against the base MLLM. Both models share the identical backbone, training data, and schedule; the baseline is trained with standard visual instruction tuning, without the Gaussian summary tokens or the reconstruction and distillation objectives.
Layer ℓ
SQA3Dtest
SPAR-Bench
7 (early)
60.7
60.2
14 (middle, Ours)
63.8
68.5
21 (late)
61.7
64.0
Table 5: Effect of the layer ℓ from which Gaussian tokens are decoded.
Method
SQA3Dtest
SPAR-Bench
Baseline (1 ep.)
56.5
60.9
Baseline (2 ep.)
57.2
61.5
Baseline (4 ep.)
52.3
55.2
Recon. only (1 ep.)
57.7
57.6
Recon. only (2 ep.)
61.9
65.1
Recon. only (4 ep.)
63.7
67.9
Table 6: Effect of distillation from the compact Gaussian teacher.
Variant
SQA3Dtest
SPAR-Bench
Baseline
56.5
60.9
+ Teacher tokens
56.4
60.8
+ Distill. only
56.6
60.5
+ Recon. (4 ep.)
63.7
67.9
Full (Ours, 1 ep.)
63.8
68.5
Table 7: Isolating the source of the gains. + Teacher tokens : the teacher’s query tokens are fed directly to the LLM as input. + Distill. only : summary tokens supervised by Ldistill but without Lrecon .
M
SQA3Dtest
SPAR-Bench
1296
61.9
64.6
2592 (Ours)
63.8
68.5
5184
60.3
58.4
Table 8: Number of Gaussian summary tokens M .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Layer-wise cross-view correspondence during training. We report cross-view correspondence scores for image-token representations extracted from middle LLM layers ℓ=11∼18 on SQA3D [ 52 ] test samples. Scores increase consistently throughout training, with pronounced gains near the Gaussian decoding layer ( ℓ=14 ), indicating that joint reconstruction training improves the view-consistency of intermediate image features. Shaded regions indicate variance across SQA3D [ 52 ] test samples.
Method
mIoU
Baseline (no summary tokens)
26.54
Imagine3D-LLM (Ours)
35.43
DINOv3 [ 65 ]
37.62
Appendix
Table 9: Unsupervised semantic segmentation probe. We cluster the LLM’s image features via K-means and measure mIoU against ground-truth ScanNet labels. DINOv3 [ 65 ] is evaluated under the same protocol as an upper-bound reference.
Method
PSNR ( ↑ )
SSIM ( ↑ )
LPIPS ( ↓ )
Gaussian teacher [ 72 ]
20.60
0.76
0.44
Imagine3D-LLM (Ours)
17.61
0.74
0.49
Appendix
Table 10: Reconstruction quality on SQA3D test scenes. We compare the rendering quality of our Gaussian summary tokens against the compact Gaussian teacher [ 72 ] used during training. Higher is better for PSNR and SSIM; lower is better for LPIPS.
Figure 5: Visualization of attention maps between multi-view images and PCA visualizations of the image features. (a) Image-to-image attention maps between a query point ( red dot) and other input views. Our model exhibits substantially stronger cross-view correspondences than the baseline, indicating that the Gaussian summary tokens implicitly drive the image features to align across views. (b) PCA visualization of the LLM’s image features (first three components shown as RGB). Our features are noticeably cleaner and more semantically structured, with corresponding objects (e.g., chairs, desks) encoded with consistent colors across views, reflecting the emergence of object-level abstractions.
Figure 6: Visualization of clustered Gaussian from summary tokens. We visualize the Gaussians clustered from a set of Gaussian summary tokens via K-means over the summary tokens. Similar summary tokens are mapped to the same object without any explicit clustering supervision.
Figure 7: Visualization of attention maps between Gaussian summary tokens and images. We visualize the attention score map between a Gaussian summary token (a red dot indicating the decoded Gaussian) and the input images. Each Gaussian token attends to corresponding regions of the same object across multiple views.
Training setting
Peak VRAM (GB)
w/o Lrecon+Ldistill
40
Imagine3D-LLM (full)
44
Appendix
Table 11: Training-time computation cost. Peak per-GPU VRAM measured during training under identical batch size and sequence length. The reconstruction and distillation losses introduce only a modest memory overhead while substantially improving 3D reasoning.
Method
Peak Memory (GB)
Inference Time (ms)
Baseline (7B)
17.09
112.1
VLM3R-7B [ 23 ]
25.29
344
Imagine3D-LLM (7B)
20.43
176.9
Appendix
Table 12: Inference cost comparison. Peak GPU memory and per-sample inference time measured on the SPAR evaluation set. VLM3R-7B [ 23 ] additionally runs an external 3D foundation model (CUT3R [ 76 ] ) on all input frames before fusion.
Figure 8: Visualization of reconstructed scene with estimated Gaussians. (Top) : input multi-view images. (Bottom) : a view rendered from the 3D Gaussians decoded by our Gaussian summary tokens. Although the rendering is coarse due to the compact M=2592 token bottleneck, it preserves the overall scene layout and the placement of major objects, indicating that the tokens encode a coherent 3D abstraction of the scene.
Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose SceneScaffold, an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases. Code is available at https://github.com/lixiangqi707/SceneScaffold.
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What'' questions in SQA3D, while maintaining strong overall accuracy (50.8%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.