Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.
Figures & tables
Figure 1: Comparison of MLLM perception paradigms. Existing methods either use text tokens (high token count, e.g., 512 tokens for masks) or special coordinate tokens (limited range and precision) to encode perception outputs. Our dynamic vector representation unifies diverse perception representations by converting them into 1D sequences and learning a unified codebook in high-dimensional space. This yields compact and high precision representations with fewer tokens, delivering consistent latency reduction across 2D and 3D tasks.
Figure 2: The illustration of the dynamic vector representation learning. We first use the 1D sequence as 2D and 3D perceptual representation. Then, we use an encoder to map these sequences to high-dimensional space to adaptively quantize and compress these 1D information. Moreover, we introduce the task-agnostic geometric loss to improve the heterogeneous perception representation learning capacity. Finally, we use a decoder to recover discrete tokens to original representation.
Figure 3: Overview of DVD. We extend the MLLM vocabulary with learned unified tokens and train the model following standard LLM paradigms without modifications. At inference, we select tokens belonging to perception tasks and use a lightweight de-tokenizer reconstructs the original perceptual outputs from the discrete tokens with minimal overhead, drastically reducing sequence length and improving efficiency.
Method
3D Grounding
2D Grounding
2D RES
SUN-RGBD
Hypersim
nuScenes
RefCOCO
RefCOCO+
RefCOCOg
RefCOCO
RefCOCO+
RefCOCOg
GroundingDINO Liu et al. (2024a)
–
–
–
90.6
88.2
86.1
–
–
–
VistaLLM-7B Pramanick et al. (2024)
–
–
–
88.1
82.9
83.6
74.5
69.1
69.0
LISA-7B Lai et al. (2024)
–
–
–
–
–
–
74.9
65.1
67.9
Text4Seg Lan et al. (2025)
–
–
–
88.3
83.5
82.4
74.7
68.5
70.7
Qwen2.5-VL-3B Bai et al. (2025b)
–
–
–
89.1
82.4
85.2
–
–
–
Table 1: Performance comparison of 2D and 3D perception.
Method
SUN-RGBD
Hypersim
ARKitScenes
KITTI
nuScenes
Gemini 2.0 Pro Team et al. (2024)
32.5
–
–
–
–
Gemini 2.5 Pro Comanici et al. (2025)
29.7
–
–
–
–
Seed1.5-VL Guo et al. (2025)
33.5
–
–
–
–
Qwen3-VL-2B Bai et al. (2025a)
33.8
12.0
–
–
–
Qwen3-VL-8B Bai et al. (2025a)
36.2
12.7
–
–
–
VST-3B Yang et al. (2025c)
37.3
–
51.7
–
–
Table 2: Performance comparison of 3D grounding on both indoor and outdoor datasets.
Method
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
GroundingDINO Liu et al. (2024a)
90.6
93.2
88.2
88.2
89.0
75.9
86.1
87.0
Qwen2.5-VL-7B Bai et al. (2025b)
90.0
92.5
85.4
84.2
89.1
76.9
87.2
87.2
Rex-Omni Jiang et al. (2026)
–
–
–
–
–
–
86.6
86.8
Seed1.5-VL Guo et al. (2025)
–
–
–
–
–
–
84.7
85.2
InternVL3.5-8B Wang et al. (2025b)
92.4
94.7
88.7
87.9
92.4
82.4
89.6
89.4
Table 3: Performance comparison of 2D grounding on the RefCOCO series.
Method
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
VLT Ding et al. (2021)
67.5
70.5
65.2
56.3
61.0
50.1
55.0
57.7
LAVT Yang et al. (2022)
72.7
75.8
68.8
62.1
68.4
55.1
61.2
62.1
VistaLLM-7B Pramanick et al. (2024)
74.5
76.0
72.7
69.1
73.7
64.0
69.0
70.9
LISA-7B Lai et al. (2024)
74.9
79.1
72.3
65.1
70.8
58.1
67.9
70.6
M 2 SA Jang et al. (2025)
74.0
76.8
69.7
63.1
67.2
56.1
67.0
68.3
Table 4: Performance comparison of 2D referring expression segmentation on the RefCOCO series.
Figure 4: Computation cost and performance comparison of different representations. Bubble area indicates the token count generated for per object of the representation.
Table 9Table 10
3D Grounding
2D Grounding
2D RES
AP3D@15 avg.
P@50 avg.
cIoU avg.
✓
84.0
–
–
✓
–
100.0
–
✓
–
–
80.3
✓
✓
✓
90.8
100.0
84.8
Table 9: Reconstruction Performance comparison of multi-task training.
Figure 5: The qualitative visualization of DVD on the 2D and 3D perception tasks. The first row is the results of 3D grounding task. The second row is the result of 2D grounding task. The third row is the results of 2D RES task. The fourth row is the results of unified perception.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Datasets
QAs
3D Grounding
SUN-RGBD Song et al. (2015) , Hypersim Roberts et al. (2021) , ARKitScenes Baruch et al. (2021) , Objectron Ahmadyan et al. (2021) , KITTI Geiger et al. (2012) , nuScenes Caesar et al. (2020)
489K
2D Grounding
RefCOCO Lin et al. (2014) , RefCOCO+ Lin et al. (2014) , RefCOCOg Mao et al. (2016) SUN-RGBD Song et al. (2015) , Hypersim Roberts et al. (2021) , ARKitScenes Baruch et al. (2021) , Objectron Ahmadyan et al. (2021) , KITTI Geiger et al. (2012) , nuScenes Caesar et al. (2020)
810K
2D RES
RefCOCO Lin et al. (2014) , RefCOCO+ Lin et al. (2014) , RefCOCOg Mao et al. (2016) , RefCLEF Kazemzadeh et al. (2014) , ReasonSeg Lai et al. (2024) , COCO Lin et al. (2014)
774K
Appendix
Table 10: Dataset Details.
Method
3D Grounding
2D Grounding
2D RES
tokens
latency (ms)
AP3D@15
tokens
latency (ms)
P@50
tokens
latency (ms)
cIoU
Text
74
4165
37.5
15
750
84.5
512
24410
48.0
Special Token
15
713
17.5
4
274
83.2
128
8771
42.9
DVD
8
499
34.1
2
143
87.1
64
1715
56.3
Appendix
Table 11: Computation cost and performance comparison of different representations.
Task
Tokens
Active Tokens
2D boxes
0.11M
2957/4096 (72.19%)
3D boxes
2.27M
4095/4096 (99.98%)
2D masks
4.26M
3940/4096 (96.19%)
Overall
6.65M
4095/4096 (99.98%)
Appendix
Table 12: The Codebook Utilization.
Figure 6: The visualization of latent space in DVD with t-SNE. (a) The vector latent space of 3D boxes. (b) The vector latent space of 3D boxes. (c) The vector latent space of 2D masks.
Figure 7: The visualization of latent space in DVD with t-SNE for unifying 2D and 3D perception tasks. There are there obvious clusters, illustrating the effectiveness of our dynamic vector representation learning process.
Figure 8: The qualitative visualization of DVD on the 3D grounding task.
Figure 9: The qualitative visualization of DVD on the 2D grounding task.
Figure 10: The qualitative visualization of DVD on the 2D RES task.
Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RGB-only design injects strong 3D inductive biases for fine-grained spatial understanding and reasoning without requiring any additional 3D inputs. Extensive experiments show that VLM-IE3D achieves superior performance consistently across various 3D tasks including 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. Code and models are available at https://github.com/Vegetebird/VLM-IE3D.
Wenhao Li, Xueying Jiang, Quanhao Qian +4
1Nanyang Technological University · 2DAMO Academy, Alibaba Group · 3HuPan Lab +1
Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatial consistency across video frames. Given the scarcity of large-scale 3D data, we present GeoVR, a novel framework that learns geometric representations using purely 2D video sequences. This approach effectively restructures the semantic latent space within MLLMs to unlock spatial intelligence. Rather than employing superficial feature mixing, GeoVR reshapes the internal representations of the MLLM by distilling geometry knowledge from pre-trained 3D foundation models. This is accomplished through a multi-objective learning strategy driven by four complementary geometric targets: (1) estimating inter-frame camera poses to embed varying viewpoint dynamics, (2) regressing dense depth maps to anchor physical distances, (3) predicting a metric scale factor for real-world calibration, and (4) distilling multi-scale 3D features to align the intermediate feature space. Guided by these explicit physical and geometric constraints, the model's internal representations naturally develop strong 3D awareness. Extensive experiments on spatial reasoning benchmarks demonstrate that GeoVR achieves state-of-the-art performance, establishing a new paradigm for endowing foundation models with spatial intelligence.