DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
Organizations: The University of Hong Kong
Abstract
Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.
Figures & tables
| Method | 3D Grounding | 2D Grounding | 2D RES | ||||||
| SUN-RGBD | Hypersim | nuScenes | RefCOCO | RefCOCO+ | RefCOCOg | RefCOCO | RefCOCO+ | RefCOCOg | |
| GroundingDINO Liu et al. (2024a) | – | – | – | 90.6 | 88.2 | 86.1 | – | – | – |
| VistaLLM-7B Pramanick et al. (2024) | – | – | – | 88.1 | 82.9 | 83.6 | 74.5 | 69.1 | 69.0 |
| LISA-7B Lai et al. (2024) | – | – | – | – | – | – | 74.9 | 65.1 | 67.9 |
| Text4Seg Lan et al. (2025) | – | – | – | 88.3 | 83.5 | 82.4 | 74.7 | 68.5 | 70.7 |
| Qwen2.5-VL-3B Bai et al. (2025b) | – | – | – | 89.1 | 82.4 | 85.2 | – | – | – |
| Method | SUN-RGBD | Hypersim | ARKitScenes | KITTI | nuScenes |
| Gemini 2.0 Pro Team et al. (2024) | 32.5 | – | – | – | – |
| Gemini 2.5 Pro Comanici et al. (2025) | 29.7 | – | – | – | – |
| Seed1.5-VL Guo et al. (2025) | 33.5 | – | – | – | – |
| Qwen3-VL-2B Bai et al. (2025a) | 33.8 | 12.0 | – | – | – |
| Qwen3-VL-8B Bai et al. (2025a) | 36.2 | 12.7 | – | – | – |
| VST-3B Yang et al. (2025c) | 37.3 | – | 51.7 | – | – |
| Method | RefCOCO | RefCOCO+ | RefCOCOg | |||||
| val | testA | testB | val | testA | testB | val | test | |
| GroundingDINO Liu et al. (2024a) | 90.6 | 93.2 | 88.2 | 88.2 | 89.0 | 75.9 | 86.1 | 87.0 |
| Qwen2.5-VL-7B Bai et al. (2025b) | 90.0 | 92.5 | 85.4 | 84.2 | 89.1 | 76.9 | 87.2 | 87.2 |
| Rex-Omni Jiang et al. (2026) | – | – | – | – | – | – | 86.6 | 86.8 |
| Seed1.5-VL Guo et al. (2025) | – | – | – | – | – | – | 84.7 | 85.2 |
| InternVL3.5-8B Wang et al. (2025b) | 92.4 | 94.7 | 88.7 | 87.9 | 92.4 | 82.4 | 89.6 | 89.4 |
| Method | RefCOCO | RefCOCO+ | RefCOCOg | |||||
| val | testA | testB | val | testA | testB | val | test | |
| VLT Ding et al. (2021) | 67.5 | 70.5 | 65.2 | 56.3 | 61.0 | 50.1 | 55.0 | 57.7 |
| LAVT Yang et al. (2022) | 72.7 | 75.8 | 68.8 | 62.1 | 68.4 | 55.1 | 61.2 | 62.1 |
| VistaLLM-7B Pramanick et al. (2024) | 74.5 | 76.0 | 72.7 | 69.1 | 73.7 | 64.0 | 69.0 | 70.9 |
| LISA-7B Lai et al. (2024) | 74.9 | 79.1 | 72.3 | 65.1 | 70.8 | 58.1 | 67.9 | 70.6 |
| M 2 SA Jang et al. (2025) | 74.0 | 76.8 | 69.7 | 63.1 | 67.2 | 56.1 | 67.0 | 68.3 |
| 3D Grounding | 2D Grounding | 2D RES | AP3D@15 avg. | P@50 avg. | cIoU avg. |
| ✓ | 84.0 | – | – | ||
| ✓ | – | 100.0 | – | ||
| ✓ | – | – | 80.3 | ||
| ✓ | ✓ | ✓ | 90.8 | 100.0 | 84.8 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Datasets | QAs |
| 3D Grounding | SUN-RGBD Song et al. (2015) , Hypersim Roberts et al. (2021) , ARKitScenes Baruch et al. (2021) , Objectron Ahmadyan et al. (2021) , KITTI Geiger et al. (2012) , nuScenes Caesar et al. (2020) | 489K |
| 2D Grounding | RefCOCO Lin et al. (2014) , RefCOCO+ Lin et al. (2014) , RefCOCOg Mao et al. (2016) SUN-RGBD Song et al. (2015) , Hypersim Roberts et al. (2021) , ARKitScenes Baruch et al. (2021) , Objectron Ahmadyan et al. (2021) , KITTI Geiger et al. (2012) , nuScenes Caesar et al. (2020) | 810K |
| 2D RES | RefCOCO Lin et al. (2014) , RefCOCO+ Lin et al. (2014) , RefCOCOg Mao et al. (2016) , RefCLEF Kazemzadeh et al. (2014) , ReasonSeg Lai et al. (2024) , COCO Lin et al. (2014) | 774K |
| Method | 3D Grounding | 2D Grounding | 2D RES | ||||||
| tokens | latency (ms) | AP3D@15 | tokens | latency (ms) | P@50 | tokens | latency (ms) | cIoU | |
| Text | 74 | 4165 | 37.5 | 15 | 750 | 84.5 | 512 | 24410 | 48.0 |
| Special Token | 15 | 713 | 17.5 | 4 | 274 | 83.2 | 128 | 8771 | 42.9 |
| DVD | 8 | 499 | 34.1 | 2 | 143 | 87.1 | 64 | 1715 | 56.3 |
| Task | Tokens | Active Tokens |
| 2D boxes | 0.11M | 2957/4096 (72.19%) |
| 3D boxes | 2.27M | 4095/4096 (99.98%) |
| 2D masks | 4.26M | 3940/4096 (96.19%) |
| Overall | 6.65M | 4095/4096 (99.98%) |