cs.CVOct 8, 2026

DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

Authors: Jinghua Hou, Zhe Liu, Hengshuang Zhao

Organizations: The University of Hong Kong

Abstract

Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

    Aug 4, 2026Lucy Lin, Ayush Jain, Yifan Liu +13D Foundation Models3D VQA

  2. 3D-Aware VLMs with Implicit and Explicit Geometries

    Jul 23, 2026Wenhao Li, Xueying Jiang, Quanhao Qian +43D Spatial Reasoning3D Object Detection