cs.CVOct 6, 2026

M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

Authors: Jinsong Zhang, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, Zhengguo Li

Organizations: School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China · Institute of Advanced Intelligence and Computation, Agency for Science, Technology and Research (A*STAR), Singapore

Abstract

Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 (δ<0.25δ< 0.25). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.

Figures & tables

Explore similar work

CardsList
  1. DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

    May 22, 2026Jie Zhu, Girish Chandar Ganesan, Xiaoming LiuMonocular Depth EstimationMonocular

  2. Grounded 3D-Aware Spatial Vision-Language Modeling

    May 28, 2026An-Chieh Cheng, Yang Fu, Yatai Ji +123D Visual GroundingSpatial Grounding

  3. AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

    May 25, 2026Cuong Huynh, Maxim Popov, Denis Gridusov +13D Visual Grounding3D Scene Understanding