M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding
Authors: Jinsong Zhang, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, Zhengguo Li
Organizations: School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China · Institute of Advanced Intelligence and Computation, Agency for Science, Technology and Research (A*STAR), Singapore
Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 (δ<0.25). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.
Figures & tables
Fig. 1: Overview of M3SunAgent. Given an RGB image and a natural-language query, the LLM selects the task and generates a corresponding DSL program, which is executed by the program interpreter using perception models and geometric operations. For instance-level metric depth estimation, point-wise depth predictions within the referred object are aggregated into an instance-level metric distance. For monocular 3D visual grounding, a VLM locates the target and predicts basic spatial attribute, which is combined with back-projection and a 2D-to-3D lifting model to construct the 3D bounding box. The result is then used to generate the final response.
Fig. 2: Overview of the instance-level metric depth estimation pipeline in M3SunAgent. LOCATE identifies the object referred to in the input description and returns its 2D bounding box. DEPTH normalizes the focal length and samples nine points within the box. It then uses a depth estimation model to predict the depth of these sampled points. AGGREGATE filters unreliable predictions and computes the median of the remaining depth estimates to obtain the target’s distance.
Fig. 3: Overview of the monocular 3D visual grounding pipeline in M3SunAgent. Given an RGB image and a target description, GROUND predicts the referred object’s 2D bounding box, the image-plane projection of its 3D center, and its depth. BACKPROJECT uses the projected center, depth, and camera intrinsics to recover the 3D center. LIFT estimates the object’s dimensions and yaw angle from the image, 2D box, and camera intrinsics. BUILD combines the recovered attributes to construct the final 3D bounding box.
Fig. 4: Examples of annotated instances in M3SI. Blue bounding boxes mark the target objects. The text beneath each image provides their descriptions and ground-truth distances from the camera in meters.
Method
Year
Valid/Total
MAE ↓
RMSE ↓
MedAE ↓
MRE (%) ↓
MedRE (%) ↓
δ<0.25 (%) ↑
General-Purpose VLMs
GPT-5
2025
1970/2910
3.31
5.07
1.80
39.22
34.43
24.33
Pixtral-12B [ 39 ]
2024
2810/2910
4.06
6.83
2.28
58.25
38.86
31.27
InternVL3-8B [ 40 ]
2025
2910/2910
4.60
6.96
2.84
51.68
53.06
16.36
InternVL3.5-8B [ 41 ]
2025
2910/2910
5.03
7.76
2.85
54.24
51.60
20.55
Qwen2.5-VL-3B [ 42 ]
2025
2910/2910
14.79
82.42
3.83
203.50
63.42
17.04
TABLE I: Comparison of instance-level metric depth estimation results on M3SI. The best and second-best results are shown in bold and underlined, respectively.
Fig. 5: Distribution of per-instance relative depth errors on M3SI. Each stacked bar shows the percentage of test instances in five mutually exclusive categories: errors below 10% , in [10%,20%) , [20%,25%) , or [25%,50%) , and errors of at least 50% or invalid predictions.
Method
Year
Acc@0.25
Acc@0.5
Unique
Multiple
Overall
Unique
Multiple
Overall
CV Methods
Cube R-CNN + Rand [ 2 ]
2024
32.76
13.36
17.02
14.61
7.21
8.60
Cube R-CNN + Best [ 2 ]
2024
35.29
60.52
55.77
16.67
32.99
29.92
ZSGNet + backproj [ 2 ]
2024
9.02
16.56
15.14
0.29
2.23
1.87
FAOA + backproj [ 2 ]
2024
11.96
13.79
13.44
2.06
2.12
2.11
TABLE II: Comparison of 3D visual grounding accuracy (%) on the Mono3DRefer [ 2 ] test set. The best and second-best results are shown in bold and underlined, respectively.
Method
Year
Near
Easy
Medium
Moderate
Acc@0.25
Acc@0.5
Acc@0.25
Acc@0.5
Acc@0.25
Acc@0.5
Acc@0.25
Acc@0.5
CV Methods
Cube R-CNN + Rand [ 2 ]
2024
17.40
11.45
21.12
11.41
18.01
8.15
17.85
8.01
Cube R-CNN + Best [ 2 ]
2024
67.76
41.45
59.66
33.05
60.69
30.35
60.56
33.45
ZSGNet + backproj [ 2 ]
2024
24.87
0.59
21.33
3.35
16.74
3.71
13.87
0.63
FAOA + backproj [ 2 ]
2024
18.03
0.53
17.51
3.43
15.64
3.95
12.18
1.34
TABLE III: Comparison of 3D visual grounding accuracy (%) on selected subsets of the Mono3DRefer test set. The best and second-best results are shown in bold and underlined, respectively.
Method
Year
mIoU (%) ↑
Object Occurrence
Distance
Difficulty
Unique
Multiple
Overall
Near
Medium
Far
Easy
Moderate
Hard
General-Purpose VLMs
MiMo-VL-7B [ 47 ]
2025
1.45
1.76
1.71
2.25
1.45
0.94
0.95
1.47
2.84
Qwen2.5-VL-72B [ 42 ]
2025
1.33
0.79
0.89
0.99
0.85
0.74
1.00
0.92
0.72
Gemini-2.5-Flash
2025
1.02
1.35
1.29
1.82
0.94
0.66
1.00
1.23
1.72
TABLE IV: Grounding mIoU (%) on the Mono3DRefer test set across object occurrence, distance, and difficulty subsets. The best and second-best results are shown in bold and underlined, respectively.
Fig. 7: Qualitative comparison on five representative samples from the Mono3DRefer test set. Each column shows a referring description and an input image, followed by the 3D bounding boxes predicted by GPT-5, Gemini-2.5-Pro, Gemini-2.5-Flash, Mono3DVG, and M3SunAgent. The boxes are shown as image projections and in 3D views. Blue indicates the ground truth, red indicates baseline predictions, and green indicates M3SunAgent predictions.
Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across diverse camera settings, such as perspective, fisheye, and panoramic images, remains challenging. Existing methods typically rely on a single depth estimator, overlooking that different models encode different camera assumptions and perform best under different input domains. In this paper, we show that depth experts exhibit strong sample-wise complementarity: model preference is highly correlated with camera geometry, and multi-model fusion brings the largest gains on difficult samples where individual experts are unreliable. Motivated by these observations, we propose \textbf{\ours}, a vision-language agent for adaptive monocular depth estimation. DepthAgent treats existing depth models as frozen tools and learns to analyze scene and camera cues, invoke suitable experts through multi-turn tool utilization, and select or fuse their predictions for each input. To optimize such discrete decision-making toward dense geometric quality, we design a multi-reward reinforcement fine-tuning scheme that jointly encourages valid tool execution, camera/scene analysis, expert-selection quality, and inference efficiency. Extensive experiments across perspective, fisheye, and panoramic benchmarks show that \ours consistently outperforms individual experts, fixed model fusion, and different selection strategies, with strong improvements on challenging samples, highlighting the critical role of expert selection and fusion. The code and model will be released upon publication.
Jie Zhu, Girish Chandar Ganesan, Xiaoming Liu
Michigan State University · University of North Carolina at Chapel Hill
We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introduces an implicit grounding mechanism that identifies entity mentions during generation and inserts the corresponding region tokens into the text stream, allowing the model to reference visual evidence on the fly when producing spatial chain-of-thought responses. In parallel, a region-prompted monocular 3D grounding design predicts 3D bounding boxes in the camera view from grounded region queries, supported by intrinsic-aware normalization and dense geometric supervision. Together, these grounding capabilities enable GR3D to decompose complex spatial understanding problems into grounded 2D perception followed by 3D inference. GR3D achieves consistent improvements across grounded and non-grounded spatial benchmarks, demonstrating grounding as an effective inductive bias for strengthening spatial understanding in VLMs. These grounding capabilities collectively enhance general spatial understanding beyond the grounding task itself.
3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero-shot methods leverage 2D vision-language models (LVLMs). However, they often rely on existing sets of multi-view images and struggle with the limited semantic and spatial details provided by standard 3D segmentation tools. We present AgentGrounder, a zero-shot 3D visual grounding framework that operates directly on colored point clouds without task-specific 3D training. Our approach follows a two-stage design: (1) an offline stage that applies 3D model to build an Object Lookup Table (OLT) with instance IDs, semantic labels, 3D bounding boxes; and (2) an online tool-driven agent that decomposes each query, retrieves only relevant candidates from the OLT, performs geometric scoring, and triggers image rendering on demand when additional visual evidence (e.g., color, material, or viewpoint-sensitive cues) is required. Compared with fixed anchor-target matching pipelines, this design reduces cascading matching errors and improves context-window efficiency by avoiding prompts overloaded with irrelevant objects. We evaluate on ScanRefer and Nr3D under a zero-shot setting and observe consistent improvements over SeeGround in our setup, including +2.5% Acc@0.5 on ScanRefer and +6.3% on Nr3D, with a notable +6.3% gain on Nr3D view-independent queries. These results show that combining selective retrieval, geometric reasoning, and adaptive visual inspection yields a practical and robust foundation for open-vocabulary 3D grounding. Our code is available at https://github.com/be2rlab/AgentGrounder.
Cuong Huynh, Maxim Popov, Denis Gridusov +1
Biomechatronics and Energy-Efficient Robotics (BE2R) Lab, ITMO University, Saint Petersburg, Russia