cs.CVOct 7, 2026

Spatial Latent Reasoning for Embodied Reference Understanding

Authors: Ling Li, Jianhui Zhong, Wei Liu, Zheng Jiang aand Yuxuan Liu, Jingyu Li, Zhidong Deng

Organizations: Tsinghua University · Dalian University of Technology · University of Science and Technology of China

Abstract

Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

    Jun 23, 2026Ling Li, Bowen Liu, Zinuo Zhan +5Visual Grounding BenchmarksMultimodal Large Language Models

  2. Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision

    Apr 23, 2026Chentao Li, Zirui Gao, Mingze Gao +3Egocentric VisionMultimodal Large Language Models

  3. GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

    Oct 1, 2026Yakun Zhu, Yi Bin, Yujuan Ding +53D Spatial ReasoningDiscriminative Congruence Transform