Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.
Figures & tables
Figure 1: Comparison of reasoning approaches on an EgoPoint-Ground example ( Li et al., 2026c ) . (a) Tool-assisted reasoning crops and re-encodes image regions during text generation ( Zheng et al., 2025 ) . (b) LVR ( Li et al., 2026a ) supervises continuous states by reconstructing annotated region features. (c) SLR combines one spatial ray state with four states supervised by parity-pooled target features. Image crops and ray overlays illustrate the reasoning operations and supervision targets; generic latent boundary markers are omitted for clarity.
Figure 2: Overview of SLR. (a) A frozen visual encoder and merger provide image features. The causal MLLM generates five recurrent states in its hidden space, then decodes the target box; s and e denote latent boundary markers. (b) A spatial readout predicts fingertip position and pointing direction, while visual states align with parity-pooled region features with stopped gradients. Auxiliary readouts and annotations are used only for training. The ray overlays and grid are schematic; readout dimensions correspond to Qwen3.5-4B.
Figure 3: Global-grid parity pooling. The target ROI spans rows and columns 1–6 of an illustrative 8×8 grid. Each branch averages nine feature vectors selected by global row and column parity and aligns the detached mean with its generated state hab=z2+2a+b . Dashed arrows connect both inputs to the cosine loss; pooled features serve only as training targets. The inset compares contiguous quadrants with interleaved parity groups.
Model
Method
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Human Study
Human Participants
0.922
0.909
0.830
0.805
Qwen2.5-VL-7B
Zero-shot
0.650
0.589
0.500
0.546
Standard SFT
0.644
0.589
0.489
0.538
Text CoT ( Wei et al., 2022 ) (adapted)
0.744
0.661
0.494
0.575
PointVG-R ( Li et al., 2026d )
0.833
0.761
0.639
0.670
SLR
0.839
0.806
0.706
0.713
Table 1: Results on EgoPoint-Ground . SLR uses one spatial and four parity-supervised visual states. Shading marks SLR; bold denotes the best grounding scores per backbone, including ties. Human scores provide a reference; token counts and times appear in Table 9 in Appendix E .
Model
Method
Hard-Similar
Hard-Complex
P@0.5 ↑
mIoU ↑
P@0.5 ↑
mIoU ↑
Qwen2.5-VL-7B
Standard SFT
0.229
0.202
0.458
0.389
Text CoT ( Wei et al., 2022 ) (adapted)
0.188
0.198
0.448
0.411
PointVG-R ( Li et al., 2026d )
0.313
0.276
0.573
0.473
SLR
0.323
0.327
0.583
0.511
Qwen3-VL-8B
Standard SFT
0.281
0.290
0.385
0.349
Table 2: Key comparisons on the two EgoPoint-Ground hard subsets. Shading identifies SLR; bold scores mark the best reported result within each backbone and subset, including ties. Table 8 in Appendix D reports all methods, localization metrics, token counts, and times.
Method
Reported Metric
@0.25 ↑
@0.50 ↑
@0.75 ↑
Human Participants ( Chen et al., 2021 )
Accuracy
0.942
0.858
0.533
YouRefIt Full ( Chen et al., 2021 )
Accuracy
0.547
0.405
0.140
REP ( Shi and Yang, 2022 )
Precision
0.588
0.457
0.188
Touch-Line (VTL) ( Li et al., 2023 )
Precision
0.711
0.635
0.390
AD-DINO (ADTL) ( Guo et al., 2024 )
Accuracy
0.763
0.724
0.554
CAPE ( Eyiokur et al., 2026b )
mAP
0.750
0.654
0.357
Table 3: Results on YouRefIt . Published baselines use their original metrics and evaluation protocols. Shading identifies SLR; the zero-shot model uses the same backbone.
Figure 4: Grounding comparisons for two pointing examples. Columns show Zero-shot, Standard SFT, Text CoT, PointVG-R, and SLR (labeled “Ours” in the supplied visualization). Green and red boxes indicate the annotated targets and predicted regions, respectively, for each of the five compared methods.
Backbone
Variant
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
EgoPoint-Ground
Qwen3.5-4B
Spatial + Parity (SLR)
0.906
0.878
0.811
0.778
Spatial-only
0.867
0.844
0.767
0.747
Target-only
0.883
0.856
0.783
0.764
SFT (no latent states)
0.861
0.839
0.761
0.750
Hard-Similar: multiple visually similar objects
Table 4: Component ablations on Qwen3.5-4B . Shading marks the full model (one spatial and four target states); bold marks the best grounding scores per subset, including ties. Spatial-only retains the ray state; Target-only retains the four visual states; SFT uses no latent segment.
Backbone
Pooling
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Qwen3.5-4B
Parity
0.883
0.856
0.783
0.764
Quadrant
0.839
0.817
0.744
0.728
Random
0.844
0.822
0.756
0.735
Max
0.850
0.833
0.756
0.746
Table 5: Pooling-operator ablations on EgoPoint-Ground with Qwen3.5-4B, four target latents, and no spatial state. Shading identifies parity pooling; bold scores mark the best grounding result, including ties. Appendix B.4 defines the operators.
Figure 5: Attention across five recurrent latent steps. Left: image-token attention for one pointing example for the spatial state hsp and target states h00,h01,h10,h11 , averaged over all 16 heads in layers 20, 24, 28, and 32 and normalized over image tokens using a shared color scale. Green and red boxes mark the ground-truth and predicted target, respectively; the dashed yellow box marks the ground-truth hand region. Right: mean attention mass assigned to context, image tokens, and previous latent states over all evaluation samples.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Reference value
Backbone / hidden dimension
Qwen3.5-4B / 2560
Spatial / target states
1 / 4
Image pixel limits
50,176 to 262,144
Optimizer / learning rate
AdamW / 10−5
Weight decay / seed
0.01 / 0
AdamW (β1,β2) / ϵ
(0.9,0.999) / 10−8
Appendix
Table 6: Reference training configuration for Qwen3.5-4B with a four-epoch training budget. These settings describe the reference implementation rather than all baseline runs.
Backbone
Pooling latents
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Tokens †
Time (s) ↓
EgoPoint-Ground
Qwen3.5-4B
4 (reference)
0.883
0.856
0.783
0.764
6.000
1.766
1
0.867
0.844
0.794
0.761
3.000
1.258
9
0.817
0.789
0.744
0.712
11.000
2.517
16
0.867
0.844
0.789
0.757
18.000
3.578
Appendix
Table 7: Ablation of target-state count on EgoPoint-Ground with Qwen3.5-4B and no spatial state. Shading marks the four-state pooling reference; the full model uses five states. Bold values denote the best grounding scores, including ties. See Appendix B.3 for token-count ( † ) conventions.
Figure 6: Effect of target-state count on EgoPoint-Ground mIoU with Qwen3.5-4B and no spatial state. Points correspond to Table 7 ; connecting lines aid visual comparison. Shading marks the four-state reference, and the bold value denotes the highest mIoU. Results are point estimates for the evaluated configurations, with no uncertainty intervals reported.
Model
Method
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Tokens †
Time (s) ↓
Hard-Similar: multiple visually similar objects
Human Study
Human Participants
0.965
0.948
0.917
0.856
N/A
N/A
Qwen2.5-VL-7B
Zero-shot
0.281
0.240
0.167
0.214
0.000
0.864
Standard SFT
0.281
0.229
0.146
0.202
0.000
0.576
Text CoT (adapted)
0.292
0.188
0.063
0.198
166.854
3.312
PointVG-R
0.354
0.313
0.177
0.276
175.521
3.511
Appendix
Table 8: Full results on the EgoPoint-Ground Hard-Similar and Hard-Complex subsets. Shading identifies SLR; bold values denote the best grounding scores within each backbone and subset, including ties. Dashes indicate unreported values; N/A indicates inapplicable entries.
Backbone
Method / variant
Evaluation set
Tokens †
Time (s)
Main comparisons
Qwen2.5-VL-7B
Zero-shot
Standard
0.000
1.181
Standard SFT
Standard
0.000
0.885
Text CoT (adapted)
Standard
173.533
4.402
PointVG-R
Standard
183.622
4.708
SLR
Standard
7.000
1.362
Appendix
Table 9: Reported token counts and elapsed times for Tables 1 , 4 , and 5 . Appendix B.3 details token-count ( † ) and timing conventions. Standard denotes EgoPoint-Ground.
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input images into static feature representations and perform reasoning primarily within the linguistic domain, often overlooking the rich perceptual cues and explicit spatial geometry inherent in images. In this study, we aim to mitigate the cognitive vulnerability of models in interpreting gestural spatial relations by proposing PointVG-R, a reasoning-guided Multi-modal Large Language Model (MLLM). PointVG-R introduces geometric-aware reasoning for pointing-based grounding, enabling the model to think with images through the strategic integration of Reinforcement Learning (RL) and cold-start data. Specifically, we design a novel geometric reasoning pipeline that simulates the iterative cognitive process humans employ when interpreting pointing gestures. Furthermore, we construct EgoPoint-CoT, a high-quality visual Chain-of-Thought (CoT) dataset featuring detailed reasoning trajectories to guide the model via Supervised Fine-Tuning (SFT) and RL. To address the varying quality of learning signals encountered during training, we further propose an Adaptive Importance Weighting strategy based on Group Variance, which dynamically adjusts reward signals to optimize the learning process. Experimental results demonstrate that PointVG-R achieves SOTA performance, outperforming the baseline by 15.86 points in mIoU. Extensive ablation studies further validate the efficacy of our proposed modules. Code: https://github.com/lingli1724/PointVG-R.
Ling Li, Bowen Liu, Zinuo Zhan +5
Tsinghua University Beijing Shi, China · Dalian University of Technology Dalian Shi, China · Northwestern Polytechnical University Xi’an Shi, China +1
Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to precisely ground the spatial semantics of pointing. Instead, they rely on spurious correlations with visual proximity or object saliency, a phenomenon we term "Referential Hallucination." To address this gap, we introduce EgoPoint-Bench, a comprehensive question-answering benchmark designed to evaluate and enhance multimodal pointing reasoning in egocentric views. Comprising over 11k high-fidelity simulated and real-world samples, the benchmark spans five evaluation dimensions and three levels of referential complexity. Extensive experiments demonstrate that while state-of-the-art proprietary and open-source models struggle with egocentric pointing, models fine-tuned on our synthetic data achieve significant performance gains and robust sim-to-real generalization. This work highlights the importance of spatially aware supervision and offers a scalable path toward precise egocentric AI assistants. Project page: https://guyyyug.github.io/EgoPoint-Bench/
Chentao Li, Zirui Gao, Mingze Gao +3
Department of Automation, Tsinghua University · Academy of Art & Design, Tsinghua University
Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
Yakun Zhu, Yi Bin, Yujuan Ding +5
Tongji University · The Hong Kong Polytechnic University