Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.
Figures & tables
Figure 1: Comparison of reasoning approaches on an EgoPoint-Ground example ( Li et al., 2026c ) . (a) Tool-assisted reasoning crops and re-encodes image regions during text generation ( Zheng et al., 2025 ) . (b) LVR ( Li et al., 2026a ) supervises continuous states by reconstructing annotated region features. (c) SLR combines one spatial ray state with four states supervised by parity-pooled target features. Image crops and ray overlays illustrate the reasoning operations and supervision targets; generic latent boundary markers are omitted for clarity.
Figure 2: Overview of SLR. (a) A frozen visual encoder and merger provide image features. The causal MLLM generates five recurrent states in its hidden space, then decodes the target box; s and e denote latent boundary markers. (b) A spatial readout predicts fingertip position and pointing direction, while visual states align with parity-pooled region features with stopped gradients. Auxiliary readouts and annotations are used only for training. The ray overlays and grid are schematic; readout dimensions correspond to Qwen3.5-4B.
Figure 3: Global-grid parity pooling. The target ROI spans rows and columns 1–6 of an illustrative 8×8 grid. Each branch averages nine feature vectors selected by global row and column parity and aligns the detached mean with its generated state hab=z2+2a+b . Dashed arrows connect both inputs to the cosine loss; pooled features serve only as training targets. The inset compares contiguous quadrants with interleaved parity groups.
Model
Method
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Human Study
Human Participants
0.922
0.909
0.830
0.805
Qwen2.5-VL-7B
Zero-shot
0.650
0.589
0.500
0.546
Standard SFT
0.644
0.589
0.489
0.538
Text CoT ( Wei et al., 2022 ) (adapted)
0.744
0.661
0.494
0.575
PointVG-R ( Li et al., 2026d )
0.833
0.761
0.639
0.670
SLR
0.839
0.806
0.706
0.713
Table 1: Results on EgoPoint-Ground . SLR uses one spatial and four parity-supervised visual states. Shading marks SLR; bold denotes the best grounding scores per backbone, including ties. Human scores provide a reference; token counts and times appear in Table 9 in Appendix E .
Model
Method
Hard-Similar
Hard-Complex
P@0.5 ↑
mIoU ↑
P@0.5 ↑
mIoU ↑
Qwen2.5-VL-7B
Standard SFT
0.229
0.202
0.458
0.389
Text CoT ( Wei et al., 2022 ) (adapted)
0.188
0.198
0.448
0.411
PointVG-R ( Li et al., 2026d )
0.313
0.276
0.573
0.473
SLR
0.323
0.327
0.583
0.511
Qwen3-VL-8B
Standard SFT
0.281
0.290
0.385
0.349
Table 2: Key comparisons on the two EgoPoint-Ground hard subsets. Shading identifies SLR; bold scores mark the best reported result within each backbone and subset, including ties. Table 8 in Appendix D reports all methods, localization metrics, token counts, and times.
Method
Reported Metric
@0.25 ↑
@0.50 ↑
@0.75 ↑
Human Participants ( Chen et al., 2021 )
Accuracy
0.942
0.858
0.533
YouRefIt Full ( Chen et al., 2021 )
Accuracy
0.547
0.405
0.140
REP ( Shi and Yang, 2022 )
Precision
0.588
0.457
0.188
Touch-Line (VTL) ( Li et al., 2023 )
Precision
0.711
0.635
0.390
AD-DINO (ADTL) ( Guo et al., 2024 )
Accuracy
0.763
0.724
0.554
CAPE ( Eyiokur et al., 2026b )
mAP
0.750
0.654
0.357
Table 3: Results on YouRefIt . Published baselines use their original metrics and evaluation protocols. Shading identifies SLR; the zero-shot model uses the same backbone.
Figure 4: Grounding comparisons for two pointing examples. Columns show Zero-shot, Standard SFT, Text CoT, PointVG-R, and SLR (labeled “Ours” in the supplied visualization). Green and red boxes indicate the annotated targets and predicted regions, respectively, for each of the five compared methods.
Backbone
Variant
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
EgoPoint-Ground
Qwen3.5-4B
Spatial + Parity (SLR)
0.906
0.878
0.811
0.778
Spatial-only
0.867
0.844
0.767
0.747
Target-only
0.883
0.856
0.783
0.764
SFT (no latent states)
0.861
0.839
0.761
0.750
Hard-Similar: multiple visually similar objects
Table 4: Component ablations on Qwen3.5-4B . Shading marks the full model (one spatial and four target states); bold marks the best grounding scores per subset, including ties. Spatial-only retains the ray state; Target-only retains the four visual states; SFT uses no latent segment.
Backbone
Pooling
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Qwen3.5-4B
Parity
0.883
0.856
0.783
0.764
Quadrant
0.839
0.817
0.744
0.728
Random
0.844
0.822
0.756
0.735
Max
0.850
0.833
0.756
0.746
Table 5: Pooling-operator ablations on EgoPoint-Ground with Qwen3.5-4B, four target latents, and no spatial state. Shading identifies parity pooling; bold scores mark the best grounding result, including ties. Appendix B.4 defines the operators.
Figure 5: Attention across five recurrent latent steps. Left: image-token attention for one pointing example for the spatial state hsp and target states h00,h01,h10,h11 , averaged over all 16 heads in layers 20, 24, 28, and 32 and normalized over image tokens using a shared color scale. Green and red boxes mark the ground-truth and predicted target, respectively; the dashed yellow box marks the ground-truth hand region. Right: mean attention mass assigned to context, image tokens, and previous latent states over all evaluation samples.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Reference value
Backbone / hidden dimension
Qwen3.5-4B / 2560
Spatial / target states
1 / 4
Image pixel limits
50,176 to 262,144
Optimizer / learning rate
AdamW / 10−5
Weight decay / seed
0.01 / 0
AdamW (β1,β2) / ϵ
(0.9,0.999) / 10−8
Appendix
Table 6: Reference training configuration for Qwen3.5-4B with a four-epoch training budget. These settings describe the reference implementation rather than all baseline runs.
Backbone
Pooling latents
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Tokens †
Time (s) ↓
EgoPoint-Ground
Qwen3.5-4B
4 (reference)
0.883
0.856
0.783
0.764
6.000
1.766
1
0.867
0.844
0.794
0.761
3.000
1.258
9
0.817
0.789
0.744
0.712
11.000
2.517
16
0.867
0.844
0.789
0.757
18.000
3.578
Appendix
Table 7: Ablation of target-state count on EgoPoint-Ground with Qwen3.5-4B and no spatial state. Shading marks the four-state pooling reference; the full model uses five states. Bold values denote the best grounding scores, including ties. See Appendix B.3 for token-count ( † ) conventions.
Figure 6: Effect of target-state count on EgoPoint-Ground mIoU with Qwen3.5-4B and no spatial state. Points correspond to Table 7 ; connecting lines aid visual comparison. Shading marks the four-state reference, and the bold value denotes the highest mIoU. Results are point estimates for the evaluated configurations, with no uncertainty intervals reported.
Model
Method
P@0.3 ↑
P@0.5 ↑
P@0.7 ↑
mIoU ↑
Tokens †
Time (s) ↓
Hard-Similar: multiple visually similar objects
Human Study
Human Participants
0.965
0.948
0.917
0.856
N/A
N/A
Qwen2.5-VL-7B
Zero-shot
0.281
0.240
0.167
0.214
0.000
0.864
Standard SFT
0.281
0.229
0.146
0.202
0.000
0.576
Text CoT (adapted)
0.292
0.188
0.063
0.198
166.854
3.312
PointVG-R
0.354
0.313
0.177
0.276
175.521
3.511
Appendix
Table 8: Full results on the EgoPoint-Ground Hard-Similar and Hard-Complex subsets. Shading identifies SLR; bold values denote the best grounding scores within each backbone and subset, including ties. Dashes indicate unreported values; N/A indicates inapplicable entries.
Backbone
Method / variant
Evaluation set
Tokens †
Time (s)
Main comparisons
Qwen2.5-VL-7B
Zero-shot
Standard
0.000
1.181
Standard SFT
Standard
0.000
0.885
Text CoT (adapted)
Standard
173.533
4.402
PointVG-R
Standard
183.622
4.708
SLR
Standard
7.000
1.362
Appendix
Table 9: Reported token counts and elapsed times for Tables 1 , 4 , and 5 . Appendix B.3 details token-count ( † ) and timing conventions. Standard denotes EgoPoint-Ground.