GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.
Figures & tables
Figure 1: Conceptual comparison of grounding paradigms. ReLaViS performs GUI-aware coarse-to-fine visual search by updating explicit spatial search states within a single model interaction.
Figure 2: Inference pipeline of ReLaViS. At step t , the spatial search head attends to visual tokens V using query Qt derived from the hidden state, producing the search distribution st . This distribution represents the search focus and aggregates V into latent visual evidence zt=st⊤V , which serves as the next input embedding to condition subsequent search. After T steps, a deterministic readout maps sT to the click coordinate. Within a single autoregressive sequence, the search focus evolves from the global interface through element groups to the target, as illustrated by the top heatmaps.
Figure 3: Left : pooled accuracy gains of ReLaViS-7B over single-step baselines across target area bins; stacked bars show instance counts and benchmark composition. Right : accuracy changes relative to predicted-state inference when target-aligned or target-opposite spatial search states are injected at t=1,2,3 individually or jointly. Curves show multi-seed means with std shading.
Figure 4: Stepwise search dynamics on ScreenSpot-Pro. Metric definitions are in Appendix B.5 . Dashed lines indicate single-step references; curves show multi-seed means with std shading.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Overview of trajectory construction and training. Top : ReLaViS converts flat GUI element annotations E into a search trajectory C1:T and maps each candidate set Ct to a target distribution dt over the visual token grid. Bottom : Stage 1 trains only the spatial search head with a frozen VLM and T=1 ; Stage 2 uses teacher-forced inputs ztgt=dt⊤V ; and Stage 3 uses self-conditioned inputs zt=st⊤V from a detached rollout. LCWC supervises all search steps, while LFGP additionally supervises the final step. “SS head” denotes the spatial search head.
Figure 6: Constructed GUI-aware search trajectories with T=4 across representative applications. The sequence above each panel reports ∣C1∣→⋯→∣C4∣ , while blue, yellow, red, and green bounding boxes mark candidate elements in C1∖C2 , C2∖C3 , C3∖C4 , and C4={e⋆} , respectively. The candidate sets contract from the full element set E through intermediate groups to the target.
Figure 7: Stepwise diagnostic LFGP (a) and LCWC (b) on the 5,000-sample GroundCUA validation set. Curves show multi-seed means with std shading.
Figure 8: Successful cases from ScreenSpot-Pro (panel (a)) and ScreenSpot-v2 (panels (b) and (c)).
Figure 9: Successful cases from MMBench-GUI-L2 (panel (a)), OSWorld-G-Refine (panel (b)), and UI-Vision (panel (c)).
Figure 10: Failure cases from ScreenSpot-Pro (panels (a) and (b)) and ScreenSpot-v2 (panel (c)).