GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.
Figures & tables
Figure 1: Conceptual comparison of grounding paradigms. ReLaViS performs GUI-aware coarse-to-fine visual search by updating explicit spatial search states within a single model interaction.
Figure 2: Inference pipeline of ReLaViS. At step t , the spatial search head attends to visual tokens V using query Qt derived from the hidden state, producing the search distribution st . This distribution represents the search focus and aggregates V into latent visual evidence zt=st⊤V , which serves as the next input embedding to condition subsequent search. After T steps, a deterministic readout maps sT to the click coordinate. Within a single autoregressive sequence, the search focus evolves from the global interface through element groups to the target, as illustrated by the top heatmaps.
Figure 3: Left : pooled accuracy gains of ReLaViS-7B over single-step baselines across target area bins; stacked bars show instance counts and benchmark composition. Right : accuracy changes relative to predicted-state inference when target-aligned or target-opposite spatial search states are injected at t=1,2,3 individually or jointly. Curves show multi-seed means with std shading.
Figure 4: Stepwise search dynamics on ScreenSpot-Pro. Metric definitions are in Appendix B.5 . Dashed lines indicate single-step references; curves show multi-seed means with std shading.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Overview of trajectory construction and training. Top : ReLaViS converts flat GUI element annotations E into a search trajectory C1:T and maps each candidate set Ct to a target distribution dt over the visual token grid. Bottom : Stage 1 trains only the spatial search head with a frozen VLM and T=1 ; Stage 2 uses teacher-forced inputs ztgt=dt⊤V ; and Stage 3 uses self-conditioned inputs zt=st⊤V from a detached rollout. LCWC supervises all search steps, while LFGP additionally supervises the final step. “SS head” denotes the spatial search head.
Figure 6: Constructed GUI-aware search trajectories with T=4 across representative applications. The sequence above each panel reports ∣C1∣→⋯→∣C4∣ , while blue, yellow, red, and green bounding boxes mark candidate elements in C1∖C2 , C2∖C3 , C3∖C4 , and C4={e⋆} , respectively. The candidate sets contract from the full element set E through intermediate groups to the target.
Figure 7: Stepwise diagnostic LFGP (a) and LCWC (b) on the 5,000-sample GroundCUA validation set. Curves show multi-seed means with std shading.
Figure 8: Successful cases from ScreenSpot-Pro (panel (a)) and ScreenSpot-v2 (panels (b) and (c)).
Figure 9: Successful cases from MMBench-GUI-L2 (panel (a)), OSWorld-G-Refine (panel (b)), and UI-Vision (panel (c)).
Figure 10: Failure cases from ScreenSpot-Pro (panels (a) and (b)) and ScreenSpot-v2 (panel (c)).
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated interfaces because a vision-language model (VLM) may recognize a requested control without locating it precisely enough for interaction. Most existing methods provide various forms of localization assistance, but still rely on a direct click prediction, allowing visual ambiguity or an inaccurate initial estimate to propagate to the final result. In this paper, we introduce GUI-Lens, a coarse-to-fine grounding framework that allows a general-purpose VLM to determine the target through active visual observations. Specifically, GUI-Lens extracts OCR text and detected UI components from the screenshot and presents their positions as coordinate references. Using the instruction, the current view, and these references, the VLM selects the region and scale of the next view, which is cropped and enlarged to provide finer visual details. This process continues over successively focused views until the target is determined. Proposed crops and clicks are checked against the instruction throughout the process, and the final local position is mapped back to the original screen coordinates. Experiments on four GUI grounding benchmarks and three general-purpose VLM backends show that GUI-Lens improves overall grounding accuracy by up to 24.9 percentage points and achieves state-of-the-art performance with GPT-5.5.
Zichuan Fu, Shirong Wang, Wenlin Zhang +10
City University of Hong Kong · Westlake University · Tencent Jarvis Lab
Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance degrades in high-resolution interfaces, where dense layouts and small interactive elements expose a resolution gap between modern displays and model input constraints. Existing zoom-in strategies rely on fixed anchors, heuristic grids, or reinforcement learning, lacking a principled mechanism to adaptively determine where refinement is needed and how much spatial uncertainty should be explored. We propose AutoFocus, a training-free, uncertainty-aware active visual search framework for GUI grounding. Our key insight is that token-level perplexity in coordinate generation naturally reflects spatial uncertainty. Rather than committing to a single prediction, AutoFocus samples multiple coordinate hypotheses and converts their axial perplexities into an anisotropic gaussian spatial probability field, explicitly modeling directional uncertainty. Based on this field, we generate global and local region proposals and introduce Shape-Aware Zooming to balance tight localization with contextual preservation. A visual prompt-based aggregation step then selects the most consistent prediction via structured comparison. Extensive experiments on ScreenSpot-Pro and ScreenSpot-V2 demonstrate consistent improvements across both general-purpose and GUI-specialized VLMs.
GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions. However, accurately grounding instruction-relevant elements from high-resolution screenshots cluttered with irrelevant UI components remains challenging for existing approaches. Inspired by how humans dynamically adjust their perceptual scope to locate task-related regions on complex screens, we propose DRS-GUI, a training-free dynamic region search framework for GUI grounding that can be seamlessly integrated into existing MLLMs. DRS-GUI introduces a lightweight UI Perceptor that performs three human-like perceptual actions (Focus, Shift, and Scatter) to progressively explore the interface and generate region proposals. To dynamically schedule these actions, we further design an Action Planner based on Monte Carlo Tree Search (MCTS). A region quality reward is employed to evaluate and select the highly instruction-relevant region, efficiently pruning redundant UI elements. Experiments demonstrate that DRS-GUI yields a 14% improvement on ScreenSpot-Pro for general and GUI-specific MLLMs (Qwen2.5-VL-7B and UGround-V1-7B), significantly enhancing grounding performance and generalization.
Yichao Liu, Huawen Shen, Liu Yu +3
1Nankai University · Institute of Information Engineering, Chinese Academy of Sciences