Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history (∼3×) at only ∼1.2× the peak retained input context of action-only history.
Figures & tables
Figure 1: Passive grounding vs. active visual probing. (a) Screenshot-coordinate policies infer actions from raw pixels, discarding page structure; (b) static DOM-SoM overlays all marks at once, cluttering dense pages and forcing one-shot alignment; (c) our Probe to Act interleaves reason – probe loops that locally ground DOM cues on demand and retain only verified evidence in memory before acting.
Figure 2: Overview of Probe to Act. At each step the agent observes the screenshot vt and DOM dt , then enters a state-preserving Reason–Probe–Operate loop: it reasons, selects a probing primitive, and applies it to the current observation, repeating until ready to emit a state-changing browser operation at . Verified rows are distilled into a sparse evidence memory Mt that persists across steps.
Figure 3: Illustration of the four probing primitives.
Table 4Figure 5
Model
SR (%)
Ops
Images
Time
Qwen3-VL-8B-Instruct
24.6
15.84
16.04
203.2 s
+ P2A (Round 2)
32.9
11.25
16.90
243.8 s
+ P2A (Prompt)
18.0
19.24
20.83
338.6 s
Gemini-3-Flash
42.4
12.80
13.20
380.1 s
+ P2A (Prompt)
48.7
13.08
20.70
432.8 s
Table 5: Representative average per-task inference accounting on VWA. “SR” denotes task success rate, and “Ops” counts state-changing browser operations under the shared 30-step budget. Image inputs and wall time include all within-step probe turns.
Symbol
Meaning
t
Outer browser-use decision step; each step ends with one browser operation.
dt,vt,ot
DOM text, screenshot, and raw multimodal observation ot=(dt,vt) at step t .
s,q,Ht−1
System prompt, user instruction, and inter-step history available before step t .
πθ
Agent policy parameterized by model parameters θ .
Aprobe
State-preserving probing action space used to augment the agent’s observation.
Abrowser
State-changing browser-operation action space, e.g., click, type, scroll, or navigate.
Table 7: Summary of notation used in the Probe-to-Act formulation.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Data Type / Source
# Gen. Traj.
# Gen. Turns (Probe %)
# Retained Turns (Probe %)
Rejection Rate (%)
Cold Start
Operation w/ Probing Primitives
3375
41587(36.6 %)
24052 (42.9 %)
42.2 %
Pure Operation
1943
21995
11771
46.5 %
GUI Grounding
−
−
40000
−
Round 1 Self-Bootstrapped
Operation w/ Probing Primitives
2260
42985 (37.3 %)
17931 (39.4 %)
58.3 %
Round 2 Self-Bootstrapped
Operation w/ Probing Primitives
1497
27245 (36.8 %)
13606 (41.3 %)
50.0 %
Appendix
Table 8: Statistics of the data generation and multi-round rejection sampling pipeline. Gen. indicates the generated data before filtering, while Retained denotes the high-quality data not rejected, kept for SFT. Probe % represents the proportion of internal probing primitive turns among all action turns.
Figure 6: We compare our framework’s performance and operation steps across three sites of VWA, with the baseline Qwen3-VL-8B as control.
Action Format
Category
Description
Standard Browser Operations
click(id)
Page Operation
Clicks on the specified DOM element.
type(id, text, enter_after)
Page Operation
Types text into a field, optionally pressing ‘Enter’.
hover(id)
Page Operation
Hovers over the specified DOM element.
press(key_comb, text)
Page Operation
Presses a key combination with optional text (e.g., ctrl+f ).
scroll(direction)
Page Operation
Scrolls the webpage up or down .
Appendix
Table 9: The complete action space for the browser-use agent. Action formats explicitly list all required parameters.
Model
SR
Browser ops
Probe turns
Image inputs
Wall time
Qwen3-VL-8B-Instruct
24.6
15.84
0
16.04
203.2 s
+ P2A (Round 2)
32.9
11.25
5.12
16.90
243.8 s
+ P2A (Prompt)
18.0
19.24
2.59
20.83
338.6 s
Gemini-3-Flash
42.4
12.80
0
13.20
380.1 s
+ P2A (Prompt)
48.7
13.08
8.14
20.70
432.8 s
Appendix
Table 10: Average per-task inference accounting on VWA. All image inputs and wall-clock measurements include the additional probe turns.
Figure 9: Comparison of Gemini-3-Pro with and without P2A. The task requires finding an item with specific visual features on a shopping site. Cluttered by global Set-of-Marks (SoM) overlays, the baseline struggles to perceive small product thumbnails, leading to a failed trial-and-error strategy of repeatedly opening and closing detail pages. In contrast, P2A identifies visual cues on uncluttered screenshots and utilizes the focus primitive to zoom in and verify the target, completing the task efficiently without any backtracking.
Figure 10: Gemini-3-Pro execution with and without P2A. The task involves finding an item and filling a form using a reference image. Relying solely on action history, the baseline loses visual context and falls into a repetitive purchasing loop (States 5–19). In contrast, P2A uses an evidence-based memory to track progress. By deploying the commit primitive (State 2, Call 2) to proactively save reference details, P2A decouples information extraction from form execution, effectively mitigating errors caused by simultaneous multi-image understanding.
Figure 11: Comparison of the baseline Qwen3-VL-8B and our P2A-tuned model. Without any web exploration, the baseline prematurely halts, falsely concluding the task is unachievable. In contrast, our agent employs the commit primitive (State 1, Call 1) to explicitly record its initial unsuccessful attempt within the task-specified section. Guided by this retained memory, the agent shifts its exploration strategy by sorting products from highest to lowest price (State 17). This efficient search successfully locates the target page, where the agent then utilizes the focus primitive (State 20, Call 0) to zoom in and accurately extract the required fine-grained visual features.
Figure 12: Comparison of the baseline Qwen3-VL-8B and our P2A-tuned model. The baseline prematurely halts without exploring the webpage, falsely concluding that the task is unachievable. Conversely, our agent uses the commit primitive (State 3, Call 0) to explicitly document its failed initial attempt of directly searching for the target forum (“f/aww”). Acknowledging this failure through its memory, the agent strategically adapts by sorting the forums alphabetically (State 12). This efficient workaround successfully locates the correct page, enabling the agent to extract the requested visual information.
Figure 13: Comparison of the baseline Qwen3-VL-8B and our P2A-tuned model. The task requires adding a product with specific visual features to the cart. Although the baseline locates the correct item, it fails to perceive critical visual feedback (a pop-up prompting to select all specifications first) and falls into a futile loop of repeatedly clicking the “Add to Cart” button (States 1–29). In contrast, our agent accurately interprets the dynamic feedback pop-up, selects the required product variants, and successfully completes the addition.
We present SUPERBROWSER, an autonomous web-navigation agent designed against a single guiding hypothesis: a web agent should browse the way a person browses. A human reading a page does not retain every pixel they have seen; they look at a few candidate targets, decide on one, and remember only what is needed to keep the goal alive. We operationalize this perception-cognition-action triad as three coupled mechanisms. First, a vision-first bounding-box pipeline labels candidate interactive regions on every screenshot and feeds them, asynchronously prefetched, to the language model so that the "eye" precedes the "hand". Second, a three-role brain -- an Orchestrator that classifies and routes, a Planner that evaluates progress every few steps, and a Worker that emits per-step actions -- separates strategic from operational reasoning. Third, a structured Ledger stores only what a person would: the goal, the last three actions, a small set of facts and dead-ends, and a handful of checkpoints; a six-phase eviction loop systematically discards stale screenshots, state blobs, and reasoning traces from the live context. Action execution is a three-tier click cascade (Chrome DevTools Protocol to Puppeteer to scripted) with humanized Bezier motion, plus a chevron-aware bounding-box snapper that resolves the "small arrow beside a large label" ambiguity. On the Mind2Web Hard benchmark (66 tasks), SUPERBROWSER attains 89.47% success, placing third overall and ahead of every published open/research browser-agent baseline by a large margin. We argue that the gain comes not from any single trick but from the consistent application of a cognitive contract throughout the system.
LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and solve them autonomously. However, evaluating proactivity is challenging; current benchmarks are constrained to localized context, limiting their ability to test reasoning across sources and longer time horizons. To address this gap, we present PROBE (Proactive Resolution Of BottlEnecks). PROBE decomposes proactivity as a pipeline of three core capabilities: (1) searching for unspecified issues, (2) identifying specific bottlenecks, and (3) executing appropriate resolutions. We apply PROBE to evaluate leading LLMs and popular agentic frameworks, showing that even state-of-the-art models struggle to solve this benchmark. Computing our consistent measurements across frontier LLMs and agents, we find that the best end-to-end performance of 40% is achieved by both GPT-5 and Claude Opus-4.1. Additionally, we demonstrate the relative capabilities of each model and analyze mutual failure modes. Our results highlight the current limitations of autonomous action in agentic systems, and expose promising future research directions.
The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in these settings is proactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or preferences unstated. However, existing benchmarks rarely evaluate whether agents can identify and act on such hidden intents before they are explicitly stated, especially in sustained multi-turn interactions where user needs emerge gradually. To address this gap, we introduce π-Bench, a benchmark for proactive assistance comprising 100 multi-turn tasks across 5 domain-specific user personas. By incorporating hidden user intents, inter-task dependencies, and cross-session continuity, π-Bench evaluates agents' ability to anticipate and address user needs over extended interactions, jointly measuring proactivity and task completion in long-horizon trajectories that better reflect real-world use. Experiments show (1) proactive assistance remains challenging, (2) a clear distinction between task completion and proactivity, and (3) the value of prior interaction for proactive intent resolution in later tasks.
Haoran Zhang, Luxin Xu, Zhilin Wang +11
Shanghai Jiao Tong University · Shanghai AI Laboratory · Fudan University +7