Many web agents operate through a reactive execution loop: they observe the current interface, reason about the next action, execute it, and repeat. This design incurs latency and cost that grow with the number of actions, while requiring agents to repeatedly rediscover how the same web application works. We present ActionEngine, a novel architecture that replaces step-by-step reasoning with programmatic execution using reusable knowledge of the application. A Crawling Agent explores the application offline and constructs an updatable state-machine memory that represents its GUI states, the operations available in each state, and the transitions between states. Unlike trajectory memory, this representation stores how the application works rather than solutions to individual tasks. At runtime, an Execution Agent uses this memory to synthesize a complete executable program in a single planning step, which is then executed deterministically without further planning calls. When the interface changes or the memory is incomplete, a reactive fallback repairs the failed action and updates the memory for future tasks. On 655 tasks across four WebArena domains, ActionEngine achieves a 91.2% success rate, outperforming the strongest reactive baseline, Claude Code, by 8.5 percentage points while reducing average task latency by 3.2x and cost by 8x.
Figures & tables
Figure 1: Overview of ActionEngine . The Crawling Agent (top, offline) systematically explores the website to construct the state-machine graph (SMG). The Execution Agent (bottom, online) uses the SMG to plan and execute tasks: the Planner synthesizes an executable program from the SMG and user task; the Executor runs it deterministically without further planning calls; and the Patcher intercepts failures, invokes a reactive fallback to recover, and updates the SMG.
Figure 2: State-machine representation of a web application. Pages with the same interface structure map to a shared state template, which consists of reusable UI atoms and exposes operations over them. Operations store how to access live application data rather than the data observed during construction; for example, ReadPosts specifies how to retrieve posts and their output schema.
Success Rate (%) ↑
Avg. Latency (s) ↓
Domain
# Tasks
AOccam
Claude
ActionEngine
AOccam
Claude
ActionEngine
Reddit
106
76.2
92.4
95.0
536
59
18
Shopping
187
42.7
79.7
87.0
363
97
28
Shop. Admin
182
58.4
74.2
95.3
1665
101
31
GitLab
180
58.3
89.0
89.0
499
79
28
Overall
655
56.8
82.7
91.2
787
87
27
Table 1: Performance on WebArena by domain. All systems use Claude Opus 4.6. Overall, ActionEngine improves success rate by 8.5 percentage points over the strongest reactive baseline while reducing average latency by 3.2× .
Figure 4: Average input tokens, output tokens, LLM calls, and estimated API cost per task. ActionEngine reduces input tokens by 14 – 22× and LLM calls by 4 – 6× compared to Claude Code.
Table 2: (a) Cost and quality of the SMG produced by offline crawling. Parenthesized operation counts report operations gained or lost after refinement; refinement does not change the number of states. Break-even is the number of tasks required to amortize the one-time crawling cost relative to Claude Code. (b) Effect of prefix caching on average API cost per task.
Computer-using agents drive real software through the screen -- clicking and typing -- but they solve every task from scratch: asked to repeat a task, an agent re-reads the screen, re-reasons every tap, and pays the full cost again. We present PreAct, which lets such an agent get faster on tasks it has done before. The first time it succeeds, PreAct compiles the run into a small state-machine program-states that check the screen, transitions that act-and on later runs replays it directly instead of invoking the agent 8.5-13x faster, with no per-step language-model calls. Replay is not blind: at each step PreAct checks that the screen matches what the program expects before acting, and hands control back to the agent the moment something is off. PreAct applies the same discipline when deciding what to keep: a freshly compiled program enters the store only if, re-run from a clean state, an independent evaluator confirms it solved the task-catching programs that replay to their last step yet leave the task undone. Across a mobile, a desktop, and a web benchmark, this store-time check separates repeated runs that improve from ones that degrade as faulty programs accumulate, worth 1.75-2.6 tasks per benchmark, the same direction on all three; a fallback that explores afresh when no program fits brings PreAct level with a strong record-and-replay baseline. We also report what did not matter: prompt wording, runtime guardrails, and whether a language model or a plain embedding retriever selects which program to reuse.
Skim is a speculative execution framework for web agents that exploits the predictable structure of purpose-built websites. Today's web-agent expense is not intrinsic to the tasks but a property of how agents are composed: frontier-model inference, browser rendering, and ReAct-style planning are applied to every step of every task regardless of complexity. Skim's key observation is that websites enforce stable URL patterns, answer formats, and task-to-trajectory mappings across queries of the same type, so most queries can bypass these heavyweight components entirely. An offline profiler captures these patterns once per site. At runtime, Skim matches each query to a template, synthesizes the destination URL, and extracts the answer with a small model. A lightweight verifier gates each fast-path output against the query and schema; rare misspeculations cascade to the full agent, warm-started by the fast path's final URL to preserve upstream trajectory progress. Across standard web-agent benchmarks paired with three backboneagents (WebVoyager, AgentOccam, BrowserUse), Skim reduces median per-task cost by 1.9x and latency by 33.4% with no accuracy loss.
Mike Wong, Kevin Hsieh, Suman Nath +1
Princeton University · Princeton, USA · Microsoft Research +1
ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default for web agents. Instead, web agents should default to plan-then-execute: commit to a task-specific program before observing runtime web content, then execute it. The reason is that web content mixes inputs from many parties. An e-commerce product page may combine a seller's listing, customer reviews and sponsored advertisements. Under ReAct, all of this content flows into the model when deciding on the next action, creating a direct path for prompt injections to steer the agent's control flow. Plan-then-execute changes this boundary: untrusted data may influence values or branches inside a predefined execution graph, but it cannot redefine the user task or cause the model to synthesize new actions at runtime. We analyze WebArena, a popular web agent benchmark, and find that all tasks are compatible with plan-then-execute, while 80% can be completed with a purely programmatic plan, without any runtime LLM subroutine. We identify the main barrier to adopting plan-then-execute on the web: For it to work well, tools must map cleanly to semantic actions, with effects known before execution, so agents have enough information to plan. The web does not naturally expose that interface. Browser tools such as click, type, and scroll have page-dependent meanings. Planning at this layer is near-sighted: the agent can only see actions on the current page, and later actions appear only after it acts. Closing this gap requires typed interfaces that turn website interactions from clicks and keystrokes to task-level operations. This is an infrastructure problem, not a modeling problem. Web tasks do not need reactivity by default; they need typed, complete, auditable website APIs.