Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent's continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217× and latency by 3.4-5.1×. On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.
Figures & tables
Success rate (%) ↑
Per run ↓
Model
Method
Pass^1
Pass^3
Cost
Time
ScienceBoard (169 tasks)
Base agent
27.0
13.6
15 ×
4.6 ×
Batch2Code
15.4
11.8
1.1 ×
1.0 ×
ASI
24.3
11.8
16 ×
5.1 ×
GPT-5.6 Terra
NSPI (ours)
29.2
26.0
0.00536
40
Table 1: Reliability across repeated executions on ScienceBoard and OSWorld. Per-run cost (USD) and wall-clock latency for NSPI and as relative multiples ( × ) of NSPI for other methods. NSPI achieves the highest Pass^3 in all four settings while operating at substantially lower per-run cost. Batch2Code, a conversion-only variant of our framework is also efficient, but with substantially lower Pass^k.
Table 2: Effectiveness of policy built across parameter and initial state variations to held-out instances of OSWorld workflows.
Cost (USD)
Wall-clock time
Benchmark
Base model
Cπ
cμ
cπ
n∗
Cπ (min)
cμ (s)
cπ (s)
n∗
ScienceBoard
GPT-5.6 Terra
2.03
0.081
0.0054
26.8
40
183
40
16.9
Claude Opus 5
5.68
0.52
0.0079
11.2
49
251
64
15.8
OSWorld v1
GPT-5.6 Terra
3.09
0.48
0.0072
6.5
33
286
56
8.6
Claude Opus 5
6.45
2.1
0.0097
3.1
43
353
102
10.3
Table 3: Cost and wall-clock latency amortization of learning a reusable NSPI in the terms of Equation 3 . Cπ is the one-time policy-construction cost, cμ the per-run cost of the base agent, cπ the per-run cost of executing the learned policy, and n∗ is the break-even count. Values are means over tasks.
Benchmark
Models
Prec.
Rec.
Spec.
ScienceBoard
Terra
41.7
11.4
94.4
Opus 5
69.6
22.2
92.7
OSWorld
Terra
76.8
40.1
87.6
Opus 5
76.7
35.3
69.3
Table 5: Agreement between NSPI judges and the benchmark evaluator on final policies.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Primitive
Mut.
Behavior
Neural: grounded GUI action G(ℓ,ot)
click( ℓ , clicks, button, hold_keys)
✓
Grounds the element description ℓ to a screen coordinate and clicks it; clicks sets double or triple clicks and hold_keys holds modifiers.
type( ℓ , text, enter, overwrite)
✓
Clicks the element grounded from ℓ if given, clears it if overwrite , types text , and presses Enter if enter .
scroll( ℓ , amount)
✓
Moves the pointer to the element grounded from ℓ and scrolls by amount .
drag_and_drop( ℓstart , ℓend , hold_keys)
✓
Grounds both descriptions and drags from the first element to the second.
Neural: condition check Q(ϕ,ot)
Appendix
Table 7: Primitive action space of policies on ScienceBoard. Each primitive is a method of an agent object that the policy calls, e.g., agent.click("File menu") . Neural primitives call a model on the current observation ot : the grounded actions use UI-Venus-1.5-30B-A3B ( Team et al., 2026 ) with a consensus over three parallel calls, and the condition check uses Gemini 3 Flash. A check mark under Mut. marks a primitive that changes the state of the computer; these count toward the per-task step budget and are gated by the pre-action verifier.
Setting
ScienceBoard
OSWorld
Base agent: output tokens, step limit
8192; task step limit
4096; 100 steps
Code rollout step limit
task limit × 5 (Terra), × 16 (Opus)
100 primitive calls
Converter / healer / continuation tokens, Terra
2000 / 4096 / 1024
4096 / 8192 / 1024
Converter / healer / continuation tokens, Opus 5
4096 / 16384 / 2048
4096 / 16384 / 2048
Continuation turns
15
15
Observer / localizer ( Jloc )
Gemini 2.5 Flash / Gemini 2.5 Pro
Appendix
Table 8: Main-experiment configuration. All judge roles use temperature 0.
ID
App
Step
Visible
t∗
Rating
Cause
R01
Chrome
27
27
27
C / C
Site is sold out; plan should report the task as infeasible
R02
Chrome
3
22
22
W / C
date -d monday returns today, not next Monday
R03
GIMP
9
9
13
A / A
Click on Mode does not open the menu; next click reuses its point
R04
GIMP
14
18
15
A / W
Condition check returns empty; Export skipped; Replace dialog left open
R05
GIMP
6
6
9
A / A
Wait too short; click misses; plan falls one step behind
R06
GIMP
2
5
5
W / C
Infeasible: plan writes a “Blue” theme that does not exist
Appendix
Table 9: The 19 annotated rollouts. Step: true divergence step. Visible: first visible divergence (–, never). t∗ rating against the true step and the first visible divergence: C correct, A adjacent, W wrong.
Setting
Value
Iterations N / rollouts per iteration K / final trials
5 / 3 / 3
Converter Φconv and healer Φheal
Claude Opus 5
Continuation agent μ
harness + DeepSeek-V4.1-Flash, 500 steps, 3 h cap
Task-completion judge Jtask
Claude Sonnet 4.6, 60 tool calls, 900 s
Observer / localizer
Gemini 2.5 Flash / Gemini 2.5 Pro
Screenshots per heal
5, evenly spaced
Appendix
Table 10: Settings of the OSWorld-V2 construction runs.
Computer-using agents drive real software through the screen -- clicking and typing -- but they solve every task from scratch: asked to repeat a task, an agent re-reads the screen, re-reasons every tap, and pays the full cost again. We present PreAct, which lets such an agent get faster on tasks it has done before. The first time it succeeds, PreAct compiles the run into a small state-machine program-states that check the screen, transitions that act-and on later runs replays it directly instead of invoking the agent 8.5-13x faster, with no per-step language-model calls. Replay is not blind: at each step PreAct checks that the screen matches what the program expects before acting, and hands control back to the agent the moment something is off. PreAct applies the same discipline when deciding what to keep: a freshly compiled program enters the store only if, re-run from a clean state, an independent evaluator confirms it solved the task-catching programs that replay to their last step yet leave the task undone. Across a mobile, a desktop, and a web benchmark, this store-time check separates repeated runs that improve from ones that degrade as faulty programs accumulate, worth 1.75-2.6 tasks per benchmark, the same direction on all three; a fallback that explores afresh when no program fits brings PreAct level with a strong record-and-replay baseline. We also report what did not matter: prompt wording, runtime guardrails, and whether a language model or a plain embedding retriever selects which program to reuse.
Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning is test-time scaling, which trains models to search over long chains of thought; but the resulting capability is entangled in model weights, is not verifiable step-by-step, and is costly at inference. We present Forethought, a neurosymbolic reasoning system that instead treats reasoning as an explicit, verifiable program, that builds from a library of symbolic and neural primitives which are composed through a domain-specific language. The result are reasoning programs, which are concrete representations of the model's work, and as such can be inspected and modified before deployment. Instantiated as a tool-calling execution kernel and evaluated across five benchmarks, Forethought improves base-model accuracy by about 30% relative and outperforms vanilla prompting, reinforcement learning scaffolds, and prompt-evolution methods, enabling small models to match or exceed frontier models capabilities. In a direct comparison, a non-reasoning model augmented with Forethought competes with a dedicated reasoning model while requiring roughly three orders of magnitude less post-training investment, and remains model-agnostic and auditable.
Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez
Computer-use agents have rapidly improved on real-world tasks such as web navigation, desktop automation, and software interaction, in some cases surpassing human performance. Yet even when the task and model are unchanged, an agent that succeeds once may fail on a repeated execution of the same task. This raises a fundamental question: if an agent can succeed at a task once, what prevents it from doing so reliably? In this work, we study the sources of unreliability in computer-use agents through three factors: stochasticity during execution, ambiguity in task specification, and variability in agent behavior. We analyze these factors on OSWorld using repeated executions of the same task together with paired statistical tests that capture task-level changes across settings. Our analysis shows that reliability depends on both how tasks are specified and how agent behavior varies across executions. These findings suggest the need to evaluate agents under repeated execution, to allow agents to resolve task ambiguity through interaction, and to favor strategies that remain stable across runs.
Gonzalo Gonzalez-Pumariega, Saaket Agashe, Jiachen Yang +2