Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent's continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217× and latency by 3.4-5.1×. On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.
Figures & tables
Success rate (%) ↑
Per run ↓
Model
Method
Pass^1
Pass^3
Cost
Time
ScienceBoard (169 tasks)
Base agent
27.0
13.6
15 ×
4.6 ×
Batch2Code
15.4
11.8
1.1 ×
1.0 ×
ASI
24.3
11.8
16 ×
5.1 ×
GPT-5.6 Terra
NSPI (ours)
29.2
26.0
0.00536
40
Table 1: Reliability across repeated executions on ScienceBoard and OSWorld. Per-run cost (USD) and wall-clock latency for NSPI and as relative multiples ( × ) of NSPI for other methods. NSPI achieves the highest Pass^3 in all four settings while operating at substantially lower per-run cost. Batch2Code, a conversion-only variant of our framework is also efficient, but with substantially lower Pass^k.
Table 2: Effectiveness of policy built across parameter and initial state variations to held-out instances of OSWorld workflows.
Cost (USD)
Wall-clock time
Benchmark
Base model
Cπ
cμ
cπ
n∗
Cπ (min)
cμ (s)
cπ (s)
n∗
ScienceBoard
GPT-5.6 Terra
2.03
0.081
0.0054
26.8
40
183
40
16.9
Claude Opus 5
5.68
0.52
0.0079
11.2
49
251
64
15.8
OSWorld v1
GPT-5.6 Terra
3.09
0.48
0.0072
6.5
33
286
56
8.6
Claude Opus 5
6.45
2.1
0.0097
3.1
43
353
102
10.3
Table 3: Cost and wall-clock latency amortization of learning a reusable NSPI in the terms of Equation 3 . Cπ is the one-time policy-construction cost, cμ the per-run cost of the base agent, cπ the per-run cost of executing the learned policy, and n∗ is the break-even count. Values are means over tasks.
Benchmark
Models
Prec.
Rec.
Spec.
ScienceBoard
Terra
41.7
11.4
94.4
Opus 5
69.6
22.2
92.7
OSWorld
Terra
76.8
40.1
87.6
Opus 5
76.7
35.3
69.3
Table 5: Agreement between NSPI judges and the benchmark evaluator on final policies.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Primitive
Mut.
Behavior
Neural: grounded GUI action G(ℓ,ot)
click( ℓ , clicks, button, hold_keys)
✓
Grounds the element description ℓ to a screen coordinate and clicks it; clicks sets double or triple clicks and hold_keys holds modifiers.
type( ℓ , text, enter, overwrite)
✓
Clicks the element grounded from ℓ if given, clears it if overwrite , types text , and presses Enter if enter .
scroll( ℓ , amount)
✓
Moves the pointer to the element grounded from ℓ and scrolls by amount .
drag_and_drop( ℓstart , ℓend , hold_keys)
✓
Grounds both descriptions and drags from the first element to the second.
Neural: condition check Q(ϕ,ot)
Appendix
Table 7: Primitive action space of policies on ScienceBoard. Each primitive is a method of an agent object that the policy calls, e.g., agent.click("File menu") . Neural primitives call a model on the current observation ot : the grounded actions use UI-Venus-1.5-30B-A3B ( Team et al., 2026 ) with a consensus over three parallel calls, and the condition check uses Gemini 3 Flash. A check mark under Mut. marks a primitive that changes the state of the computer; these count toward the per-task step budget and are gated by the pre-action verifier.
Setting
ScienceBoard
OSWorld
Base agent: output tokens, step limit
8192; task step limit
4096; 100 steps
Code rollout step limit
task limit × 5 (Terra), × 16 (Opus)
100 primitive calls
Converter / healer / continuation tokens, Terra
2000 / 4096 / 1024
4096 / 8192 / 1024
Converter / healer / continuation tokens, Opus 5
4096 / 16384 / 2048
4096 / 16384 / 2048
Continuation turns
15
15
Observer / localizer ( Jloc )
Gemini 2.5 Flash / Gemini 2.5 Pro
Appendix
Table 8: Main-experiment configuration. All judge roles use temperature 0.
ID
App
Step
Visible
t∗
Rating
Cause
R01
Chrome
27
27
27
C / C
Site is sold out; plan should report the task as infeasible
R02
Chrome
3
22
22
W / C
date -d monday returns today, not next Monday
R03
GIMP
9
9
13
A / A
Click on Mode does not open the menu; next click reuses its point
R04
GIMP
14
18
15
A / W
Condition check returns empty; Export skipped; Replace dialog left open
R05
GIMP
6
6
9
A / A
Wait too short; click misses; plan falls one step behind
R06
GIMP
2
5
5
W / C
Infeasible: plan writes a “Blue” theme that does not exist
Appendix
Table 9: The 19 annotated rollouts. Step: true divergence step. Visible: first visible divergence (–, never). t∗ rating against the true step and the first visible divergence: C correct, A adjacent, W wrong.
Setting
Value
Iterations N / rollouts per iteration K / final trials
5 / 3 / 3
Converter Φconv and healer Φheal
Claude Opus 5
Continuation agent μ
harness + DeepSeek-V4.1-Flash, 500 steps, 3 h cap
Task-completion judge Jtask
Claude Sonnet 4.6, 60 tool calls, 900 s
Observer / localizer
Gemini 2.5 Flash / Gemini 2.5 Pro
Screenshots per heal
5, evenly spaced
Appendix
Table 10: Settings of the OSWorld-V2 construction runs.