GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Organizations: The Chinese University of Hong Kong, Shenzhen · Tianjin University · Harbin Institute of Technology (Shenzhen) · East China Normal University · Shenzhen Research Institute of Big Data
Abstract
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.
Figures & tables
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Initial harness | GUI-HARVEST | |||||||
| Backbone | Search | Val. | Test | Full | Search | Val. | Test | Full |
| Qwen3-VL-8B-Instruct | 31.01 | 29.72 | 28.79 | 29.49 | 41.43 (+10.41) | 37.22 (+7.50) | 34.84 (+6.05) | 36.83 (+7.34) |
| Qwen3-VL-32B-Instruct | 31.07 | 35.06 | 43.02 | 38.61 | 46.25 (+15.18) | 50.00 (+14.94) | 53.18 (+10.16) | 50.94 (+12.33) |
| OpenCUA-32B | 27.40 | 28.75 | 32.82 | 30.71 | 33.69 (+6.29) | 34.63 (+5.88) | 34.31 (+1.49) | 34.24 (+3.53) |
| OpenCUA-72B | 36.25 | 30.00 | 41.25 | 37.65 | 42.50 (+6.25) | 37.50 (+7.50) | 44.18 (+2.93) | 42.33 (+4.68) |
| Gemini 3.1 Pro | 65.33 | 67.69 | 66.42 | 66.46 | 72.68 (+7.35) | 75.92 (+8.23) | 72.63 (+6.20) | 73.37 (+6.91) |
| Method / model | Step | OS | Office | Daily | Prof. | Multi | Avg. |
|---|---|---|---|---|---|---|---|
| Max 15 Steps | |||||||
| GUI-HARVEST / Qwen3-VL-8B-Instruct | 15 | 58.33 | 32.47 | 38.64 | 67.35 | 19.16 | 36.83 |
| GUI-HARVEST / Qwen3-VL-32B-Instruct | 15 | 83.33 | 48.68 | 57.62 | 75.51 | 26.88 | 50.94 |
| Max 50 Steps | |||||||
| Qwen / Qwen3-VL-8B-Instruct ( Yang et al., 2026a ) | 50 | – | – | – | – | – | 33.90 |
| OS-Symphony / Qwen3-VL-8B-Instruct ( Yang et al., 2026a ) | 50 | – | – | – | – | – | 33.90 |
| Method / model | Step | OS | Office | Daily | Prof. | Multi | Avg. |
|---|---|---|---|---|---|---|---|
| Max 15 Steps | |||||||
| UI-TARS-1.5-7B ( Song et al., 2026 ) | 15 | 34.78 | 27.19 | 27.99 | 61.45 | 5.38 | 25.76 |
| OpenCUA / OpenCUA-32B ( Wang et al., 2025 ) | 15 | – | – | – | – | – | 29.71 |
| OpenCUA / OpenCUA-72B ( Wang et al., 2025 ) | 15 | – | – | – | – | – | 39.03 |
| GUI-HARVEST / OpenCUA-32B | 15 | 58.33 | 29.05 | 44.76 | 65.31 | 9.36 | 34.24 |
| GUI-HARVEST / OpenCUA-72B | 15 | 45.83 | 41.01 | 48.67 | 75.51 | 20.28 | 42.33 |
| Method / model | Step | OS | Office | Daily | Prof. | Multi | Avg. |
|---|---|---|---|---|---|---|---|
| Max 15 Steps | |||||||
| OpenAI o3 ( Song et al., 2026 ) | 15 | 37.50 | 1.45 | 8.02 | 12.29 | 11.82 | 9.09 |
| OpenAI CUA / GPT-4o ( Song et al., 2026 ) | 15 | 45.83 | 22.17 | 37.65 | 41.22 | 10.75 | 26.01 |
| Jedi-7B w/ GPT-4o ( OSWorld Team, 2026 ) | 15 | – | – | – | – | – | 26.80 |
| Agent S2.5 / OpenAI o3 ( Song et al., 2026 ) | 15 | 70.83 | 42.85 | 44.61 | 57.10 | 17.82 | 38.98 |
| CoAct-1 / GPT-5 ( Song et al., 2026 ) | 15 | 66.67 | 47.18 | 42.30 | 47.74 | 23.82 | 39.81 |
| Qwen3-VL- 32B-Instruct | Qwen3-VL- 8B-Instruct | Gemini 3.1 Pro | GPT-5 | OpenCUA-72B | OpenCUA-32B | |
| Flip rate (%) | 20.4 | 16.6 | 12.7 | 12.4 | 15.6 | 11.8 |
| Backbone | Verified pattern changes ( ) | Selected mechanisms | Residual bottleneck |
| General-purpose open models | |||
| Qwen3-VL-8B-Instruct | redundant action loop ; inefficient execution strategy ; infeasibility misjudgment | loop rejection and fresh restart; code routing; budget verdict; save nudge | after a blocked action, the model may choose another ineffective route or emit done() again without correcting the state |
| Qwen3-VL-32B-Instruct | inefficient execution strategy ; premature completion ; redundant action loop | application-aware code routing; route ledger; save gate; completion adjudication; code receipt | wrong targets or routes despite executed actions; incomplete work after code execution |
| GUI-specialized models | |||
| OpenCUA-32B | zero-duration drag ; unsupported mouse API ; infeasibility misjudgment | drag-duration normalization; computer.* -to- pyautogui.* mapping; infeasibility remapping | correctly executed actions can still select the wrong target or drag distance; no alternative execution path is introduced |
| OpenCUA-72B | unpersisted LibreOffice final state ; dropped terminal action ; infeasibility misjudgment | pre-completion save; preservation of the last-step DONE ; text-grounded infeasibility handling | ineffective or wrong GUI actions and visually plausible false completion |
| Behavioral pattern | Affected backbones | Shared intervention principle | Model-specific realizations |
| Execution and control | |||
| Inefficient execution strategy | Qwen3-VL-8B and 32B; Gemini 3.1 Pro; GPT-5 | route bulk or repetitive operations through a shorter compatible execution path | document maps and application references for Qwen; code routing and Office finalization for Gemini and GPT-5 |
| Redundant action loop | Qwen3-VL-8B and 32B; GPT-5 | revise or withhold an action when repeated execution produces no useful state change | loop rejection and fresh restart; route retirement; coordinate re-query and alternative code routing |
| Completion judgment | |||
| Premature completion | Qwen3-VL-8B and 32B; OpenCUA-72B; Gemini 3.1 Pro | require the task-relevant edit to be committed or persisted before accepting completion | save nudge; save gate and completion adjudication; pre-completion save; Office-file normalization |
| Infeasibility misjudgment | all six backbones | use sufficiently supported infeasibility evidence to select the correct terminal outcome | budget verdict and capability checks for general models; constrained feasibility probes for frontier models; text-grounded terminal remapping for OpenCUA |
| Backbone | Harness | WAA |
| Qwen3-VL-32B-Instruct | Qwen | 31.68 |
| Agent S3 | 38.21 | |
| GUI-HARVEST | 44.68 | |
| GPT-5 | Agent S3 | 50.88 |
| GUI-HARVEST | 64.75 |
| Method / backbone | Steps | Office | Web | System | Code | Media | Utility | Avg. |
| Max 50 Steps | ||||||||
| Qwen / Qwen3-VL-32B-Instruct ( Yang et al., 2026a ) | 50 | 19.05 | 49.66 | 54.17 | 21.05 | 42.19 | 25.00 | 31.68 |
| OS-Symphony / Qwen3-VL-32B-Instruct ( Yang et al., 2026a ) | 50 | 26.19 | 46.33 | 75.00 | 47.37 | 27.90 | 41.67 | 45.32 |
| VLAA-GUI / Gemini 3 Flash ( Han et al., 2026 ) | 50 | 32.60 | 73.30 | 87.50 | 66.70 | 52.40 | 75.00 | 60.40 |
| OS-Symphony / GPT-5-Mini ( Yang et al., 2026a ) | 50 | 42.86 | 73.00 | 79.17 | 68.42 | 48.66 | 66.67 | 62.15 |
| OS-Symphony / GPT-5 ( Yang et al., 2026a ) | 50 | 54.76 | 73.00 | 75.00 | 42.11 | 70.09 | 75.00 | 63.45 |