GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Authors: Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai
Organizations: The Chinese University of Hong Kong, Shenzhen · Tianjin University · Harbin Institute of Technology (Shenzhen) · East China Normal University · Shenzhen Research Institute of Big Data
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.
Figures & tables
Figure 1: OSWorld-Verified results for six backbones at 15, 50, and 100 maximum environment steps. Axes fix the backbone; markers compare harnesses and agent frameworks at each budget. GUI-HARVEST uses only 15-step Search rollouts for optimization and remains frozen for evaluation, achieving the highest score in every model–budget comparison shown.
Figure 2: GUI-HARVEST self-improvement loop. A frozen backbone and the current harness produce repeated search executions ( K=3 ), which support task-local diagnosis, cross-task clustering, and source-code edits with explicit behavioral predictions. The Validator combines score and behavior checks to promote or roll back each edit; the test set remains sealed until optimization stops.
Table 3
Figure 3: Harness evolution for Qwen3-VL-32B, OpenCUA-72B, and Gemini 3.1 Pro. Solid and dashed curves show Search and Validation scores, respectively; filled circles mark promoted harnesses, and the rightmost markers show Test (T) and Full (F) scores after optimization stops. The numbered tiles report candidate-edit attempts in each optimization round, and the red cross marks the terminal round. Results for the other three backbones are plotted in Appendix Figure 8 .
Figure 4: OSWorld accuracy versus full-suite API cost for systems with available cost records. We show the 15-step GUI-HARVEST points for Qwen3-VL-32B-Instruct and GPT-5 and the 100-step point for Gemini 3.1 Pro. Published points retain the models, harnesses, budgets, protocols, and per-task costs reported by Wei et al. (2026) ; Gonzalez-Pumariega et al. (2026b) ; full-suite costs are reconstructed as detailed in Appendix F .
Table 3: Component ablations and optimization traces on Qwen3-VL-32B-Instruct. Test and Full gains are measured against the initial harness. Each numbered trace cell denotes a promoted round; its value and shade encode the number of edit attempts required for promotion. X marks termination. All selected harnesses use the same K=3 final evaluation protocol.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Initial harness H0
GUI-HARVEST H∗
Backbone
Search
Val.
Test
Full
Search
Val.
Test
Full
Qwen3-VL-8B-Instruct
31.01
29.72
28.79
29.49
41.43 (+10.41)
37.22 (+7.50)
34.84 (+6.05)
36.83 (+7.34)
Qwen3-VL-32B-Instruct
31.07
35.06
43.02
38.61
46.25 (+15.18)
50.00 (+14.94)
53.18 (+10.16)
50.94 (+12.33)
OpenCUA-32B
27.40
28.75
32.82
30.71
33.69 (+6.29)
34.63 (+5.88)
34.31 (+1.49)
34.24 (+3.53)
OpenCUA-72B
36.25
30.00
41.25
37.65
42.50 (+6.25)
37.50 (+7.50)
44.18 (+2.93)
42.33 (+4.68)
Gemini 3.1 Pro
65.33
67.69
66.42
66.46
72.68 (+7.35)
75.92 (+8.23)
72.63 (+6.20)
73.37 (+6.91)
Appendix
Table 4: Complete paired 15-step results before and after harness optimization. Search and Validation guide optimization, whereas Test remains sealed until the final harness is selected. Full reports all 361 OSWorld-Verified tasks. Parentheses show absolute percentage-point gains over the corresponding initial harness.
Method / model
Step
OS
Office
Daily
Prof.
Multi
Avg.
Max 15 Steps
GUI-HARVEST / Qwen3-VL-8B-Instruct
15
58.33
32.47
38.64
67.35
19.16
36.83
GUI-HARVEST / Qwen3-VL-32B-Instruct
15
83.33
48.68
57.62
75.51
26.88
50.94
Max 50 Steps
Qwen / Qwen3-VL-8B-Instruct ( Yang et al., 2026a )
50
–
–
–
–
–
33.90
OS-Symphony / Qwen3-VL-8B-Instruct ( Yang et al., 2026a )
50
–
–
–
–
–
33.90
Appendix
Table 5: OSWorld-Verified comparison with general open-weight backbones. Public rows retain their original runtime and evaluation budget; GUI-HARVEST is optimized with 15-step Search rollouts and frozen before evaluation at all three budgets ( Xie et al., 2024 ; Yang et al., 2026a ; Bai et al., 2025 ; OSWorld Team, 2026 ) . “–” denotes an unreported domain score; bold marks the best reported value within each step-budget block.
Method / model
Step
OS
Office
Daily
Prof.
Multi
Avg.
Max 15 Steps
UI-TARS-1.5-7B ( Song et al., 2026 )
15
34.78
27.19
27.99
61.45
5.38
25.76
OpenCUA / OpenCUA-32B ( Wang et al., 2025 )
15
–
–
–
–
–
29.71
OpenCUA / OpenCUA-72B ( Wang et al., 2025 )
15
–
–
–
–
–
39.03
GUI-HARVEST / OpenCUA-32B
15
58.33
29.05
44.76
65.31
9.36
34.24
GUI-HARVEST / OpenCUA-72B
15
45.83
41.01
48.67
75.51
20.28
42.33
Appendix
Table 6: OSWorld-Verified comparison with GUI-specialized open models. OpenCUA rows use the official coordinate-action runtime; LFF and GUI-HARVEST modify that runtime family ( Qin et al., 2025 ; Wang et al., 2025 ; Sun et al., 2026 ; Yang et al., 2026a ; OSWorld Team, 2026 ) . “–” denotes an unreported domain score; bold marks the best reported value within each step-budget block.
Method / model
Step
OS
Office
Daily
Prof.
Multi
Avg.
Max 15 Steps
OpenAI o3 ( Song et al., 2026 )
15
37.50
1.45
8.02
12.29
11.82
9.09
OpenAI CUA / GPT-4o ( Song et al., 2026 )
15
45.83
22.17
37.65
41.22
10.75
26.01
Jedi-7B w/ GPT-4o ( OSWorld Team, 2026 )
15
–
–
–
–
–
26.80
Agent S2.5 / OpenAI o3 ( Song et al., 2026 )
15
70.83
42.85
44.61
57.10
17.82
38.98
CoAct-1 / GPT-5 ( Song et al., 2026 )
15
66.67
47.18
42.30
47.74
23.82
39.81
Appendix
Table 7: OSWorld-Verified comparison with proprietary frontier backbones and systems. Each public row preserves its original framework and step budget and provides literature-level benchmark context ( Agashe et al., 2025b ; Song et al., 2026 ; Yang et al., 2026c ; Gonzalez-Pumariega et al., 2026b ; Yang et al., 2026a ; Han et al., 2026 ; OSWorld Team, 2026 ) . “–” denotes an unreported domain score; bold marks the best reported value within each step-budget block.
Figure 5: Selected OSWorld-Verified operating points at 15, 50, and 100 maximum environment steps. Colored bars show frozen GUI-HARVEST harnesses optimized with 15-step Search rollouts; gray bars show published systems in their reported model and runtime configurations. This visualization summarizes the domain-level tables above; published values are drawn from the corresponding papers and the official OSWorld leaderboard ( Song et al., 2026 ; Wang et al., 2025 ; Sun et al., 2026 ; Yang et al., 2026a ; Han et al., 2026 ; OSWorld Team, 2026 ) .
Figure 6: Execution variability in desktop GUI tasks. (a) Clean resets one minute apart already differ in the initial frame. (b) Task-irrelevant system and web events enter the observation stream during execution. (c) Two runs share the first four steps but diverge at step five and receive opposite scores. These perturbations enter the multimodal feedback loop rather than remaining isolated pixel noise.
Qwen3-VL- 32B-Instruct
Qwen3-VL- 8B-Instruct
Gemini 3.1 Pro
GPT-5
OpenCUA-72B
OpenCUA-32B
Flip rate (%)
20.4
16.6
12.7
12.4
15.6
11.8
Appendix
Table 8: Task-level flip rates under each baseline harness over K=3 Full-361 executions. A task flips when its repeated runs contain both a zero score and a positive score.
Figure 7: Effect of the repeat count K for Qwen3-VL-32B-Instruct. (a) Standard error of the Full mean decreases with K , while rollout API cost scales approximately linearly. (b) Task-level flip rate as K varies, with always-positive and never-positive shares shown for context. (c) Full score of the harness selected by each setting; the dashed line is the initial harness.
correctly executed actions can still select the wrong target or drag distance; no alternative execution path is introduced
OpenCUA-72B
unpersisted LibreOffice final state 24→18 ; dropped terminal action 4→0 ; infeasibility misjudgment 5→3
pre-completion save; preservation of the last-step DONE ; text-grounded infeasibility handling
ineffective or wrong GUI actions and visually plausible false completion
Appendix
Table 9: Representative verified behavioral-pattern changes and the corresponding runtime adaptations for each backbone. Counts denote Search-set tasks matching each pattern under the initial and final harnesses ( H0→H∗ ). Patterns are not mutually exclusive, so the counts do not form a partition of failures.
Behavioral pattern
Affected backbones
Shared intervention principle
Model-specific realizations
Execution and control
Inefficient execution strategy
Qwen3-VL-8B and 32B; Gemini 3.1 Pro; GPT-5
route bulk or repetitive operations through a shorter compatible execution path
document maps and application references for Qwen; code routing and Office finalization for Gemini and GPT-5
Redundant action loop
Qwen3-VL-8B and 32B; GPT-5
revise or withhold an action when repeated execution produces no useful state change
loop rejection and fresh restart; route retirement; coordinate re-query and alternative code routing
Completion judgment
Premature completion
Qwen3-VL-8B and 32B; OpenCUA-72B; Gemini 3.1 Pro
require the task-relevant edit to be committed or persisted before accepting completion
save nudge; save gate and completion adjudication; pre-completion save; Office-file normalization
Infeasibility misjudgment
all six backbones
use sufficiently supported infeasibility evidence to select the correct terminal outcome
budget verdict and capability checks for general models; constrained feasibility probes for frontier models; text-grounded terminal remapping for OpenCUA
Appendix
Table 10: Cross-backbone organization of recurring, harness-addressable behavioral patterns. The affected-backbone column lists models in which each pattern was verified. A pattern groups model-specific manifestations with the same runtime consequence; it does not imply an identical underlying cause.
Figure 8: Harness evolution for Qwen3-VL-8B, GPT-5, and OpenCUA-32B, complementing Figure 3 . Solid and dashed curves show Search and Validation scores, respectively; filled circles mark promoted harnesses, and the rightmost markers show Test (T) and Full (F) scores after optimization stops. The numbered tiles report candidate-edit attempts in each optimization round, and the red cross marks the terminal round.
Backbone
Harness
WAA ↑
Qwen3-VL-32B-Instruct
Qwen
31.68
Agent S3
38.21
GUI-HARVEST
44.68
GPT-5
Agent S3
50.88
GUI-HARVEST
64.75
Appendix
Table 11: Same-backbone transfer to WindowsAgentArena at 50 steps. The GUI-HARVEST harnesses are selected on OSWorld and frozen before WAA evaluation; no WAA trajectory enters optimization. The published Qwen operating point is from Yang et al. (2026a) .
Method / backbone
Steps
Office
Web
System
Code
Media
Utility
Avg.
Max 50 Steps
Qwen / Qwen3-VL-32B-Instruct ( Yang et al., 2026a )
50
19.05
49.66
54.17
21.05
42.19
25.00
31.68
OS-Symphony / Qwen3-VL-32B-Instruct ( Yang et al., 2026a )
50
26.19
46.33
75.00
47.37
27.90
41.67
45.32
VLAA-GUI / Gemini 3 Flash ( Han et al., 2026 )
50
32.60
73.30
87.50
66.70
52.40
75.00
60.40
OS-Symphony / GPT-5-Mini ( Yang et al., 2026a )
50
42.86
73.00
79.17
68.42
48.66
66.67
62.15
OS-Symphony / GPT-5 ( Yang et al., 2026a )
50
54.76
73.00
75.00
42.11
70.09
75.00
63.45
Appendix
Table 12: WindowsAgentArena results. Published rows retain their reported model, runtime, and step budget; GUI-HARVEST rows transfer the frozen OSWorld-derived harness with platform adapters only. WAA contains six task domains ( Bonatti et al., 2025 ) . Published comparison values are from Yang et al. (2026a) and Han et al. (2026) . For OS-Symphony rows, Avg. is the published score over all 154 tasks, including its separately reported 13-task infeasible subset; the six displayed columns are the WAA application domains.
Figure 9: Full-suite API cost and OSWorld-Verified score for three target backbones. Open markers denote the 15-step Agent S3 baseline; filled markers denote the frozen GUI-HARVEST harness evaluated at 15, 50, and 100 steps. Costs cover target-model API calls for the 361-task evaluation.
Figure 10: GAFT evidence for a pointing action. The toolkit aligns the executed click with before/after screenshots and reports pixel change over the full screen and a local window around the landing point. These measurements are mechanical evidence supplied to the Evidence Analyst, not semantic judgments.
Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modular GUI agentic framework built around three integrated components that guide the system on when to Stop, Recover, and Search. First, a mandatory Completeness Verifier enforces UI-observable success criteria and verification at every finish step -- with an agent-level verifier that cross-examines completion claims with decision rules, rejecting those lacking direct visual evidence. Second, a mandatory Loop Breaker provides multi-tier filtering: switching interaction mode after repeated failures, forcing strategy changes after persistent screen-state recurrence, and binding reflection signals to strategy shifts. Third, an on-demand Search Agent searches online for unfamiliar workflows by directly querying a capable LLM with search ability, returning results as plain text. We additionally integrate a Coding Agent for code-intensive actions and a Grounding Agent for precise action grounding, both invoked on demand when required. We evaluate VLAA-GUI across five top-tier backbones, including Opus 4.5, 4.6 and Gemini 3.1 Pro, on two benchmarks with Linux and Windows tasks, achieving top performance on both (77.5% on OSWorld and 61.0% on WindowsAgentArena). Notably, three of the five backbones surpass human performance (72.4%) on OSWorld in a single pass. Ablation studies show that all three proposed components consistently improve a strong backbone, while a weaker backbone benefits more from these tools when the step budget is sufficient. Further analysis also shows that the Loop Breaker nearly halves wasted steps for loop-prone models.
Long-horizon GUI automation remains challenging due to error accumulation over extended interaction sequences. Process Reward Models (PRMs) provide dense step-level supervision for mitigating error accumulation, yet standard PRMs are poorly suited to GUI verification. Standard PRM judgments often rely on superficial visual alignment rather than functional correctness, reflecting an evaluative knowledge gap caused by missing domain-specific adjudication logic. Standard PRMs also perform passive, single-pass visual assessment, which creates Visual Ambiguity when reliable judgment requires actively locating, parsing, or inspecting task-relevant UI evidence. We introduce GUI-PRA, a Process Reward Agent that transforms GUI process evaluation from passive scoring into active investigation. GUI-PRA couples Experience-Injected Criterion Synthesis, which distills generalized verification principles into state-specific criteria, with Criterion-Guided Autoregressive Perception, which uses these criteria to navigate multi-granularity visual tools and gather grounded evidence. On AndroidWorld and Mobile-MiniWoB++, GUI-PRA achieves improvements of 5.0% and 6.5% over standard PRMs on the Qwen-VL series, with Qwen3-VL attaining 54.74% success rate on AndroidWorld. On the offline OS-Critic Bench, GUI-PRA demonstrates strong competitiveness against fully trained critic models.
While progress in GUI agents has been largely driven by industrial-scale training, ungrounded hallucinations often trigger cascading failures in real-world deployments.Unlike general VLM domains, the GUI agent field lacks a hallucination-focused suite for fine-grained diagnosis, reliable evaluation, and targeted mitigation.To bridge this gap, we introduce HalluClear, a comprehensive suite for hallucination mitigation in GUI agents as a complement to computation-intensive scaling. HalluClear comprises: (1) a GUI-specific hallucination taxonomy derived from empirical failure analysis; (2) a calibrated three-stage evaluation workflow which enhances VLM-as-a-judge reliability via expert-annotated benchmarking and ensemble credibility estimation; and (3) a mitigation scheme based on closed-loop structured reasoning, enabling lightweight continual post-training with cold-start initialization for both generalist and GUI-specialist agents. Experiments across representative agents and public benchmarks demonstrate that post-training on only 9K samples within our suite can significantly reduce hallucinations, thereby improving grounding and action fidelity, offering a compute-efficient pathway to robust GUI automation.
Chao Jin, Wenkui Yang, Hao Sun +6
MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences