Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
Figures & tables
Figure 1: Motivation and preliminaries. (a) Comparison of GRPO, OPSD and ComputerSD. During online training, we measure (b) suppressed ratio (the ratio of tokens at correct steps whose log-probability is lowered) under fixed guidance and real-time feedback and (c) conflict ratio (the ratio of tokens whose shift contradicts the step-level judgment) in correct and incorrect steps.
Figure 2: Overview of ComputerSD. The base policy samples trajectories online, while a GUI analyzer provides step-level value scores and guidance. Privileged rescoring produces token-level supervision, which is combined with trajectory-level GRPO to update the policy.
Figure 3: Illustration of the fully asynchronous online training framework.
Model
Type
Max Steps
Success Rate (%)
Proprietary Models
OpenAI CUA ( OpenAI, 2025 )
Specialized
50
31.3
Seed1.5-VL ( Guo et al., 2025 )
General
100
36.7
Step-GUI-8B ( Yan et al., 2025 )
Specialized
100
40.2
Qwen3-VL-Flash ( Bai et al., 2025 )
General
100
41.6
UI-TARS-1.5 ( Qin et al., 2025 )
Specialized
100
42.5
Table 1: Performance comparison on OSWorld-Verified. Success rate is reported as Pass@1. Our experimental results are averaged over three independent evaluation runs to mitigate variance in online environments, while results for other models are taken from their official reports.
Model
OSWorld-Verified
Windows AgentArena
In-Domain
OOD
Qwen3-VL-8B-Thinking
40.5
23.0
19.2
w/ GRPO
48.2 +7.7
21.6 -1.4
23.3 +4.1
[][1.6em] w/ ComputerSD
47.5 +7.0
27.5 +4.5
24.5 +5.3
EvoCUA-8B
47.3
31.6
24.4
w/ GRPO
52.2 +4.9
30.2 -1.4
24.2 -0.2
Table 2: In-domain and out-of-distribution performance on OSWorld-Verified and cross-platform benchmark WindowsAgentArena.
Figure 6Figure 7Table 8
Figure 8: Training dynamics of ComputerSD and its reverse-gate variant. Left: Teacher–student gap. Right: Policy entropy.
λOPSD
SR (%)
0.1
35.6
0.01
39.8
0.001
37.1
Table 5: Effect of the OPSD loss coefficient on OSWorld-Verified.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Table 11
Category
Count
Percentage (%)
Successful trajectories
770
43.4
Failed trajectories
1,006
56.6
Correct steps
13,717
40.1
Incorrect steps
20,471
59.9
Appendix
Table 8: Statistics of the GUI analyzer SFT dataset.
Figure 9: System prompt for expert annotation.
Model
Expert agreement (%)
Valid output rate (%)
Qwen3-VL-8B-Thinking
29.9
41.6
GUI analyzer (SFT)
85.3
99.7
Appendix
Table 9: GUI analyzer evaluation on 341 held-out samples.
Method
SR (%)
Base model
33.8
Prompt-only method
34.6
GRPO with PRM
37.4
ComputerSD
39.8
Appendix
Table 10: Additional baseline comparisons on OSWorld-Verified.
Figure 10: System prompt used for trajectory sampling and evaluation.
Computer use agents (CUAs) have shown strong potential for automating complex digital workflows, yet their training remains constrained by costly live environment interaction and limited high-quality supervision. Existing filtered behavior cloning pipelines suffer from imitation bottlenecks, including distribution shift from the expert demonstration and the absence of negative learning signals. Meanwhile, standard trajectory-level reinforcement learning struggles with sparse rewards, ambiguous credit assignment, and high infrastructure costs for long-horizon GUI interaction. In this work, we propose PRO-CUA, a process-reward optimization framework for training CUAs with iterative step-level reinforcement learning. PRO-CUA decouples on-policy environment interaction from policy optimization: the current policy collects states through live rollouts, generates diverse candidate actions for each state, receives step-level feedback from a process reward model (PRM), and is optimized with group-relative advantages. This design enables dense and flexible credit assignment without relying on golden answers or offline expert trajectories, while reducing distribution shift by training on the agent's own execution states. Experiments on live web benchmarks demonstrate the effectiveness of PRO-CUA and the reliability of PRM-guided step-level training.
Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-managed observations, sparse terminal rewards, variable-length trajectories, and slow environment feedback. EvoCUA-1.5 addresses these challenges with Step-Level Policy Optimization (STEPO), which preserves trajectory-level advantage balance after decomposition into step-level samples; policy-aware filtering and pass-rate calibration over verifiable synthesized tasks; Dynamic Tri-Adaptive Curriculum (DTAC), which combines learnable tasks, difficult positive replay, and controlled infeasible-task exposure; and a fully asynchronous RL infrastructure with staleness control and mini-group batching. Experiments show that these components improve training stability and downstream performance. EvoCUA-1.5 achieves 63.2% success on OSWorld-Verified, outperforming comparable 32B/35B-scale open-weight baselines and even approaching models with significantly larger parameter counts. Overall, EvoCUA-1.5 provides a practical framework for scaling online RL in multi-turn computer-use agents.
Mianqiu Huang, Taofeng Xue, Chong Peng +12
1Meituan · 2Fudan University · 3Shanghai Jiao Tong University +1
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.
Chunji Lv, Yangguang Wei, Junlin Liu +6
Beijing Institute of Technology · Meituan · Institute of Automation, Chinese Academy of Sciences +1