Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
Figures & tables
Figure 1: Motivation and preliminaries. (a) Comparison of GRPO, OPSD and ComputerSD. During online training, we measure (b) suppressed ratio (the ratio of tokens at correct steps whose log-probability is lowered) under fixed guidance and real-time feedback and (c) conflict ratio (the ratio of tokens whose shift contradicts the step-level judgment) in correct and incorrect steps.
Figure 2: Overview of ComputerSD. The base policy samples trajectories online, while a GUI analyzer provides step-level value scores and guidance. Privileged rescoring produces token-level supervision, which is combined with trajectory-level GRPO to update the policy.
Figure 3: Illustration of the fully asynchronous online training framework.
Model
Type
Max Steps
Success Rate (%)
Proprietary Models
OpenAI CUA ( OpenAI, 2025 )
Specialized
50
31.3
Seed1.5-VL ( Guo et al., 2025 )
General
100
36.7
Step-GUI-8B ( Yan et al., 2025 )
Specialized
100
40.2
Qwen3-VL-Flash ( Bai et al., 2025 )
General
100
41.6
UI-TARS-1.5 ( Qin et al., 2025 )
Specialized
100
42.5
Table 1: Performance comparison on OSWorld-Verified. Success rate is reported as Pass@1. Our experimental results are averaged over three independent evaluation runs to mitigate variance in online environments, while results for other models are taken from their official reports.
Model
OSWorld-Verified
Windows AgentArena
In-Domain
OOD
Qwen3-VL-8B-Thinking
40.5
23.0
19.2
w/ GRPO
48.2 +7.7
21.6 -1.4
23.3 +4.1
[][1.6em] w/ ComputerSD
47.5 +7.0
27.5 +4.5
24.5 +5.3
EvoCUA-8B
47.3
31.6
24.4
w/ GRPO
52.2 +4.9
30.2 -1.4
24.2 -0.2
Table 2: In-domain and out-of-distribution performance on OSWorld-Verified and cross-platform benchmark WindowsAgentArena.
Figure 6Figure 7Table 8
Figure 8: Training dynamics of ComputerSD and its reverse-gate variant. Left: Teacher–student gap. Right: Policy entropy.
λOPSD
SR (%)
0.1
35.6
0.01
39.8
0.001
37.1
Table 5: Effect of the OPSD loss coefficient on OSWorld-Verified.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Table 11
Category
Count
Percentage (%)
Successful trajectories
770
43.4
Failed trajectories
1,006
56.6
Correct steps
13,717
40.1
Incorrect steps
20,471
59.9
Appendix
Table 8: Statistics of the GUI analyzer SFT dataset.
Figure 9: System prompt for expert annotation.
Model
Expert agreement (%)
Valid output rate (%)
Qwen3-VL-8B-Thinking
29.9
41.6
GUI analyzer (SFT)
85.3
99.7
Appendix
Table 9: GUI analyzer evaluation on 341 held-out samples.
Method
SR (%)
Base model
33.8
Prompt-only method
34.6
GRPO with PRM
37.4
ComputerSD
39.8
Appendix
Table 10: Additional baseline comparisons on OSWorld-Verified.
Figure 10: System prompt used for trajectory sampling and evaluation.