cs.CVSep 30, 2026

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Authors: Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen

Organizations: Zhejiang University · Ant Group

Abstract

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PRO-CUA: Process-Reward Optimization for Computer Use Agents

    May 27, 2026Yifei He, Rui Yang, Hao Bai +2Process Reward ModelOn-Policy

  2. EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

    Jul 7, 2026Mianqiu Huang, Taofeng Xue, Chong Peng +12Computer-Use AgentsOffline Reinforcement Learning

  3. PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

    Aug 3, 2026Chunji Lv, Yangguang Wei, Junlin Liu +6Unsupervised On-Policy Self-DistillationAgentic Reinforcement Learning