PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
Organizations: University of Chinese Academy of Sciences · Harbin Institute of Technology · Shenzhen Loop Area Institute · The Hong Kong University of Science and Technology · The Chinese University of Hong Kong
Abstract
Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next symbolic UI layout given a candidate action, enabling early anomaly avoidance and ranking candidate actions by predicted reliability. Finally, a Pre-cognitive Execution Controller (PEC) fuses these priors and predictions, prioritizes handling of foreseen anomalies, and ensures execution robustness through a closed-loop error correction mechanism. For robust evaluation, we develop AutoTraj, an automatic data-generation engine, to construct InterfereBench, a benchmark for long-horizon tasks with strong disturbances. Experiments demonstrate that PrecogUI surpasses state-of-the-art methods on InterfereBench while maintaining competitive performance on public benchmarks. The code will be publicly available.
Figures & tables
| Method | InterfereBench-Low | InterfereBench-High | Average | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o | 76.7 | 22.1 | 69.0 | 9.3 | 73.4 | 3.9 | 65.0 | 1.2 | 71.0 | 9.1 |
| Gemini-2.5-Pro | 87.1 | 28.4 | 80.0 | 12.5 | 83.8 | 17.7 | 76.0 | 7.5 | 81.7 | 16.5 |
| Qwen-2.5-VL | 90.6 | 33.7 | 82.5 | 18.9 | 71.7 | 36.5 | 64.0 | 18.0 | 77.2 | 26.8 |
| OmniParser | 84.1 | 70.9 | 76.0 | 41.5 ( 29.4) | 70.6 | 27.3 | 63.0 | 12.8 ( 14.5) | 73.4 | 38.1 |
| InfiGUI-R1 | 88.9 | 73.6 | 81.5 | 45.8 ( 27.8) | 77.4 | 37.3 | 68.0 | 19.5 ( 17.8) | 79.0 | 44.0 |
| Method | AndroidControl-Low | AndroidControl-High | GUI-Odyssey | Average | ||||
|---|---|---|---|---|---|---|---|---|
| GPT-4o ( OpenAI et al. 2024 ) | 74.3 | 19.4 | 63.1 | 21.2 | 37.5 | 5.4 | 58.3 | 15.3 |
| Qwen-2.5-VL ( Bai et al. 2025b ) | 94.1 | 85.0 | 75.1 | 62.9 | 59.5 | 46.3 | 76.2 | 64.7 |
| UI-TARS-7B ( Qin et al. 2025 ) | 98.0 | 90.8 | 83.7 | 72.5 | 94.6 | 87.0 | 92.1 | 83.4 |
| SeeClick | 93.0 | 75.0 | 82.9 | 59.1 | 71.0 | 53.9 | 82.3 | 62.7 |
| OS-Atlas-4B ( Wu et al. 2025 ) | 91.9 | 80.6 | 84.7 | 67.5 | 83.5 | 56.4 | 86.7 | 68.2 |
| Method | Mobile | Desktop | Web | Avg | |||
|---|---|---|---|---|---|---|---|
| Text | Icon | Text | Icon | Text | Icon | ||
| GPT-4o | 30.5 | 23.2 | 20.6 | 19.4 | 11.1 | 7.8 | 18.8 |
| Gemini-2.0 | – | – | – | – | – | – | 84.0 |
| Qwen-2.5-VL | – | – | – | – | – | – | 84.7 |
| SeeClick | 78.0 | 52.0 | 72.5 | 30.0 | 55.7 | 32.5 | 53.4 |
| ShowUI | 92.3 | 75.5 | 76.3 | 61.1 | 81.7 | 63.6 | 75.1 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Type Acc. (%) | Bbox IoU (%) | Element F1 (%) |
|---|---|---|---|
| InterfereBench | 91.3 | 82.6 | 85.4 |
| AndroidControl | 88.7 | 79.1 | 82.3 |
| GUI-Odyssey | 87.2 | 77.8 | 80.9 |
| Overall | 89.4 | 80.2 | 83.2 |
| Method | 15-step | 50-step |
|---|---|---|
| GPT-4o ( OpenAI et al. 2024 ) | 5.0 | – |
| CogAgent-9B ( Hong et al. 2024 ) | 8.1 | – |
| OS-Atlas-7B ( Wu et al. 2025 ) | 14.6 | – |
| Aguvis-72B ( Xu et al. 2025b ) | 17.0 | – |
| UI-TARS-7B-SFT ( Qin et al. 2025 ) | 17.7 | – |
| UI-TARS-7B-DPO ( Qin et al. 2025 ) | 18.7 | – |
| Dataset | Episodes w/ recovery | Success after recovery | Episodes requiring restart |
|---|---|---|---|
| AndroidControl-High | 12.4% | 69.7% | 1.8% |
| GUI-Odyssey | 15.3% | 73.1% | 2.4% |
| InterfereBench | 18.6% | 71.2% | 2.7% |