AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
Organizations: Hunyuan AI Data Team
Abstract
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
Figures & tables
| Component | Configuration |
| Qwen representation | Qwen3.5-9B PatchMerger output; merge factor 2; 4096-dimensional merged tokens; image-level mean pooling |
| Image preprocessing | smart_resize with max_pixels , matching the data pipeline |
| MMD | Per-domain joint z-score; RBF kernel; one bandwidth selected from the median heuristic over the union of sources in that domain |
| C2ST | GroupKFold logistic-regression probe; AUC 0.5 denotes chance-level discrimination |
| Sampling | About 1,500 images per source for Qwen embeddings; 500 for DINOv2; 80 for OCR; 200–500 for element detection |
| Pixel diagnostics | 300 images/source 20 random patches; 250 images/source 12 artifact-probe patches |
| Domain | Steps | Action types | Entropy | Open apps | Rollout failure | |
| Ubuntu | 2168 | 11.9 (4/10/24) | 4.33 | .79 | 1.62 (2/2) | 0% |
| Windows | 1394 | 11.4 (2/8/25) | 4.05 | .82 | 2.06 (2/3) | .29% |
| macOS | 591 | 7.1 (2/6/14) | 3.95 | .92 | 2.50 (2/3) | .68% |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Action | Definition | Arguments |
| key | Presses keys in order and releases them in reverse order. | keys |
| key_down | Holds the specified keys until release. | keys |
| key_up | Releases the specified keys in reverse order. | keys |
| type | Enters the specified text. | text |
| mouse_move | Moves the pointer to the target location. | |
| left_click | Left-clicks at the target location. |
| Benchmark | Tasks | Evaluation subset | Metric (%) |
| OSWorld | 361 | test_nogdrive : excludes eight Google Drive tasks | Mean task score |
| WAA | 154 | test_all | Mean task score |
| macOSWorld | 231 | English task and interface, including 29 safety tasks | Task success |
| ScienceBoard | 143 | Five application domains, excluding Lean | Task success |
| ScreenSpot-Pro | 1,581 | English instructions, positive targets | Grounding accuracy |
| Reported issue | Task identifiers | Effect on evaluation |
| False theorem statements | A-02, A-04, B-05, B-06, B-07, C-05, D-02, D-03, D-07 | The formal statements admit counterexamples. |
| Incorrect objective | B-01 | The predicate formalises a different coprimality condition. |
| Vacuous proof | B-02 | Contradictory premises permit a proof without the intended argument. |
| Executability not enforced | E-01, E-02 | The evaluators accept noncomputable completions for tasks intended to test executable decision procedures. |
| Domain | Real–real MMD | Synthetic–ScaleCUA | Synthetic–other real | |
| Ubuntu | .117/.134/.154 | .166 [.157,.208] | .177/.266 | 1.24 [1.18,1.56] |
| Windows | .307 | .261 [.244,.284] | .293 | .85 [.79,.93] |
| Web | .092/.145/.150 | .165 [.154,.185] | .184/.188 | 1.14 [1.07,1.28] |
| macOS | .351 | .208 [.193,.237] | .355 | .59 [.55,.67] |
| Domain | Real–real MMD | Synthetic–real MMD | Relation |
| Ubuntu | .035–.071 | .092–.197 | 1.3–5.6 real–real |
| Web | .050 | .047–.059 | Same scale |
| Windows | .156 | .076–.124 | Below real–real |
| macOS | .087–.240 | .139–.202 | Inside real–real range |
| Domain | High-frequency power | Peak kurtosis | Edge width | OCR confidence |
| Ubuntu | 16.85 / 17.11–17.33 | 141.9 / 118–132 | 2.07 / .74–1.27 | .825 / .69–.83 |
| Web | 18.05 / 17.19–17.81 | 110.6 / 75–92 | .88 / .70–1.34 | .864 / .63–.80 |
| Windows | 16.93 / 17.04–17.08 | 90.6 / 105–134 | 1.55 / 1.43–5.02 | .786 / .48–.49 |
| macOS | 17.12 / 16.74–17.62 | 85.4 / 95–132 | 1.76 / 1.45–1.94 | .757 / .55–.88 |
| Domain | Flat background | Text | Icon edge |
| Ubuntu | .65 | .44 | .49 |
| Windows | .54 | .57 | .50 |
| Union coverage | Element count | |||
| Domain | Synthetic | ScaleCUA | Synthetic | ScaleCUA |
| Ubuntu | .360/.292/.257 | .221/.170/.145 | 137/103/90 | 130/98/86 |
| Windows | .414/.351/.322 | .275/.212/.184 | 184/142/126 | 200/154/134 |
| Web | .425/.332/.277 | .286/.200/.162 | 103/68/56 | 69/48/41 |
| macOS | .432/.359/.321 | .394/.319/.292 | 165/122/105 | 110/85/77 |
| Group | Dimension | Criterion |
| Grounding | target_exists | The intended target is visible in the pre-action screenshot. |
| Grounding | box_hits_target | The annotated box encloses the intended target rather than another element or empty space. |
| Grounding | box_tightness | The target box has a reasonable spatial extent around the interactive element. |
| Action | action_valid_here | The action is executable and meaningful in the current GUI state. |
| Transition | obs_transition_ok | The generated post-action screenshot shows a visual consequence consistent with the action. |
| OS | Evaluable | Fail | Rate | Filtered | Fail | Rate |
| Ubuntu | 23,625 | 731 | 3.09% | 21,762 | 516 | 2.37% |
| Windows | 14,429 | 479 | 3.32% | 13,541 | 286 | 2.11% |
| macOS | 4,472 | 83 | 1.86% | 4,048 | 16 | 0.40% |
| All | 42,526 | 1,293 | 3.04% | 39,351 | 818 | 2.08% |
| Stage | Input | Output |
| Seed realization | Platform, appearance, sampled environment | Initial-screen description and visible elements |
| Task generation | Seed context, task history, sampling directives | Task instruction and task attributes |
| Meta Planner | Task, fixed seed, action vocabulary | High-level plan and atomic action sequence |
| Voyager and Image2 | Current screenshot, action, rollout progress | Step text, rendering instruction, next screenshot |
| Target grounding | Clean screenshot and element description | Target region or point |
| Quality checks | Seed, task, or transition evidence | Structured defect and consistency judgments |