GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
Figures & tables
Figure 1 : Overview of AutoGUIWorld. The method first samples a structured GUI world from an OS substrate, visual appearance space, and initial GUI state space. It then realizes a seed screenshot, generates seed-conditioned tasks, plans atomic actions, and rolls out a clean sequence of GUI visual states using Image2 as a visual world model.
Figure 2 : Voyager-guided closed-loop rendering. At each step, Voyager observes the current GUI image and receives the planned action. It produces a first-person thought, an action abstract, and an after-action rendering prompt. Image2 uses the current image as an image-to-image reference and renders the next GUI state, which is returned to Voyager for the next iteration. The bottom trajectory shows how one planned action sequence becomes a temporally consistent chain of rendered GUI states.
Figure 3 : Distribution of the AutoGUIWorld training set. The inner ring shows 79,266 training samples across Chrome, Ubuntu, Windows, and macOS. The outer ring and accompanying tables summarize the functional-category composition within each environment.
Table 1: Experimental configuration. Qwen features follow the visual path used to construct the model input. Auxiliary encoders and pixel statistics test whether the result depends on that representation.
Figure 4 : Distribution alignment in the Qwen3.5-9B visual-token space. Panel a reports the matched synthetic–ScaleCUA MMD normalized by the median real–real MMD; horizontal bars give group-bootstrap 95% confidence intervals and ρ=1 marks the empirical real-data scale. Panel b shows every available real–real distance, the matched synthetic–ScaleCUA estimate with its interval, and synthetic distances to the remaining real sources.
Figure 5 : Exploratory DINOv2 MMD using 500 images per source. Gray and blue spans collect the available real–real and synthetic–real source-pair distances within each domain; single available comparisons appear as points. Ubuntu is the only domain whose synthetic–real range lies wholly above the observed real–real range.
Figure 6 : Low-level appearance and generation-residual diagnostics. Panel a places each synthetic random-patch statistic against the range of real sources in the same domain; the gray spans show real-source ranges, and OCR denotes recognition confidence. Panel b reports content-matched ROI C2ST AUC; 0.5 is chance. Text and icon patches remain near chance, while Ubuntu flat backgrounds retain the clearest residual.
Figure 7 : Static interface complexity across OmniParser confidence thresholds. Panel a compares the rasterized union coverage of synthetic and ScaleCUA screenshots; the annotated percentage is the relative gain at the middle threshold, 0.15. Panel b reports the element-count difference, synthetic minus ScaleCUA. Coverage is higher in all twelve comparisons, including Windows where the detector returns fewer synthetic elements.
Figure 8 : Element detection and OCR examples on Windows interfaces. Each row shows the same full-screen view without annotations, with blue OmniParser boxes at confidence threshold 0.15, and with teal OCR regions. Row summaries report full-image counts and mean OCR confidence.
Domain
N
Steps
Action types
Entropy
Open apps
Rollout failure
Ubuntu
2168
11.9 (4/10/24)
4.33
.79
1.62 (2/2)
0%
Windows
1394
11.4 (2/8/25)
4.05
.82
2.06 (2/3)
.29%
macOS
591
7.1 (2/6/14)
3.95
.92
2.50 (2/3)
.68%
Table 2 : Synthetic trajectory structure. Step percentiles are P10/P50/P90; application percentiles are P50/P90. Rollout failure is the fraction of rendered terminal states not marked as success.
Figure 9 : Task execution scores for Qwen3.5-35B-A3B and AGW-35B (checkpoint 469), with all OSWorld and WAA domains. The top Overall averages the four benchmark scores equally. Within each benchmark, Overall averages task scores, retaining partial credit on OSWorld and WAA. Gray shows base scores, blue shows gains, and hatching marks decreases. Blue ticks mark AGW-35B scores. n counts tasks, and gains are computed before rounding.
Figure 10 : All task domains on macOSWorld and ScienceBoard. Gray bars show base success rates, blue extensions show gains, and gray hatching marks decreases. Blue ticks identify AGW-35B scores. Overall uses all tasks in each benchmark. ScienceBoard uses the 143-task subset in Appendix C.2 .
Figure 11 : Post-training performance of AGW-35B (blue circles) and the AgentNet-trained model (orange dashed lines and squares). Gray marks the base, shared across all four benchmarks. Curves use actual training steps and panel-specific score ranges.
Figure 12 : ScreenSpot-Pro accuracy for Qwen3.5-35B-A3B and AGW-35B (checkpoint 469). Gray shows base scores and blue shows training gains. The top group reports overall, icon, and text accuracy; each domain reports icon and text separately. n counts all examples in each group.
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Definition
Arguments
key
Presses keys in order and releases them in reverse order.
keys
key_down
Holds the specified keys until release.
keys
key_up
Releases the specified keys in reverse order.
keys
type
Enters the specified text.
text
mouse_move
Moves the pointer to the target location.
C(x,y)
left_click
Left-clicks at the target location.
C(x,y)
Appendix
Table 3: Unified action space for desktop and Chrome browser trajectories. C(x,y) denotes a screen-coordinate pair; C1 and C2 denote the drag start and end points, respectively. The keys argument is an array, text is a string, pixels is a signed scroll amount, and time is a duration in seconds. The status argument is either success or failure .
Benchmark
Tasks
Evaluation subset
Metric (%)
OSWorld
361
test_nogdrive : excludes eight Google Drive tasks
Mean task score
WAA
154
test_all
Mean task score
macOSWorld
231
English task and interface, including 29 safety tasks
Task success
ScienceBoard
143
Five application domains, excluding Lean
Task success
ScreenSpot-Pro
1,581
English instructions, positive targets
Grounding accuracy
Appendix
Table 4: Evaluation task sets and scoring units. Each checkpoint uses the same task set within a benchmark. Counts describe the evaluated samples, not model results.
The predicate formalises a different coprimality condition.
Vacuous proof
B-02
Contradictory premises permit a proof without the intended argument.
Executability not enforced
E-01, E-02
The evaluators accept noncomputable completions for tasks intended to test executable decision procedures.
Appendix
Table 5: Lean task issues reported in ScienceBoard issue #7. The 13 flagged tasks are part of the 26-task Lean domain excluded from our evaluation. Task identifiers are local to Lean.
Domain
Real–real MMD
Synthetic–ScaleCUA
Synthetic–other real
ρ
Ubuntu
.117/.134/.154
.166 [.157,.208]
.177/.266
1.24 [1.18,1.56]
Windows
.307
.261 [.244,.284]
.293
.85 [.79,.93]
Web
.092/.145/.150
.165 [.154,.185]
.184/.188
1.14 [1.07,1.28]
macOS
.351
.208 [.193,.237]
.355
.59 [.55,.67]
Appendix
Table 6 : MMD in the Qwen3.5-9B visual-token space. Real–real entries list all available pairwise distances in ascending order. “Other real” lists distances from the synthetic source to non-ScaleCUA real sources. Brackets give group-bootstrap 95% confidence intervals.
Domain
Real–real MMD
Synthetic–real MMD
Relation
Ubuntu
.035–.071
.092–.197
1.3–5.6 × real–real
Web
.050
.047–.059
Same scale
Windows
.156
.076–.124
Below real–real
macOS
.087–.240
.139–.202
Inside real–real range
Appendix
Table 7: Exploratory DINOv2 MMD using 500 images per source. Ranges collect the available source-pair distances within each domain.
Domain
High-frequency power
Peak kurtosis
Edge width
OCR confidence
Ubuntu
16.85 / 17.11–17.33
141.9 / 118–132
2.07 / .74–1.27
.825 / .69–.83
Web
18.05 / 17.19–17.81
110.6 / 75–92
.88 / .70–1.34
.864 / .63–.80
Windows
16.93 / 17.04–17.08
90.6 / 105–134
1.55 / 1.43–5.02
.786 / .48–.49
macOS
17.12 / 16.74–17.62
85.4 / 95–132
1.76 / 1.45–1.94
.757 / .55–.88
Appendix
Table 8 : Random-patch diagnostics, reported as synthetic / real-source range. OCR confidence is reported for each domain.
Table 10 : Static complexity at confidence thresholds 0.05/0.15/0.25. Each cell lists values in that order. Coverage is the rasterized union of detected boxes.
Figure 13 : Detection and OCR outputs for four selected AutoGUIWorld Windows screenshots. All three columns use the same full-screen field of view.
Figure 14 : Detection and OCR outputs for four selected ScaleCUA Windows screenshots: Illustrator, Unreal Engine, Excel, and Photoshop.
Figure 15 : Step-level quality control with a Gemini VLM checker. Each request combines the task and exact action context with an ordered set of visual evidence. Gemini evaluates grounding, action validity, and visual-transition fidelity, then returns a strict structured verdict. Code validates the per-dimension values, recomputes the aggregate label, and routes the step to target-box repair, removal, transition filtering, or retention. Repaired target boxes are evaluated again before retention.
Group
Dimension
Criterion
Grounding
target_exists
The intended target is visible in the pre-action screenshot.
Grounding
box_hits_target
The annotated box encloses the intended target rather than another element or empty space.
Grounding
box_tightness
The target box has a reasonable spatial extent around the interactive element.
Action
action_valid_here
The action is executable and meaningful in the current GUI state.
Transition
obs_transition_ok
The generated post-action screenshot shows a visual consequence consistent with the action.
Appendix
Table 11: Image–action dimensions used by the transition-quality audit.
Figure 16 : AutoGUIWorld training-data filtering. The proportional ribbons track Ubuntu, Windows, macOS, and Chrome samples through VLM quality control, target-box repair, and transition filtering. Desktop counts are measured directly. Chrome contributes 35,209 retained samples; its preceding stage counts are inferred using the desktop-stage retention ratios. The resulting training set contains 79,266 samples.
OS
Evaluable
Fail
Rate
Filtered
Fail
Rate
Ubuntu
23,625
731
3.09%
21,762
516
2.37%
Windows
14,429
479
3.32%
13,541
286
2.11%
macOS
4,472
83
1.86%
4,048
16
0.40%
All
42,526
1,293
3.04%
39,351
818
2.08%
Appendix
Table 12: Transition-fidelity audit by operating system. Overall rates use all evaluable transitions. Filtered rates remove steps with an invalid action, absent target, or incorrect target grounding.
Figure 17 : Representative Image2 transition failures. Each case shows the same 16:9 crop before and after an atomic action, with the inconsistent region highlighted in blue. The examples show terminal-command substitution, premature rendering of a styled artifact, invented table structure and values, and persistent code mutation following a cursor-movement shortcut.
Figure 18 : Application and website coverage in the AutoGUIWorld training set. The desktop panels include all application entries for Ubuntu, Windows, and macOS; the Chrome panel shows 67 websites across all 14 categories. Items are grouped by environment and function. Numbers beneath icons indicate training samples; website-category headings report the number of domains.
Figure 19 : Ubuntu desktop trajectories (1/2).
Figure 20 : Ubuntu desktop trajectories (2/2).
Figure 21 : Windows 11 trajectories (1/2).
Figure 22 : Windows 11 trajectories (2/2).
Figure 23 : macOS trajectories (1/2).
Figure 24 : macOS trajectories (2/2).
Figure 25 : Chrome trajectories (1/2).
Figure 26 : Chrome trajectories (2/2).
Figure 27 : Hierarchical coverage of professional application domains.
Figure 28 : Engineering CAD and 3D authoring workflows (1/2).
Figure 29 : Engineering CAD and 3D authoring workflows (2/2).
Figure 30 : Visual and media design workflows (1/2).
Figure 31 : Visual and media design workflows (2/2).
Figure 32 : Data analysis and econometrics workflows.