Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire. While human demonstrations are unscalable, existing GUI world models rely on text descriptions or HTML rendering, discarding crucial pixel-level visual details like icons and layout styles. To address this issue, we introduce Infinite-Dreamer, a simulation-free data synthesis method powered by a pixel-level Image Editing World Model. By conceptualizing GUI transitions as image editing tasks, we leverage Vision-Language Models (VLMs) to describe action-induced UI changes as structured delta-text. We then fine-tune an image editing backbone to controllably synthesize realistic screenshot transitions. We utilize this model to generate both single-frame visual robustness data and multi-step imaginary trajectories. To validate the effectiveness of our approach, we fine-tune the Qwen3-VL baseline solely on the synthesized data to obtain Infinite-Actor, and evaluate it on AndroidWorld, MobileWorld, and AndroidControl-Curated benchmarks. Infinite-Actor consistently outperforms the Qwen3-VL baselines across scales: Infinite-Actor-8B improves AndroidWorld Pass@1 by +4.45 and nearly doubles the MobileWorld Pass@3 success rate, while Infinite-Actor-2B improves Pass@1 by +9.05. Code is available at https://github.com/swaydy-n/Infinite-Dreamer.
Figures & tables
Figure 1: Motivation. Post-action screenshots predicted by different world models: (1) ground truth; (2) HTML-vision rendering, lacking visual detail; (3) the pretrained FireRed-Image-Edit-1.0, failing to capture GUI-specific transitions; and (4) our Infinite-Dreamer-9B, producing high-fidelity, pixel-level screenshots faithful to the ground truth.
Figure 2: Overview of our proposed Infinite-Dreamer : (a) Data Construction and Delta-Text Image-Editing World Model Training. We extract structured delta-texts from real UI trajectories and fine-tune an instruction-guided image-editing model as a GUI world model; (b) Data Generative Synthesis and GUI Agent Training. The trained world model enables two synthesis regimes, namely single-frame augmentation and multi-step imagination, whose outputs are filtered by a VLM-based quality checker and used to train the GUI agent.
Method
AndroidWorld
MobileWorld
AC-Curated-Hard
AC-Curated-Easy
Pass@1
Pass@3
Pass@1
Pass@3
Type Acc
Box Acc
Step SR
Type Acc
Box Acc
Step SR
General-Purpose Models
Claude-Sonnet-4.5
19.25
25.86
33.92
42.11
73.40
22.20
30.60
80.90
22.60
35.00
Gemini-3-Flash
34.91
56.03
43.27
52.63
73.10
48.10
46.00
84.40
72.40
68.20
MiniMax-M3
10.34
17.24
14.03
23.68
70.90
47.40
41.60
75.00
45.60
42.50
MiMo-V2.5
30.46
49.14
36.55
52.63
73.20
63.40
53.70
81.40
77.10
68.00
Table 1: GUI agent performance evaluation. Infinite-Actor (Ours) is trained solely on Infinite-Dreamer synthesis data and paired with its base-model Qwen3-VL counterpart of the same scale; green ↑ / red ↓ indicates whether Infinite-Actor is higher / lower than its counterpart. The best and second-best results in each column are bolded and underlined , respectively.
Single-Step Fidelity
Multi-Step Rollout
Method
Acc.
Comp.
Cont.
Vis.
1 Step
5 Steps
10 Steps
FLUX.1 Kontext (12B)
33.33
33.01
36.05
34.34
34.18
22.16
27.69
Step1X-Edit-v1.2 (12B)
34.27
34.11
36.85
30.60
33.96
48.95
39.61
LongCat-Image-Edit (6B)
33.79
32.66
36.34
40.65
35.86
20.89
32.69
Qwen-Image-Edit-2511 (20B)
50.30
60.50
57.44
55.52
55.94
39.44
26.15
FireRed-Image-Edit-1.0 (20B)
64.34
67.73
71.24
58.89
65.55
47.46
40.96
Table 2: World-model synthesis quality. Single-step instruction fidelity on the AndroidControl validation set and multi-step rollout quality, both scored across four VLM-judged dimensions (Accuracy, Completeness, Contextual alignment, Visual preservation). Higher is better.
Figure 3: AndroidWorld analyses for base-4b, real-4b, and mix-4b.
Real Data
Gen Data
Aug Data
SR (%)
×
×
×
30.17
×
✓
✓
39.37
✓
×
×
42.67
✓
✓
×
44.83
✓
×
✓
47.41
✓
✓
✓
48.28
Table 3: Data source ablation on AndroidWorld Pass@1 (%), using Qwen3-VL-4B as the fixed backbone. Real Data: real collected trajectories; Gen Data: multi-step imagination from the world model; Aug Data: single-frame visual augmentation.
Figure 4: Qualitative synthesis comparison. (a) Ground Truth trajectory (top) versus (b) Infinite-Dreamer-9B (Ours) (bottom) over a 6-step GUI interaction sequence (from opening “Abercrombie” to clicking the search icon). Our model faithfully reproduces the real UI state transitions, maintaining layout accuracy, text fidelity, and visual coherence across all steps.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Backbone
FLUX.2-Klein (4B / 9B)
Training data
AndroidControl train split
Fine-tuning method
LoRA
LoRA rank
32
Training epochs
3
Batch size
1
Appendix
Table 5: World model training configuration. Both 4B and 9B models share the same hyperparameters unless otherwise noted.
Configuration Item
Value
Image Width
672 px
Image Height
1536 px
Inference Steps
40 steps
Guidance Scale
4.0
Appendix
Table 6: Inference configuration for performance testing.
Metric
Infinite-Dreamer-9B
Infinite-Dreamer-4B
Average Time per Image
57.202 s
28.779 s
Average Time per Step
1430.0 ms
719.5 ms
Throughput (Img/s)
0.01748
0.03475
Pixel Throughput (Mpx/s)
0.01804
0.03587
Appendix
Table 7: Inference performance comparison between FLUX.2-Klein 9B and 4B models.
Configuration
Value
Backbone
Qwen3-VL (2B / 4B / 8B)
Training data
Infinite-Dreamer synthesis data (aug, gen)
Fine-tuning method
Full fine-tuning
Training steps
1,000
Batch size
32
Learning rate
1.0×10−5
Appendix
Table 8: GUI agent training configuration. All model sizes (2B/4B/8B) share the same hyperparameters.
Figure 5: Performance comparison between zero-shot Qwen3-VL and Infinite-Actor agents across model scales (2B/4B/8B) on six evaluation metrics from Table 1 : AW Pass@1 (AndroidWorld Pass@1), AW Pass@3, MW Pass@1 (MobileWorld Pass@1), MW Pass@3, AC-Easy Step SR (AC-Curated-Easy Step SR), and AC-Hard Step SR (AC-Curated-Hard Step SR). Infinite-Actor models consistently outperform their zero-shot counterparts on task-level benchmarks.
Figure 6: Qualitative effect of delta-text conditioning. Given the source UI (leftmost) and the delta-text action “Open LeafSnap App”, Infinite-Dreamer-9B (third) closely matches the ground-truth transition (second), whereas the variant without delta-text (rightmost, which only receives the trigger action/operation without any fine-grained descriptions of added or removed elements) collapses into an unrelated screenshot that diverges substantially from the intended target UI.
Table 10: Action space for mobile GUI agents. Each action type is listed with its required and optional parameters. Coordinates are specified as (x,y) pixel positions on a 1000×1000 screen.
Rollout Step ( t )
1
2
3
4
5
6
7
8
9
10
Step Validity (%)
84.6
84.9
82.8
84.7
79.0
78.4
80.4
79.4
81.2
79.5
Cumulative Survival (%)
84.6
72.0
60.6
50.3
39.8
30.4
24.1
18.7
14.1
9.3
Appendix
Table 11: Survival of synthesized rollouts under VLM quality filtering ( η=90 ) across horizons t=1,…,10 , measured over 1,599 imagined trajectories. Step validity : share of evaluated steps at horizon t that pass. Cumulative survival : share of trajectories reaching horizon t whose steps 1,…,t all pass. Step 0 (the initial page) is excluded.
Figure 7: Qualitative comparison of single-step UI editing. Three examples (rows) show the source UI and outputs from FLUX.1 Kontext, Step1X-Edit-v1.2, LongCat-Image-Edit, Qwen-Image-Edit-2511, FireRed-Image-Edit-1.0, and Infinite-Dreamer-9B (rightmost). Infinite-Dreamer-9B most faithfully executes the requested edits while preserving layout and unmodified regions.
Figure 8: Qualitative comparison of multi-step rollout. (a) Part I shows outputs from FLUX.1 Kontext, Step1X-Edit-v1.2, and LongCat-Image-Edit for the five-step LeafSnap watering task; each column is a step. These baselines drift away from the intended UI state.
Figure 8: (b) Part II shows outputs from Qwen-Image-Edit-2511, FireRed-Image-Edit-1.0, and Infinite-Dreamer-4B (Ours) for the same rollout. Infinite-Dreamer-4B preserves visual coherence while the baselines introduce artifacts and layout drift.
Figure 8: (c) Part III compares the ground-truth trajectory with Infinite-Dreamer-9B (Ours). Infinite-Dreamer-9B closely follows the ground-truth layout and text across all five steps.
Figure 9: Single-frame augmentation examples. For each pair, the left column shows the source screenshot and the right column shows a visually distinct variant produced by Infinite-Dreamer. The pairs span multiple app categories and cover dark/light theme swap, accent-color and background recoloring, and icon/keyboard restyling, while widget positions and text content are held fixed.
Figure 10: Multi-step imagination examples. Each row is an independent trajectory: starting from an Initial screenshot, Infinite-Dreamer autoregressively applies the delta-text action printed above each column (e.g., “Click the Create new folder Button”, “Input the Text daily_goals ”, “Click the Send Button”) to produce Step 1 through Step 5. The imagined screenshots keep the surrounding UI chrome stable while correctly rendering the widget or text that the action introduces.
Failure Mode
Share
Operational Definition (from the evaluation log)
Step-budget exhaustion
41.2% (7/17)
Ran to the allocated interaction budget ( 10× task complexity) while still acting; no terminal action was issued.
False completion
35.3% (6/17)
Issued a terminate call with goal_status = complete , yet the task verifier marked the episode failed.
Incorrect final answer
17.6% (3/17)
Ended with an answer action whose response failed verification (three information-retrieval tasks).
Abrupt stop
5.9% (1/17)
Episode ended mid-interaction; the log records no terminate or answer call.
Appendix
Table 12: Measured failure outcomes of mix-4b on the 19 AndroidWorld Hard tasks (17 failures, one episode per task). Categories are mutually exclusive and are assigned from the terminal event recorded in the evaluation log.
Figure 11: Representative failure cases of Infinite-Dreamer. (a) Long-horizon drift: over many generated steps, visual inconsistencies such as garbled text, shifted buttons, and spurious elements accumulate. (b) Dense UI challenges: tightly packed widgets (e.g., grid layouts, complex lists) can cause the model to merge adjacent elements or misplace insertions and removals.
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. Existing pipelines often rely on costly human demonstrations or on-policy framework, which tends to over-sample common flows while missing rare transitions and complex multi-step procedures. To address this problem, we propose SEE, a two-stage data synthesis framework consisting of (i) an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and (ii) a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling. This design yields reproducible and explainable data generation, while explicitly preventing spurious cycles and enabling long-horizon composition. Across multiple real-world apps, SEE produces trajectories with an average length of 14.8 steps while avoiding spurious loops, and agents fine-tuned on SEE achieve improved task success and generalization to unseen screens. We will publicly release our synthesis code and dataset.
Zhuohang Fan, Beichen Zhang, Yuanfa Li +4
Harbin Institute of Technology, Weihai, China · Harbin Institute of Technology (Weihai) Qingdao Research Institute, Qingdao, China
Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarcity of large-scale training data spanning diverse real-world applications. Existing datasets rely heavily on costly manual annotations and are typically confined to narrow domains. To address this challenge, we propose Video2GUI, a fully automated framework that extracts grounded GUI interaction trajectories directly from unlabeled Internet videos. Video2GUI employs a coarse-to-fine filtering strategy to identify high-quality GUI tutorial videos and convert them into structured agent trajectories. Applying this pipeline to 500 million video metadata entries, we construct WildGUI, a large-scale dataset containing 12 million interaction trajectories spanning over 1,500 applications and websites. Pre-training Qwen2.5-VL and Mimo-VL on WildGUI yields consistent improvements of 5-20% across multiple GUI grounding and action benchmarks, matching or surpassing state-of-the-art performance. We will release both the WildGUI dataset and the Video2GUI pipeline to support future research of GUI agents.
Weimin Xiong, Shuhao Gu, Bowen Ye +5
National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University