GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
Figures & tables
Approach
GT supervision
Extra rollouts
Adaptation stage
Retained state
Related work
Demonstration SFT
Yes
No
Pre-deployment
Weights
Wu et al. (2025) ; Qin et al. (2025)
RL / self-evolution
Varies
Yes
Pre-deployment
Weights
Bai et al. (2024) ; Qi et al. (2025) ; Xiao et al. (2025)
Exploration memory
Varies
Yes
Pre-deployment
Memory
Zhang et al. (2025) ; Sun et al. (2026)
Retry-based improvement
Varies
Yes
Deployment
Memory
Shinn et al. (2023) ; Li et al. (2026a) ; He et al. (2026)
Multi-rollout inference
No
Yes
Deployment
None
Snell et al. (2024) ; Yang et al. (2026)
Multi-rollout adaptation
No
Yes
Deployment
Weights
Akyürek et al. (2025) ; Zuo et al. (2025)
Table 1: Where fully test-time adaptation sits among ways to improve a GUI agent. Each row is placed by the resources its learning uses: ground-truth supervision (demonstrations, an environment check or a human; Varies : some members use it, or a trained or VLM evaluator stands in), rollouts beyond the single attempt that counts, the stage at which adaptation happens, and the state retained across tasks. Our setting withholds the first two and admits only the deployment stage, with memory or weights as the retained state. The row marked † satisfies it. The last column lists representative works, and Sec. 5 discusses the rest.
Figure 1: Fully test-time adaptation for GUI agents, and Solo . Top: each task occurrence gets one attempt, and a judge decides from the instruction, screens and actions whether it was completed. Bottom: a judged success enters a sliding window as a full episode; for a judged failure, a proposer–verifier pair relabels a completed prefix, which enters the same window. While the window holds a judged success, one top- K self-distillation step updates the LoRA adapter.
WebArena
VisualWebArena
MobileWorld
Method
UI-TARS-7B
Qwen3-VL-8B
UI-TARS-7B
Qwen3-VL-8B
UI-TARS-7B
Qwen3-VL-8B
Frozen agent
20.8 ± 0.6
20.1 ± 0.8
16.7 ± 1.4
14.8 ± 0.6
6.1 ± 1.0
9.7 ± 1.0
AWM-online
22.7 ± 1.2
18.7 ± 1.6
16.3 ± 1.3
14.7 ± 1.2
9.7 ± 1.0
10.8 ± 1.7
DMS
22.4 ± 0.6
19.2 ± 0.5
17.0 ± 0.9
15.7 ± 0.9
7.5 ± 0.8
9.7 ± 0.5
Solo (ours)
25.8 ± 1.6
24.9 ± 0.9
20.0 ± 0.8
20.9 ± 1.9
9.2 ± 0.8
13.3 ± 2.2
Table 2: Main result. Success rate (%) over the stream under the benchmark’s own evaluator; mean ± std over three runs per cell. Rows per stream: WebArena 324, VisualWebArena 411, MobileWorld 120. AWM-online and DMS are the in-setting memory baselines of Sec. 2 , run on the identical streams.
Variant
VisualWebArena
WebArena
Mean
Frozen agent
14.8 ± 0.6
20.1 ± 0.8
17.5
Solo (full)
20.9 ± 1.9
24.9 ± 0.9
22.9
w/o window (one step per episode)
18.2 ± 1.2
23.0 ± 1.2
20.6
w/o top- K self-distillation (one-hot targets)
20.6 ± 2.6
22.0 ± 1.5
21.3
w/o hindsight relabeling (success branch only)
17.0 ± 1.4
21.9 ± 0.6
19.4
Table 3: Ablations. One component removed at a time from the full method, on the Qwen3-VL-8B agent: the window, replaced by one step per admitted episode; top- K self-distillation, replaced by one-hot imitation of the executed tokens; and hindsight relabeling, leaving the success branch alone. Success rate (%), mean ± std over three runs, and the mean over the two streams.
VisualWebArena
WebArena
Auxiliaries
Success (%)
Judge prec.
Success (%)
Judge prec.
Frozen agent (no auxiliaries)
14.8 ± 0.6
–
20.1 ± 0.8
–
gpt-5-mini
20.9 ± 1.9
0.51
24.9 ± 0.9
0.66
gpt-5
18.9 ± 1.4
0.59
23.0 ± 0.7
0.76
gpt-5-nano
19.1 ± 0.6
0.43
23.9 ± 1.7
0.46
the agent’s own model (Qwen3-VL-8B)
17.3 ± 0.8
0.35
24.2 ± 1.0
0.54
Table 4: Dependence on the auxiliary models. Solo on the Qwen3-VL-8B agent with the judge, the proposer and the verifier all replaced by the same alternative model. Success rate (%), mean ± std over three runs, and the judge’s precision, the share of the episodes it admitted that the benchmark evaluator scores as successes, averaged over the runs.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Stream
Arm
Wins per round (total)
One-step episodes
Entropy
VisualWebArena
frozen (mean of 3 runs)
22 / 18 / 21 (61)
36 / 35 / 33
–
entropy min., run 1
18 / 4 / 1 (23)
55 / 90 / 117
0.116 → 0.014
entropy min., run 2
22 / 10 / 4 (36)
48 / 53 / 83
0.113 → 0.012
entropy min., run 3
20 / 6 / 8 (34)
48 / 87 / 76
0.114 → 0.014
WebArena
frozen (mean of 3 runs)
20 / 22 / 23 (65)
5 / 6 / 6
–
entropy min., run 1
21 / 24 / 21 (66)
6 / 2 / 2
0.111 → 0.016
Appendix
Table 5: Entropy minimization as the test-time signal. Qwen3-VL-8B, both web streams, three runs each, against the frozen agent. Wins per round with the total, episodes ended at their first step per round, and the mean per-episode response entropy in nats over the first quarter of the stream and over the last.
Stream
Agent
Episodes
Succ.
Fail.
No prefix
Abstain
Guards
Verifier
Admitted
Prefix steps
WebArena
UI-TARS-7B
324
96
228
1
42
63 (53)
22
101 (44%)
3.7
WebArena
Qwen3-VL-8B
324
74
250
6
23
69 (51)
19
134 (54%)
3.0
VisualWebArena
UI-TARS-7B
411
94
317
2
67
133 (100)
24
91 (29%)
3.3
VisualWebArena
Qwen3-VL-8B
411
127
284
49
45
100 (83)
20
69 (24%)
3.0
MobileWorld
UI-TARS-7B
120
23
97
0
12
19 (13)
10
56 (58%)
3.9
MobileWorld
Qwen3-VL-8B
120
25
95
0
6
22 (16)
5
62 (65%)
3.6
Appendix
Table 6: What becomes of a judged failure. Per stream and agent, mean over the three Solo runs: episodes, judged successes, judged failures, and of the failures those with no prefix (ended at the first step), those on which the proposer abstained, those rejected by the guards (in parentheses, rejected as a single-action prefix), those rejected by the verifier, and those admitted as relabeled prefixes, with the admitted share of the failures and the mean number of steps in an admitted prefix.
Benchmark
Method
Agent
Run 1
Run 2
Run 3
Mean ± std
SR (%)
WebArena
Frozen agent
UI-TARS-7B
68 (26/19/23)
69 (25/21/23)
65 (19/20/26)
67.33 ± 2.08
20.8
Frozen agent
Qwen3-VL-8B
63 (15/23/25)
64 (23/19/22)
68 (23/23/22)
65.00 ± 2.65
20.1
AWM-online
UI-TARS-7B
72 (23/28/21)
78 (27/27/24)
71 (19/28/24)
73.67 ± 3.79
22.7
AWM-online
Qwen3-VL-8B
62 (19/21/22)
55 (19/17/19)
65 (20/23/22)
60.67 ± 5.13
18.7
DMS
UI-TARS-7B
72 (22/28/22)
71 (25/25/21)
75 (20/27/28)
72.67 ± 2.08
22.4
DMS
Qwen3-VL-8B
62 (21/20/21)
61 (22/18/21)
64 (23/18/23)
62.33 ± 1.53
19.2
Appendix
Table 7: Every run behind Tab. 2 . Wins per run with the round split in parentheses, then mean ± std of wins and success rate (%).
Stream
Variant
Run 1
Run 2
Run 3
Mean ± std
SR (%)
Updates
VisualWebArena
Frozen agent
64 (24/20/20)
59 (19/18/22)
60 (22/16/22)
61.00 ± 2.65
14.8
–
Solo (full)
82 (22/31/29)
95 (30/32/33)
81 (21/27/33)
86.00 ± 7.81
20.9
196
w/o window
74 (20/28/26)
80 (23/29/28)
70 (20/25/25)
74.67 ± 5.03
18.2
176
w/o top- K self-distillation
75 (19/28/28)
83 (21/30/32)
96 (26/37/33)
84.67 ± 10.60
20.6
180
w/o hindsight relabeling
64 (19/18/27)
71 (24/22/25)
75 (24/25/26)
70.00 ± 5.57
17.0
109
WebArena
Frozen agent
63 (15/23/25)
64 (23/19/22)
68 (23/23/22)
65.00 ± 2.65
20.1
–
Appendix
Table 8: Every run behind Tab. 3 . Qwen3-VL-8B on both web streams. Wins per run with the round split in parentheses, then mean ± std of wins, success rate (%), and the mean number of updates per run.
Stream
Auxiliaries
Wins per run
Judge precision per run
SR (%)
VisualWebArena
gpt-5-mini
82 / 95 / 81
0.52 / 0.54 / 0.48
20.9
VisualWebArena
gpt-5
76 / 73 / 84
0.59 / 0.59 / 0.60
18.9
VisualWebArena
gpt-5-nano
80 / 76 / 80
0.46 / 0.41 / 0.42
19.1
VisualWebArena
the agent’s own model (Qwen3-VL-8B)
69 / 75 / 69
0.36 / 0.36 / 0.34
17.3
WebArena
gpt-5-mini
80 / 78 / 84
0.63 / 0.67 / 0.67
24.9
WebArena
gpt-5
72 / 76 / 76
0.74 / 0.75 / 0.78
23.0
Appendix
Table 9: Every run behind Tab. 4 . Wins per run, the judge’s precision on that run (share of admitted episodes that the evaluator scores as successes), and the success rate over the runs.
Stream
Agent
Method
Steps per episode (r1 / r2 / r3)
Budget-ended % (r1 / r2 / r3)
WebArena
UI-TARS-7B
frozen agent
17.0 / 16.5 / 17.5
41 / 40 / 43
AWM-online
17.4 / 16.4 / 17.5
42 / 36 / 40
DMS
16.7 / 16.0 / 17.5
41 / 36 / 43
Solo
17.8 / 15.7 / 16.1
43 / 34 / 35
Qwen3-VL-8B
frozen agent
9.1 / 9.5 / 9.1
6 / 8 / 8
AWM-online
9.0 / 8.1 / 8.0
9 / 7 / 5
Appendix
Table 10: Episode length along the stream. Mean steps per episode and share of episodes ending on the step budget, per round, mean over the three runs. Every cell of Tab. 2 on the web streams and every variant of Tab. 3 .
Stream
Agent
Judged
True among judged
Evaluator
Precision
Recall
WebArena
UI-TARS-7B
96
58
84
0.60
0.69
WebArena
Qwen3-VL-8B
74
49
81
0.66
0.60
VisualWebArena
UI-TARS-7B
94
57
82
0.61
0.70
VisualWebArena
Qwen3-VL-8B
127
65
86
0.51
0.76
MobileWorld
UI-TARS-7B
23
10
11
0.44
0.94
MobileWorld
Qwen3-VL-8B
25
16
16
0.63
0.98
Appendix
Table 11: The gpt-5-mini judge on Solo ’s episodes. Mean over the three runs per cell: judged successes, the true successes among them, the evaluator’s successes, precision (true among judged) and recall (judged among true).
Stream
Judge
Wins per run
Judge precision
Success (%)
VisualWebArena
gpt-5-mini (default)
82 / 95 / 81
0.51
20.9 ± 1.9
VisualWebArena
benchmark evaluator
80 / 78 / 95
1.00
20.5 ± 2.3
WebArena
gpt-5-mini (default)
80 / 78 / 84
0.66
24.9 ± 0.9
WebArena
benchmark evaluator
78 / 71 / 78
1.00
23.4 ± 1.2
Appendix
Table 12: A perfect judge against the default judge. Solo on Qwen3-VL-8B with the judge replaced by the benchmark evaluator, proposer and verifier unchanged. Wins per run, judge precision, and success rate (%) mean ± std over three runs.
Judge
Proposer and verifier
Wins per run
Judge precision
Success (%)
gpt-5-mini
gpt-5-mini
82 / 95 / 81
0.51
20.9 ± 1.9
Qwen3-VL-8B
gpt-5-mini
71 / 70 / 83
0.36
18.2 ± 1.8
gpt-5-mini
Qwen3-VL-8B
76 / 69 / 70
0.48
17.4 ± 0.9
Qwen3-VL-8B
Qwen3-VL-8B
69 / 75 / 69
0.35
17.3 ± 0.8
Appendix
Table 13: The auxiliary roles swapped one at a time on VisualWebArena. Solo on Qwen3-VL-8B. Wins per run, judge precision and success rate (%) mean ± std over three runs.