One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
Organizations: Concordia University · Mila – Qu´ebec AI Institute · University of Toronto · Shanghai University
Abstract
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.
Figures & tables
| Approach | GT supervision | Extra rollouts | Adaptation stage | Retained state | Related work |
|---|---|---|---|---|---|
| Demonstration SFT | Yes | No | Pre-deployment | Weights | Wu et al. (2025) ; Qin et al. (2025) |
| RL / self-evolution | Varies | Yes | Pre-deployment | Weights | Bai et al. (2024) ; Qi et al. (2025) ; Xiao et al. (2025) |
| Exploration memory | Varies | Yes | Pre-deployment | Memory | Zhang et al. (2025) ; Sun et al. (2026) |
| Retry-based improvement | Varies | Yes | Deployment | Memory | Shinn et al. (2023) ; Li et al. (2026a) ; He et al. (2026) |
| Multi-rollout inference | No | Yes | Deployment | None | Snell et al. (2024) ; Yang et al. (2026) |
| Multi-rollout adaptation | No | Yes | Deployment | Weights | Akyürek et al. (2025) ; Zuo et al. (2025) |
| WebArena | VisualWebArena | MobileWorld | ||||
|---|---|---|---|---|---|---|
| Method | UI-TARS-7B | Qwen3-VL-8B | UI-TARS-7B | Qwen3-VL-8B | UI-TARS-7B | Qwen3-VL-8B |
| Frozen agent | 20.8 0.6 | 20.1 0.8 | 16.7 1.4 | 14.8 0.6 | 6.1 1.0 | 9.7 1.0 |
| AWM-online | 22.7 1.2 | 18.7 1.6 | 16.3 1.3 | 14.7 1.2 | 9.7 1.0 | 10.8 1.7 |
| DMS | 22.4 0.6 | 19.2 0.5 | 17.0 0.9 | 15.7 0.9 | 7.5 0.8 | 9.7 0.5 |
| Solo (ours) | 25.8 1.6 | 24.9 0.9 | 20.0 0.8 | 20.9 1.9 | 9.2 0.8 | 13.3 2.2 |
| Variant | VisualWebArena | WebArena | Mean |
|---|---|---|---|
| Frozen agent | 14.8 0.6 | 20.1 0.8 | 17.5 |
| Solo (full) | 20.9 1.9 | 24.9 0.9 | 22.9 |
| w/o window (one step per episode) | 18.2 1.2 | 23.0 1.2 | 20.6 |
| w/o top- self-distillation (one-hot targets) | 20.6 2.6 | 22.0 1.5 | 21.3 |
| w/o hindsight relabeling (success branch only) | 17.0 1.4 | 21.9 0.6 | 19.4 |
| VisualWebArena | WebArena | |||
| Auxiliaries | Success (%) | Judge prec. | Success (%) | Judge prec. |
| Frozen agent (no auxiliaries) | 14.8 0.6 | – | 20.1 0.8 | – |
| gpt-5-mini | 20.9 1.9 | 0.51 | 24.9 0.9 | 0.66 |
| gpt-5 | 18.9 1.4 | 0.59 | 23.0 0.7 | 0.76 |
| gpt-5-nano | 19.1 0.6 | 0.43 | 23.9 1.7 | 0.46 |
| the agent’s own model (Qwen3-VL-8B) | 17.3 0.8 | 0.35 | 24.2 1.0 | 0.54 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Stream | Arm | Wins per round (total) | One-step episodes | Entropy |
|---|---|---|---|---|
| VisualWebArena | frozen (mean of 3 runs) | 22 / 18 / 21 (61) | 36 / 35 / 33 | – |
| entropy min., run 1 | 18 / 4 / 1 (23) | 55 / 90 / 117 | 0.116 0.014 | |
| entropy min., run 2 | 22 / 10 / 4 (36) | 48 / 53 / 83 | 0.113 0.012 | |
| entropy min., run 3 | 20 / 6 / 8 (34) | 48 / 87 / 76 | 0.114 0.014 | |
| WebArena | frozen (mean of 3 runs) | 20 / 22 / 23 (65) | 5 / 6 / 6 | – |
| entropy min., run 1 | 21 / 24 / 21 (66) | 6 / 2 / 2 | 0.111 0.016 |
| Stream | Agent | Episodes | Succ. | Fail. | No prefix | Abstain | Guards | Verifier | Admitted | Prefix steps |
|---|---|---|---|---|---|---|---|---|---|---|
| WebArena | UI-TARS-7B | 324 | 96 | 228 | 1 | 42 | 63 (53) | 22 | 101 (44%) | 3.7 |
| WebArena | Qwen3-VL-8B | 324 | 74 | 250 | 6 | 23 | 69 (51) | 19 | 134 (54%) | 3.0 |
| VisualWebArena | UI-TARS-7B | 411 | 94 | 317 | 2 | 67 | 133 (100) | 24 | 91 (29%) | 3.3 |
| VisualWebArena | Qwen3-VL-8B | 411 | 127 | 284 | 49 | 45 | 100 (83) | 20 | 69 (24%) | 3.0 |
| MobileWorld | UI-TARS-7B | 120 | 23 | 97 | 0 | 12 | 19 (13) | 10 | 56 (58%) | 3.9 |
| MobileWorld | Qwen3-VL-8B | 120 | 25 | 95 | 0 | 6 | 22 (16) | 5 | 62 (65%) | 3.6 |
| Benchmark | Method | Agent | Run 1 | Run 2 | Run 3 | Mean std | SR (%) |
|---|---|---|---|---|---|---|---|
| WebArena | Frozen agent | UI-TARS-7B | 68 (26/19/23) | 69 (25/21/23) | 65 (19/20/26) | 67.33 2.08 | 20.8 |
| Frozen agent | Qwen3-VL-8B | 63 (15/23/25) | 64 (23/19/22) | 68 (23/23/22) | 65.00 2.65 | 20.1 | |
| AWM-online | UI-TARS-7B | 72 (23/28/21) | 78 (27/27/24) | 71 (19/28/24) | 73.67 3.79 | 22.7 | |
| AWM-online | Qwen3-VL-8B | 62 (19/21/22) | 55 (19/17/19) | 65 (20/23/22) | 60.67 5.13 | 18.7 | |
| DMS | UI-TARS-7B | 72 (22/28/22) | 71 (25/25/21) | 75 (20/27/28) | 72.67 2.08 | 22.4 | |
| DMS | Qwen3-VL-8B | 62 (21/20/21) | 61 (22/18/21) | 64 (23/18/23) | 62.33 1.53 | 19.2 |
| Stream | Variant | Run 1 | Run 2 | Run 3 | Mean std | SR (%) | Updates |
|---|---|---|---|---|---|---|---|
| VisualWebArena | Frozen agent | 64 (24/20/20) | 59 (19/18/22) | 60 (22/16/22) | 61.00 2.65 | 14.8 | – |
| Solo (full) | 82 (22/31/29) | 95 (30/32/33) | 81 (21/27/33) | 86.00 7.81 | 20.9 | 196 | |
| w/o window | 74 (20/28/26) | 80 (23/29/28) | 70 (20/25/25) | 74.67 5.03 | 18.2 | 176 | |
| w/o top- self-distillation | 75 (19/28/28) | 83 (21/30/32) | 96 (26/37/33) | 84.67 10.60 | 20.6 | 180 | |
| w/o hindsight relabeling | 64 (19/18/27) | 71 (24/22/25) | 75 (24/25/26) | 70.00 5.57 | 17.0 | 109 | |
| WebArena | Frozen agent | 63 (15/23/25) | 64 (23/19/22) | 68 (23/23/22) | 65.00 2.65 | 20.1 | – |
| Stream | Auxiliaries | Wins per run | Judge precision per run | SR (%) |
|---|---|---|---|---|
| VisualWebArena | gpt-5-mini | 82 / 95 / 81 | 0.52 / 0.54 / 0.48 | 20.9 |
| VisualWebArena | gpt-5 | 76 / 73 / 84 | 0.59 / 0.59 / 0.60 | 18.9 |
| VisualWebArena | gpt-5-nano | 80 / 76 / 80 | 0.46 / 0.41 / 0.42 | 19.1 |
| VisualWebArena | the agent’s own model (Qwen3-VL-8B) | 69 / 75 / 69 | 0.36 / 0.36 / 0.34 | 17.3 |
| WebArena | gpt-5-mini | 80 / 78 / 84 | 0.63 / 0.67 / 0.67 | 24.9 |
| WebArena | gpt-5 | 72 / 76 / 76 | 0.74 / 0.75 / 0.78 | 23.0 |
| Stream | Agent | Method | Steps per episode (r1 / r2 / r3) | Budget-ended % (r1 / r2 / r3) |
| WebArena | UI-TARS-7B | frozen agent | 17.0 / 16.5 / 17.5 | 41 / 40 / 43 |
| AWM-online | 17.4 / 16.4 / 17.5 | 42 / 36 / 40 | ||
| DMS | 16.7 / 16.0 / 17.5 | 41 / 36 / 43 | ||
| Solo | 17.8 / 15.7 / 16.1 | 43 / 34 / 35 | ||
| Qwen3-VL-8B | frozen agent | 9.1 / 9.5 / 9.1 | 6 / 8 / 8 | |
| AWM-online | 9.0 / 8.1 / 8.0 | 9 / 7 / 5 |
| Stream | Agent | Judged | True among judged | Evaluator | Precision | Recall |
|---|---|---|---|---|---|---|
| WebArena | UI-TARS-7B | 96 | 58 | 84 | 0.60 | 0.69 |
| WebArena | Qwen3-VL-8B | 74 | 49 | 81 | 0.66 | 0.60 |
| VisualWebArena | UI-TARS-7B | 94 | 57 | 82 | 0.61 | 0.70 |
| VisualWebArena | Qwen3-VL-8B | 127 | 65 | 86 | 0.51 | 0.76 |
| MobileWorld | UI-TARS-7B | 23 | 10 | 11 | 0.44 | 0.94 |
| MobileWorld | Qwen3-VL-8B | 25 | 16 | 16 | 0.63 | 0.98 |
| Stream | Judge | Wins per run | Judge precision | Success (%) |
|---|---|---|---|---|
| VisualWebArena | gpt-5-mini (default) | 82 / 95 / 81 | 0.51 | 20.9 1.9 |
| VisualWebArena | benchmark evaluator | 80 / 78 / 95 | 1.00 | 20.5 2.3 |
| WebArena | gpt-5-mini (default) | 80 / 78 / 84 | 0.66 | 24.9 0.9 |
| WebArena | benchmark evaluator | 78 / 71 / 78 | 1.00 | 23.4 1.2 |
| Judge | Proposer and verifier | Wins per run | Judge precision | Success (%) |
|---|---|---|---|---|
| gpt-5-mini | gpt-5-mini | 82 / 95 / 81 | 0.51 | 20.9 1.9 |
| Qwen3-VL-8B | gpt-5-mini | 71 / 70 / 83 | 0.36 | 18.2 1.8 |
| gpt-5-mini | Qwen3-VL-8B | 76 / 69 / 70 | 0.48 | 17.4 0.9 |
| Qwen3-VL-8B | Qwen3-VL-8B | 69 / 75 / 69 | 0.35 | 17.3 0.8 |