Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Organizations: Department of Electrical Engineering and Computer Science, South Dakota State University
Abstract
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.
Figures & tables
| MobileGym, gain over the base agent | ||||
| Arm | GUI-Owl-8B | Qwen3-VL-8B | ScaleCUA-7B | |
| Procedures to the context | +5.3 | +5.2 | +5.1 | – |
| state condition removed | +1.4 | +1.8 | +0.8 | 3.9 |
| Procedures to the weights | +0.3 | +2.0 | +0.6 | – |
| items copied | +3.3 | +6.1 | +3.9 | 3.5 |
| State facts to the context | +4.6 | +5.4 | +4.9 | – |
| MobileGym | AndroidWorld | ||||||
| Arm | GUI-Owl | Qwen3-VL | ScaleCUA | GUI-Owl | Qwen3-VL | ScaleCUA | |
| Base agent | 21.6 | 16.7 | 19.2 | 23.6 | 17.8 | 24.0 | 14.1 |
| Self-retry, three attempts | 26.4 | 20.6 | 24.6 | 26.1 | 22.7 | 27.2 | 9.3 |
| Success-filtered fine-tuning | 22.2 | 17.3 | 20.0 | 24.1 | 19.3 | 23.3 | 13.4 |
| Retrieve one trajectory | 21.0 | 15.1 | 18.1 | 22.8 | 18.1 | 22.1 | 15.2 |
| Retrieve ten step records | 22.9 | 17.9 | 20.8 | 24.3 | 19.5 | 25.3 | 12.6 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| MobileGym | AndroidWorld | |||||||
| Component | Items | Valued | Items | Valued | ||||
| Locators | 9907 | 0.39 | 0.10 | 0.21 | 5544 | 0.38 | 0.09 | 0.20 |
| Procedures | 51052 | 0.05 | 0.66 | 0.06 | 25519 | 0.05 | 0.63 | 0.05 |
| State facts | 3980 | 0.32 | 0.56 | 0.53 | 1972 | 0.31 | 0.56 | 0.56 |
| Lessons | 24651 | 0.45 | 0.53 | 0.15 | 12775 | 0.44 | 0.50 | 0.15 |
| MobileGym | Training samples | Dev. accuracy | ||||
| Arm | GUI-Owl | Qwen3-VL | ScaleCUA | GUI-Owl | Qwen3-VL | ScaleCUA |
| Locators to the weights | 1211 | 1035 | 1056 | 0.952 | 0.949 | 0.959 |
| Procedures to the weights | 5721 | 5498 | 5798 | 0.932 | 0.946 | 0.946 |
| State facts to the weights | 476 | 450 | 401 | 0.954 | 0.954 | 0.961 |
| Lessons to the weights | 2854 | 2600 | 2763 | 0.931 | 0.934 | 0.940 |
| All to the weights | 10263 | 9583 | 10018 | 0.942 | 0.930 | 0.941 |
| Routing coefficients | Agreement | ||||
| Fit | Cells | Sign | |||
| All 24 cells | 24 | 11.29 | 5.73 | 1.28 | 24/24 |
| Held out: GUI-Owl-8B | 16 | 9.32 | 4.96 | 0.89 | 8/8 |
| Held out: Qwen3-VL-8B | 16 | 9.01 | 5.28 | 0.70 | 8/8 |
| Held out: ScaleCUA-7B | 16 | 9.16 | 4.86 | 0.94 | 8/8 |
| MobileGym | Pool coordinates | Route gain | Held-out rule | ||||
| Component | Family | [95%] | Pred. / obs. | ||||
| Locators | GUI-Owl-8B | 0.397 | 0.121 | +5.9 [+4.9, +7.0] | 1 | +2.21 | W / W |
| Qwen3-VL-8B | 0.388 | 0.064 | +5.5 [+4.5, +6.6] | 1 | +2.47 | W / W | |
| ScaleCUA-7B | 0.374 | 0.106 | +3.2 [+2.2, +4.3] | 1 | +1.97 | W / W | |
| Procedures | GUI-Owl-8B | 0.052 | 0.673 | 5.0 [ 6.0, 4.0] | 1 | -3.74 | C / C |
| Qwen3-VL-8B | 0.051 | 0.645 | 3.2 [ 4.2, 2.2] | 1 | -3.64 | C / C | |
| Assignment of the held-out family | MobileGym | AndroidWorld | |||||
| Fit score | GUI-Owl | Qwen3-VL | ScaleCUA | Single | Reverse | Single | Reverse |
| Half 1 half 2 | W CC W | W CCC / W CC W | W CC W | +3.7 | +8.5 | +2.4 | +6.8 |
| Half 2 half 1 | W CC W | W CC W | W CC W | +5.1 | +9.3 | +2.6 | +6.6 |
| MobileGym | Half of the pool | Route gain | Held-out rule | ||||
| Half | Family | Items | [95%] | Pred. / obs. | |||
| Lessons, high | GUI-Owl-8B | 4090 | 0.796 | 0.557 | +2.4 [+1.4, +3.5] | +3.77 | W / W |
| Qwen3-VL-8B | 3525 | 0.782 | 0.535 | +2.2 [+1.2, +3.2] | +3.52 | W / W | |
| ScaleCUA-7B | 4092 | 0.806 | 0.506 | +2.8 [+1.7, +3.8] | +3.99 | W / W | |
| Lessons, low | GUI-Owl-8B | 4472 | 0.134 | 0.549 | 1.6 [ 2.6, 0.7] | -2.36 | C / C |
| Qwen3-VL-8B | 4276 | 0.138 | 0.540 | 2.1 [ 3.0, 1.1] | -2.30 | C / C | |
| MobileGym, gain over the base agent | ||||
| Arm | GUI-Owl-8B | Qwen3-VL-8B | ScaleCUA-7B | |
| Procedures to the context | +5.3 | +5.2 | +5.1 | – |
| state condition removed | +1.4 | +1.8 | +0.8 | 3.9 |
| signature read from the screen elements | +4.0 | +5.5 | +3.9 | 0.7 |
| Procedures to the weights | +0.3 | +2.0 | +0.6 | – |
| items copied | +1.2 | – | – | 0.9 |
| AndroidWorld: routing gain (points) | |||
| Family | vs. all to the context | vs. all to the weights | vs. reverse assignment |
| GUI-Owl-8B | 0.3 [ 2.6, +2.0] | +4.3 [+1.7, +6.8] | +5.6 [+2.9, +8.2] |
| Qwen3-VL-8B | +5.2 [+2.7, +7.6] | +5.0 [+2.4, +7.6] | +9.8 [+6.8, +12.6] |
| ScaleCUA-7B | +3.6 [+1.3, +5.9] | +3.2 [+0.7, +5.6] | +4.7 [+2.3, +7.3] |
| Producer experience | Own experience | ||||
| Consumer | Producer | Notes | Weights | Mixed | Practice |
| Qwen3-VL-4B | Qwen3-VL-4B | +1.5 | 0.3 | – | 0.3 |
| Qwen3-VL-8B | +1.9 | 2.7 | 0.5 | – | |
| Qwen3-VL-32B | +3.1 | 1.9 | +0.0 | – | |
| Qwen3-VL-8B | Qwen3-VL-4B | +2.2 | 2.3 | 0.8 | – |
| Qwen3-VL-8B | +3.6 | +0.9 | – | +0.6 | |
| Locators | Procedures | State facts | Lessons | |||||
| Consumer | Before | After | Before | After | Before | After | Before | After |
| GUI-Owl-8B | 45 | 17 | 33 | 23 | 43 | 17 | 50 | 15 |
| Qwen3-VL-8B | 45 | 17 | 32 | 22 | 41 | 19 | 46 | 14 |
| ScaleCUA-7B | 46 | 16 | 35 | 22 | 42 | 19 | 50 | 15 |
| Gain over base (points) | Difference | ||||
| Setting | Variant | Arm | Default | Variant | Shift |
| Jaccard thresholds | exact 0.8 | C_F | +5.2 | +4.7 | 0.5 |
| C_G | +1.9 | +1.5 | 0.4 | ||
| C_L | +4.0 | +3.4 | 0.5 | ||
| C_P | +5.6 | +5.3 | 0.3 | ||
| relaxed 0.6 | C_F | +5.2 | +4.9 | 0.3 | |
| Adapter training | Inference | |||
| Arm | Samples | GPU-h | Tokens (k) | Ratio |
| Base agent | – | – | 13.2 | 1.00 |
| Locators to the context | – | – | 14.2 | 1.07 |
| Procedures to the context | – | – | 14.2 | 1.07 |
| State facts to the context | – | – | 14.2 | 1.07 |
| Lessons to the context | – | – | 14.2 | 1.07 |