GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.
Figures & tables
Figure 1: Performance overview across GUI benchmarks. Panel (a) reports ScreenSpot-Pro percentages on a fixed 0–50 scale, averaging the Text and Icon columns within each domain. Panel (b) reports the Type, GR, and SR percentages on a 0–100 scale. Panel (c) summarizes the arithmetic mean of the four ScreenSpot subcolumns in Table 1 , plus the Low-level Overall and GUI-Odyssey scores in Table 2 . GA, OW, and OD denote GUI-Act-Web, OmniAct-Web, and OmniAct-Desktop.
Figure 2: Spatial credit within grouped GUI training. At state st , the policy samples actions and observes their rewards. The upper panel shows group credit Atg and the spatial correction AtS=AtF−Atg . Below, Prox ranks all-miss clicks by proximity scores pi , while Residual predicts mixed-hit rewards from held-out fits, forms residual credit, and blends it with group credit according to predictive skill. All-hit groups have zero spatial correction.
ScreenSpot-Pro
ScreenSpot
Model
Dev
CAD
Creative
Scientific
Office
OS
Web
Desktop
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Supervised fine-tuning
SeeClick
0.6
0.0
2.5
0.0
1.0
0.0
3.5
0.0
1.1
0.0
2.8
0.0
55.7
32.5
72.2
30.0
OS-Atlas-4B
7.1
0.0
2.0
0.0
3.0
1.4
9.0
5.5
5.1
3.8
5.6
0.0
82.6
63.1
72.1
45.7
ShowUI-2B
16.9
1.4
2.5
0.0
9.1
0.0
13.2
7.3
15.3
7.5
10.3
2.2
–
–
–
–
Table 1: GUI grounding results on ScreenSpot-Pro and ScreenSpot. ScreenSpot-Pro contains six domains with text/icon subsets; ScreenSpot contains Web and Desktop text/icon subsets. SCA reports mean ± sample standard deviation over three independent training runs; other rows reproduce reported values. All RFT rows use a 3B-scale policy. Bold marks the highest RFT value.
Model
GUI-Act-Web
OmniAct-Web
OmniAct-Desktop
Low-lvl
GUI-Odyssey
Type
GR
SR
Type
GR
SR
Type
GR
SR
Overall
SR
Supervised fine-tuning
OS-Atlas-4B
79.22
58.57
42.62
46.74
49.24
22.99
63.30
42.55
26.94
50.71
64.58
OS-Atlas-7B
86.95
75.61
57.02
85.63
69.35
59.15
90.24
62.87
56.73
70.07
73.00
QwenVL2.5-3B
76.95
66.34
61.69
66.24
56.91
53.02
77.62
62.54
63.76
65.79
62.03
QwenVL2.5-7B
87.66
84.77
79.89
81.62
73.45
73.39
86.23
80.17
79.80
80.09
84.00
Table 2: Offline GUI action-prediction results. Action-type accuracy (Type), click-grounding accuracy (GR), and step success rate (SR). SCA reports mean ± sample standard deviation over three independent training runs; comparison rows reproduce reported values. Bold marks the highest RFT value.
Domain
GUI-R1
SCA
Δ
Dev
19.30
20.05
+0.75
CAD
17.10
17.85
+0.75
Creative
23.25
23.95
+0.70
Scientific
39.55
40.50
+0.95
Office
35.30
36.40
+1.10
OS
16.85
17.75
+0.90
Table 3: ScreenSpot-Pro domain means (%). Text and icon scores are equally weighted; Δ denotes the gain over GUI-R1.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Constraint summary
Required
Selection
Holdout
Maximum null active-rollout rate
≤0.1000
0.0801
0.0793
Maximum null rollout rate with w≥0.9
≤0.0200
0.0171
0.0155
Minimum clean paired activation lift
≥0.0100
0.0139
0.0159
Minimum clean mean gate weight
≥0.0100
0.0179
0.0189
Appendix
Table 5: Selection and one-shot holdout study for the fixed N=5 gate (10,000 groups per scenario; selection seed 20260807; holdout seed 20260808). “Null active” and “null strong” are simultaneous 95% upper bounds; “clean lift” and “clean mean weight” are simultaneous lower bounds.
Rule
Spatial reference
Target geometry
Binary GRPO
None; normalize binary group rewards
In the evaluator
Distance shaping
Fixed function of action-to-target distance
Read explicitly
Distance LOO
Reward trend fitted against target distance
Read explicitly
SCA-Residual
Reward trend fitted against sampled coordinates
In the evaluator
SCA-Prox
Exponential proximity, scaled after normalization
Read explicitly
Appendix
Table 6: Information used to construct credit. The task evaluator supplies scalar rewards; spatial rules use the additional quantities listed.
Click ai
Reward ri
GRPO
Prediction r^i
Gate wi
Residual
(0,0)
0
−0.730
−0.423
1.000
0.770
(1,2)
0
−0.730
0.309
0.190
−1.095
(2,1)
0
−0.730
0.309
0.190
−1.095
(4,5)
1
1.095
–
0.000
0.710
(5,4)
1
1.095
–
0.000
0.710
Appendix
Table 7: A deterministic five-click calculation. Values are rounded to three decimals. A dash denotes an invalid prediction, with zero gate weight.
Scenario
cx
cy
hx
hy
P(hit)
∇xJ
∇yJ
Dense offset
0.35
−0.20
1.20
1.20
0.56418
0.12015
−0.06842
Balanced offset
0.45
0.25
0.85
0.70
0.28076
0.09896
0.05948
Sparse offset
0.90
0.55
0.55
0.45
0.08733
0.07110
0.04489
Rare/far
1.45
−0.80
0.45
0.35
0.02615
0.03550
−0.02009
Appendix
Table 8: Exact-gradient contextual-bandit scenarios. c is target center, h is target half-extent, and P(hit) is analytic at μ=(0,0) .
Estimator
Mean gx± MCSE
Mean gy± MCSE
Exact reference
0.08143
0.00396
Binary GRPO
0.13220±0.00045
0.00770±0.00044
2-D OLS-LOO
−0.00307±0.00034
−0.00489±0.00034
2-D Ridge-LOO
−0.00304±0.00034
−0.00490±0.00034
SCA-Residual
0.12396±0.00043
0.00818±0.00042
Target-distance LOO
−0.01560±0.00045
−0.00416±0.00045
Appendix
Table 9: Estimated mean gradient and component-wise Monte Carlo standard error on the uniform contextual mixture. The exact row has no sampling error.
Context
Corruption
Estimator
Drift
103 MSE Δ
Balanced
reward flip
median ensemble
0.146
−0.36
Balanced
reward flip
OLS ensemble
0.145
−1.32
Balanced
coordinate spike
median ensemble
0.042
9.24
Balanced
coordinate spike
OLS ensemble
0.038
8.21
Sparse
reward flip
median ensemble
0.021
−2.42
Sparse
reward flip
OLS ensemble
0.022
−2.61
Appendix
Table 10: Controlled corruption study; smaller is better for both metrics.
Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat all training samples equally, although their training value to the baseline model varies with difficulty. Overlooking this can greatly reduce training efficiency or even cause collapse. At the strategy level, existing frameworks struggle to balance the trade-off between cropping larger regions for sufficient context and smaller ones for reduced redundancy, a tension inherent to tool-augmented grounding agents. In addition, overly complex decision-making is difficult for small-parameter models and significantly increases inference time. To address these issues, at the data level, we propose GUI-D, a data mining and difficulty scoring pipeline that identifies the training-worthy samples by proper testing and assigns difficulty scores to guide subsequent training weights. At the strategy level, we propose GUI-C2, which employs an area-gated coarse-to-fine refinement mechanism that progressively narrows the visual field via model-internal uncertainty signals, adaptively reserving context for large targets while amplifying precision for small ones, reinforced by improvement-aware stage rewards that ensure each refinement genuinely advances grounding. Meanwhile, we simplify the decision-making process to greatly reduce additional inference time. Finally, extensive experiments show that our method achieves state-of-the-art performance. The code and data will be publicly available.
Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tuning and reinforcement learning methods are constrained by the high cost of annotation, creating a scalability bottleneck. In this paper, we introduce a label-free test-time training paradigm driven by two key insights: (1) confidence patterns in coordinate tokens are a better indicator than full-sequence confidence, and (2) in sparse GUI coordinate spaces, negative samples offer more reliable learning signals than potentially noisy positive ones. We first propose Confidence-Anchored Learning (CAL), which utilizes coordinate-token confidence to filter pseudo-labels and assign distance-based binary rewards. Building on this, we develop Confidence-Anchored Negative Learning (CANL), which exclusively optimizes the model using negative samples to bypass the risks of incorrect positive samples. Experimental results demonstrate that CANL-7B achieves 92.1% on ScreenSpot-V2. On more challenging ScreenSpot-Pro, CANL-7B reaches 33.8%, an 8.9% absolute improvement over the base model. Our findings establish coordinate-token confidence as a powerful alternative to manual annotations for scalable GUI agent development.
Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.