GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.
Figures & tables
Figure 1: Performance overview across GUI benchmarks. Panel (a) reports ScreenSpot-Pro percentages on a fixed 0–50 scale, averaging the Text and Icon columns within each domain. Panel (b) reports the Type, GR, and SR percentages on a 0–100 scale. Panel (c) summarizes the arithmetic mean of the four ScreenSpot subcolumns in Table 1 , plus the Low-level Overall and GUI-Odyssey scores in Table 2 . GA, OW, and OD denote GUI-Act-Web, OmniAct-Web, and OmniAct-Desktop.
Figure 2: Spatial credit within grouped GUI training. At state st , the policy samples actions and observes their rewards. The upper panel shows group credit Atg and the spatial correction AtS=AtF−Atg . Below, Prox ranks all-miss clicks by proximity scores pi , while Residual predicts mixed-hit rewards from held-out fits, forms residual credit, and blends it with group credit according to predictive skill. All-hit groups have zero spatial correction.
ScreenSpot-Pro
ScreenSpot
Model
Dev
CAD
Creative
Scientific
Office
OS
Web
Desktop
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Supervised fine-tuning
SeeClick
0.6
0.0
2.5
0.0
1.0
0.0
3.5
0.0
1.1
0.0
2.8
0.0
55.7
32.5
72.2
30.0
OS-Atlas-4B
7.1
0.0
2.0
0.0
3.0
1.4
9.0
5.5
5.1
3.8
5.6
0.0
82.6
63.1
72.1
45.7
ShowUI-2B
16.9
1.4
2.5
0.0
9.1
0.0
13.2
7.3
15.3
7.5
10.3
2.2
–
–
–
–
Table 1: GUI grounding results on ScreenSpot-Pro and ScreenSpot. ScreenSpot-Pro contains six domains with text/icon subsets; ScreenSpot contains Web and Desktop text/icon subsets. SCA reports mean ± sample standard deviation over three independent training runs; other rows reproduce reported values. All RFT rows use a 3B-scale policy. Bold marks the highest RFT value.
Model
GUI-Act-Web
OmniAct-Web
OmniAct-Desktop
Low-lvl
GUI-Odyssey
Type
GR
SR
Type
GR
SR
Type
GR
SR
Overall
SR
Supervised fine-tuning
OS-Atlas-4B
79.22
58.57
42.62
46.74
49.24
22.99
63.30
42.55
26.94
50.71
64.58
OS-Atlas-7B
86.95
75.61
57.02
85.63
69.35
59.15
90.24
62.87
56.73
70.07
73.00
QwenVL2.5-3B
76.95
66.34
61.69
66.24
56.91
53.02
77.62
62.54
63.76
65.79
62.03
QwenVL2.5-7B
87.66
84.77
79.89
81.62
73.45
73.39
86.23
80.17
79.80
80.09
84.00
Table 2: Offline GUI action-prediction results. Action-type accuracy (Type), click-grounding accuracy (GR), and step success rate (SR). SCA reports mean ± sample standard deviation over three independent training runs; comparison rows reproduce reported values. Bold marks the highest RFT value.
Domain
GUI-R1
SCA
Δ
Dev
19.30
20.05
+0.75
CAD
17.10
17.85
+0.75
Creative
23.25
23.95
+0.70
Scientific
39.55
40.50
+0.95
Office
35.30
36.40
+1.10
OS
16.85
17.75
+0.90
Table 3: ScreenSpot-Pro domain means (%). Text and icon scores are equally weighted; Δ denotes the gain over GUI-R1.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Constraint summary
Required
Selection
Holdout
Maximum null active-rollout rate
≤0.1000
0.0801
0.0793
Maximum null rollout rate with w≥0.9
≤0.0200
0.0171
0.0155
Minimum clean paired activation lift
≥0.0100
0.0139
0.0159
Minimum clean mean gate weight
≥0.0100
0.0179
0.0189
Appendix
Table 5: Selection and one-shot holdout study for the fixed N=5 gate (10,000 groups per scenario; selection seed 20260807; holdout seed 20260808). “Null active” and “null strong” are simultaneous 95% upper bounds; “clean lift” and “clean mean weight” are simultaneous lower bounds.
Rule
Spatial reference
Target geometry
Binary GRPO
None; normalize binary group rewards
In the evaluator
Distance shaping
Fixed function of action-to-target distance
Read explicitly
Distance LOO
Reward trend fitted against target distance
Read explicitly
SCA-Residual
Reward trend fitted against sampled coordinates
In the evaluator
SCA-Prox
Exponential proximity, scaled after normalization
Read explicitly
Appendix
Table 6: Information used to construct credit. The task evaluator supplies scalar rewards; spatial rules use the additional quantities listed.
Click ai
Reward ri
GRPO
Prediction r^i
Gate wi
Residual
(0,0)
0
−0.730
−0.423
1.000
0.770
(1,2)
0
−0.730
0.309
0.190
−1.095
(2,1)
0
−0.730
0.309
0.190
−1.095
(4,5)
1
1.095
–
0.000
0.710
(5,4)
1
1.095
–
0.000
0.710
Appendix
Table 7: A deterministic five-click calculation. Values are rounded to three decimals. A dash denotes an invalid prediction, with zero gate weight.
Scenario
cx
cy
hx
hy
P(hit)
∇xJ
∇yJ
Dense offset
0.35
−0.20
1.20
1.20
0.56418
0.12015
−0.06842
Balanced offset
0.45
0.25
0.85
0.70
0.28076
0.09896
0.05948
Sparse offset
0.90
0.55
0.55
0.45
0.08733
0.07110
0.04489
Rare/far
1.45
−0.80
0.45
0.35
0.02615
0.03550
−0.02009
Appendix
Table 8: Exact-gradient contextual-bandit scenarios. c is target center, h is target half-extent, and P(hit) is analytic at μ=(0,0) .
Estimator
Mean gx± MCSE
Mean gy± MCSE
Exact reference
0.08143
0.00396
Binary GRPO
0.13220±0.00045
0.00770±0.00044
2-D OLS-LOO
−0.00307±0.00034
−0.00489±0.00034
2-D Ridge-LOO
−0.00304±0.00034
−0.00490±0.00034
SCA-Residual
0.12396±0.00043
0.00818±0.00042
Target-distance LOO
−0.01560±0.00045
−0.00416±0.00045
Appendix
Table 9: Estimated mean gradient and component-wise Monte Carlo standard error on the uniform contextual mixture. The exact row has no sampling error.
Context
Corruption
Estimator
Drift
103 MSE Δ
Balanced
reward flip
median ensemble
0.146
−0.36
Balanced
reward flip
OLS ensemble
0.145
−1.32
Balanced
coordinate spike
median ensemble
0.042
9.24
Balanced
coordinate spike
OLS ensemble
0.038
8.21
Sparse
reward flip
median ensemble
0.021
−2.42
Sparse
reward flip
OLS ensemble
0.022
−2.61
Appendix
Table 10: Controlled corruption study; smaller is better for both metrics.