Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4% of failures occur in 20% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8% and retains a 1.88% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR
Figures & tables
Figure 1 : Limitations of step-level evaluation and the need for structured failure diagnosis. Left: step-level success rate aggregates performance across all steps, obscuring where failures occur. Right: GUITAR maps execution outcomes onto the STG, enabling state-level and transition-level analysis to localize bottleneck states and failure transitions for targeted diagnosis and improvement.
Method
Cross-Task Aggregation
Abstraction Level
Primary Purpose
Failure Localization
Step-level SR [ 5 ]
✗
N/A
Performance Measurement
✗
PageAgent [ 18 ]
✓
Perceptual
Navigation Planning
✗
WebGraphEval [ 19 ]
✗
Behavioral
Cross-Agent Comparison
✗
GUITAR (Ours)
✓
Semantic-Functional
Failure Diagnosis
✓
Table 1 : Comparison of GUITAR with related graph-based GUI approaches. Cross-Task Aggregation : cross-run trajectory aggregation into a unified graph. Abstraction Level : semantic depth of state unification, from perceptual similarity to functional equivalence. Failure Localization : whether structured failure diagnoses are produced.
Figure 2 : Analysis of performance and SR contribution across state frequency intervals. We observe that long-tail states contribute marginally to the overall SR, regardless of their individual performance.
Figure 4
Figure 4 : Examples of screen-to-functional-role mappings. Left: visually different screens mapped to the same functional role (item purchasing). Right: visually similar screens mapped to different functional roles, e.g., Directions Setup Page vs. Route Map Page.
Figure 5 : Cumulative distribution of failures across states ranked by failure frequency. A small fraction of states accounts for a disproportionately large portion of total errors, indicating strong failure concentration.
Figure 6 : Analysis of bottleneck states in Maps : (a) state-level accuracy of the five most challenging states across all models, and (b) error transition rates originating from bottleneck states ( AGUVIS-7B ). GUITAR clearly reveals concentrated failure patterns at both the state and transition levels.
Model
Maps
Kitchen
CNN
Booking
Carmax
Kayak
Avg.
Uni.
Tgt.
Uni.
Tgt.
Uni.
Tgt.
Uni.
Tgt.
Uni.
Tgt.
Uni.
Tgt.
Uni.
Tgt.
Qwen2.5-VL-7B
-3.5
+1.0
+8.3
+4.6
-7.6
+2.6
-12.3
+1.4
-13.1
-0.7
-17.3
+1.1
-7.6
+1.7
Qwen2.5-VL-72B
0.0
+0.6
+3.0
+1.8
-17.9
0.0
-9.6
+1.4
-9.2
+3.9
+3.9
+3.9
-5.0
+2.0
OS-Atlas-7B
0.0
+3.5
-1.4
+2.3
+1.3
+2.5
+0.4
+0.3
+3.0
+8.0
+9.5
+9.5
+2.1
+4.4
GUI-Owl-7B
+5.5
+6.0
+7.5
+1.5
-7.6
+1.3
+1.0
+3.1
-0.9
-8.7
-2.7
-6.0
+0.5
-0.5
GUI-Owl-32B
+8.5
+8.5
+11.2
+12.5
-6.4
+2.5
+0.5
+5.3
-4.8
+6.3
-11.5
-12.0
-0.4
+3.9
Table 3 : SR Improvement ( Δ SR%) of Targeted ( Tgt. ) vs. Uniform ( Uni. ) Guidance over Original Baseline. Tgt. denotes GUITAR-guided inference targeting bottleneck states (Threshold = 0.6); Uni. denotes uniform guidance applied to all states (Threshold = 0). Best results per model are bolded .
Model
Maps
Kitchen
CNN
Booking
Carmax
Kayak
Avg.
Orig.
Tgt.
Orig.
Tgt.
Orig.
Tgt.
Orig.
Tgt.
Orig.
Tgt.
Orig.
Tgt.
Orig.
Tgt.
Qwen2.5-VL-7B
23.6
38.3
10.5
52.4
0.0
33.3
10.0
40.0
25.0
8.3
14.3
14.3
13.9
31.1
Qwen2.5-VL-72B
27.8
39.3
22.0
47.6
0.0
0.0
23.5
47.1
16.0
40.0
14.3
14.3
17.3
31.4
OS-Atlas-7B
27.9
31.2
22.9
36.4
27.8
44.4
34.2
35.5
16.2
27.6
28.4
49.1
26.2
37.4
GUI-Owl-7B
24.1
54.0
16.7
29.3
30.4
36.4
30.7
35.6
34.1
41.7
0.0
0.0
22.7
32.8
GUI-Owl-32B
25.4
62.8
17.4
60.7
26.7
34.9
30.2
36.1
34.4
47.8
16.7
16.7
25.1
43.2
Table 4 : Bottleneck State Success Rate Change ΔC(s) (%) before and after targeted guidance across six tasks. Values represent GUITAR-guided − Original (%).
Panel A: Fully automatic vs. manually verified STGs
Dataset
Top-20% Jaccard
Spearman ρ
∣ΔGini∣
∣ΔCoverage∣ (pp)
AndroidControl
0.932±0.214
0.989±0.029
0.017±0.024
1.89±2.70
Mind2Web †
0.500±0.136
0.942±0.035
0.043±0.025
6.26±3.00
Table 5: Robustness to manual verification and trajectory reuse on the six primary tasks. Panel A compares fully automatic and manually verified STGs; Panel B reports three-fold held-out guidance using fully automatic STGs.
Metric
Score
Gwet’s AC1
0.848
Raw Agreement
83.1%
Table 6: Inter-annotator agreement for state abstraction evaluation (5 annotators, 71 states).
App
Method
States
State Image Similarity
Action Similarity
Mean ↑
Std ↓
Ratio ↑
Intra ↑
Inter ↓
Gap ↑
p -value
Kitchen Stories
ImageHash
12
75.4
14.8
57.1
54.0
37.4
16.6
4.7e-10
PageAgent
45
82.2
9.7
83.3
78.6
29.0
27.3
6.1e-06
GUITAR
7
80.9
9.7
85.7
75.8
41.0
34.8
5.2e-09
Maps
ImageHash
12
68.2
12.1
37.9
42.1
42.7
− 0.7
0.6679
PageAgent
63
71.5
13.4
48.0
37.0
39.9
− 2.9
0.8907
Table 7: Comparison of STGs constructed by different methods, including PageAgent [ 18 ] . GUITAR achieves a more compact state abstraction while preserving both visual similarity and behavioral consistency,. Ratio denotes the proportion of image pairs with similarity above 70%. Intra and Inter represent the textual similarity within and across transitions, respectively. ↑ indicates higher is better, and ↓ indicates lower is better. Bold denotes the best result.
Method
VLM Calls
Avg. States ( N )
PageAgent
1,167
≈ 49
GUITAR
933
≈ 9
Reduction
20.1%
82.6%
Table 8: VLM call efficiency and graph compactness comparison on AndroidControl.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : The fully automatic STG construction pipeline. It transforms raw UI interaction trajectories into a State Transition Graph through five VLM-based stages: (1) extracting semantic captions via Mcap , (2) abstracting trajectories into state chains via Mabs , (3) aggregating chains into a global STG via Mgraph , (4) reconstructing chains for semantic consistency via Mrecon , and (5) performing fine-grained screenshot-to-state assignment via Massign . Manual verification, when used, is a separate post-hoc step and is not part of this pipeline.
Figure 8 : An illustrative example of STG construction. Stage 1: extract screen captions from raw screenshots. Stage 2: abstract state trajectories from captions and actions. Stage 3: merge state trajectories into a unified STG. Stage 4: refine state trajectories guided by the constructed STG. Stage 5: assign each screen to its best-matching state based on the screenshot, caption, and executed action.
Platform
Task
Correction Rate (%)
Android
Kitchen Stories
5.1
Maps
25.3
CNN
0.0
Web
Booking
20.8
Carmax
22.7
Kayak
23.7
Appendix
Table 9 : Post-hoc screenshot-to-state correction rates across tasks. A correction changes the Stage 5 assignment but does not redefine state semantics or graph topology.
Figure 9 : Human evaluation of STG state classification accuracy. Annotators are shown screenshots and asked whether each matches its assigned state label and description. Numbers indicate the count of confirmed matches per state.
Figure 10 : Analysis of SR contribution across state frequency intervals for 8 models over 6 tasks. Across all settings, low-frequency states contribute marginally to the overall SR regardless of their individual performance, confirming that SR systematically underrepresents agent behavior in the long tail.
Figure 11 : Cumulative distribution of failures across states ranked by failure frequency for 8 models over 6 tasks. The top 20% of states account for the majority of total errors across all settings, indicating strong and consistent failure concentration.
Task
Gini (SFT)
Gini (DPO)
Jaccard
Spearman ρ
Maps
0.648
0.571
0.333
0.972
Kitchen Stories
0.464
0.449
1.000
0.905
CNN
0.654
0.644
1.000
0.863
Mean
0.589
0.555
0.778
0.913
Appendix
Table 10: Controlled comparison of failure patterns between UI-TARS-7B-SFT and UI-TARS-7B-DPO. Jaccard measures the overlap between their Top-20% bottleneck states, and Spearman ρ compares the full state-level failure rankings.
Panel A: Results by task
Task
N
Mean Δ SR (%)
I/T/D
CNN
7
+0.87
4/3/0
Kitchen Stories
7
+3.29
5/2/0
Maps
7
+2.98
4/3/0
Arts & Culture
7
+6.04
5/0/2
Artsy
7
+4.37
5/0/2
Appendix
Table 11: Full three-fold trajectory-held-out guidance results with fully automatic STGs, summarized by task (Panel A) and agent (Panel B). N is the number of evaluated task–agent pairs, and I/T/D denotes improved/tied/decreased. Overall values average all 57 available pairs.
Model
Maps
Kitchen
CNN
Booking
Carmax
Kayak
Mean
Qwen3-VL (2B/4B/8B avg.)
+9.05
+7.75
+3.79
+2.85
+0.17
+0.40
+4.00
UI-Venus-1.5-2B
+3.11
+2.32
+7.95
+1.67
+3.04
−3.32
+2.46
UI-Venus-1.5-8B
−0.38
+1.44
+9.97
+4.72
+4.30
−1.52
+3.09
GLM-4.6V-Flash
−0.61
+3.60
+11.49
+2.62
−1.69
−2.28
+2.19
Appendix
Table 12: Guidance performance ( Δ SR, %) on recent model families over the six primary tasks. The Qwen3-VL row is averaged across its 2B, 4B, and 8B variants; the remaining rows report individual models.
Categories
TCF
SPF
IGF
IKF
Others
Percentage
28.7
27.3
18.2
11.2
14.6
Appendix
Table 13: Distribution of failure categories (%) observed in GUI agent trajectories. The results reveal a highly structured failure space, where Task Completion Failures (TCF), State Perception Failures (SPF), Instruction Grounding Failures (IGF), and Interaction Knowledge Failures (IKF) account for the majority of errors, while a small portion falls into the Others category.
Figure 12 : Prompt for State Trajectory Abstraction.
Figure 13 : Prompt for State Transition Graph Construction.
Figure 14 : Prompt for State Chain Reconstruction.
Figure 27
Figure 15 : Prompt for State Assignment. (Cont.)
Figure 16 : Prompt for GUITAR-Guided Inference. (Part 1)
Figure 17 : Prompt for GUITAR-Guided Inference. (Cont.)