Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.
Figures & tables
1. Step inputs Task: What are the top three best selling products in Jan. 2023? Guide Text: Open Reports → Bestsellers; set Period to Month; enter 01/01/23–01/31/23; run the report and verify three rows. Current state: URL, simplified HTML, action history, previous action, and current memory. ⟶
2. Action prompt and model output Task Instruction: {task} Action History: {history} Agent Memory: {memory} Simplified HTML: {page_text} HINT / workflow data: {guide_text} Output: do(action="Click", element="549")
Figure 1: Prompt and state flow for Guide Text and MASM. Guide Text supports action planning and memory update, while action result analysis is grounded in the executed transition and environment evidence.
Model and support
Steps
Raw n (%)
Corr. n (%)
Rec. n (pp)
Qwen3.5 9B, untrained
25
23 (13.90)
–
–
Qwen3.5 9B, untrained + MASM
25
31 (18.80)
–
–
GPT 5.5 + MASM
15
49 (29.70)
58 (35.15)
9 (+5.45)
GPT 5.5 + Guide Text + MASM
15
49 (29.70)
61 (36.97)
12 (+7.27)
GPT 5.5 + MASM
25
43 (26.06)
57 (34.55)
14 (+8.49)
GPT 5.5 + Guide Text + MASM
25
51 (30.91)
63 (38.18)
12 (+7.27)
Table 1: Retained WebArena Lite evaluation results. Raw and Corr. report successful tasks as count (percentage). Rec. is the number recovered from raw failures and the percentage point gain. The Qwen rows form an evaluator only MASM ablation and were not included in the GPT 5.5 human correction study.
Subset
Tasks
MASM
Guide + MASM
Gain/loss
All tasks
165
58 (35.15%)
61 (36.97%)
8 / 5
Guide available
143
53 (37.06%)
56 (39.16%)
7 / 4
No guide
22
5 (22.73%)
5 (22.73%)
1 / 1
Table 2: Corrected GPT 5.5 outcomes by Guide Text availability at 15 steps. Gain and loss count paired task flips relative to MASM only.
Measure
15 steps
25 steps
Corrected MASM success
58 (35.15%)
57 (34.55%)
Corrected Guide Text + MASM success
61 (36.97%)
63 (38.18%)
Guide Text only
8 (4.85%)
20 (12.12%)
MASM only
5 (3.03%)
14 (8.49%)
Net change
+3 (+1.82 pp)
+6 (+3.64 pp)
Exact paired test
p=.549
p=.392
Table 3: Matched GPT 5.5 comparison with and without Guide Text.
Task type
MASM
Guide Text + MASM
Net
Action
11 (6.67%)
16 (9.70%)
+5
Mixed
26 (15.76%)
29 (17.58%)
+3
Retrieval
20 (12.12%)
18 (10.91%)
−2
Table 4: Corrected GPT 5.5 success by task type at 25 steps.
Site
Net change
Shopping administration
+4 (+2.42 pp)
Reddit
+2 (+1.21 pp)
GitLab
+1 (+0.61 pp)
Shopping
+1 (+0.61 pp)
Map
−2 ( −1.21 pp)
Table 5: Net corrected Guide Text effect by site at 25 steps.
Failure mode
n
Share
Scroll oscillation
32
31.4%
Exploration unfinished
13
12.7%
Premature answer
13
12.7%
Invalid action format
9
8.8%
Incomplete form or edit
8
7.8%
Exact action repetition
7
6.9%
Table 6: Primary failure modes for 102 failed GPT 5.5 trajectories at 25 steps with Guide Text and MASM.
Measure
Value
Structured progress extractable
91/165 (55.2%)
Failed runs reporting ≥80%
18/102 (17.6%)
Failed runs reporting 100%
7/102 (6.9%)
Analyzable action limit failures
40
Mean maximum verified progress
56.2%
Action limit failures at ≥80%
8/40 (20.0%)
Table 7: GPT 5.5 progress evidence for the 25 step Guide Text and MASM condition.
Step 1
Step 2
Step 3
Step 4
Step 5
Step 6
Increment (pp)
23.6
19.0
9.9
8.6
7.4
4.2
Table 8: Mean reported GPT 5.5 progress increments over the first six actions.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Failure mode
Operational definition
Scroll oscillation
Repeated up and down scrolling without extracting new evidence or changing strategy.
Exploration unfinished
The run exhausts its budget before reaching the page or state needed for the task.
Premature answer
A final answer is emitted before required evidence or action confirmation is obtained.
Wrong value or entity
The agent selects a similar but incorrect item, person, date, count, or destination.
Click without state change
A click leaves the relevant URL, page, and task state unchanged and is not repaired.
Exact action repetition
The same executable action is repeated in the same state.
Appendix
Table 10
Task index
Guide Text effect
Observed mechanism
0
gain: January top three best sellers
Direct route to the report and correct date interval.
18
gain: two segment travel time
Preserves segment direction and travel modes.
29
gain: product search for bruxism
Supplies a useful query and explicit completion condition.
130
gain: add two GitLab maintainers
Provides the correct members page workflow.
132
gain: subscribe to a trending forum thread
Preserves forum context and verifies subscription.
12
regression: nearest cafe to Hunt Library
Incomplete guide anchors the model to an unsupported cafe.
Appendix
Table 9: Representative Guide Text gains and regressions from paired GPT 5.5 runs.
Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, and terminate confidently while still missing fields, over-including unsupported items, or relying on stale evidence. We study these failures with Parallel WebBench, a parallel web-exploration benchmark containing 1,679 verified records: 350 manually curated parallel tasks and 1,329 reconstructed records with verified URL-based trajectories. We train WebExplorer-style agents with GRPO under human-only, balanced human-synthetic, and synthetic-heavy data mixtures. At 16k context and 16 interaction rounds, the best GRPO model improves completion over WebExplorer-8B from 50.7% to 96.0% and GPT-4.1-mini-judged element-wise F1 from 0.2489 to 0.4529, but binary accuracy remains far below completion. Trace-level analysis identifies three persistent failure modes: context-bound search loops, premature termination on partial answers, and synthesis collapse after relevant evidence has already been retrieved. These results show that synthetic-data GRPO reduces abstention and improves partial correctness, but leaves a completion-correctness gap that requires evidence-grounded coverage and synthesis diagnostics.
Aagam Sogani, Botao Rui, Swetha Vaidyanathan +3
University of Wisconsin–Madison, Madison, WI, USA.
Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task success. AgentAtlas reframes agent evaluation as a diagnostic vocabulary and audit protocol for separating outcome success from control-decision quality and trajectory quality. The paper contributes: (i) a six-state control-decision taxonomy (Act / Ask / Refuse / Stop / Confirm / Recover); (ii) a trajectory-failure vocabulary with primary error source and downstream impact; (iii) a 0/1/2 benchmark-coverage audit over fifteen agent benchmarks; and (iv) an illustrative protocol study on a synthetic 1,342-item set evaluated with eight models under taxonomy-aware and taxonomy-blind prompt formats. The synthetic demonstration is not a public benchmark release and should not be read as a definitive model comparison. Instead, it illustrates two measurement risks: mapped label agreement can change substantially when the explicit label menu is removed, and axis choice can change apparent rankings. AgentAtlas is intended to help benchmark designers state what behavior they cover, and to help evaluators diagnose failures that outcome-only leaderboards hide.
Parsa Mazaheri, Kasra Mazaheri
University of California, Santa Cruz · Massachusetts Institute of Technology
HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods reduce overall agent latency while maintaining performance. The main obstacle is the high cost of end-to-end evaluation: in our experiments, evaluating 11 methods across 32 configurations on 33 tasks of WorkArena L1 required 232.4 cumulative hours. To address this, we propose a lightweight evaluation framework based on the Minimal Failure Set (MFS), the minimal set of HTML elements whose removal causes task failure. We define coverage as the fraction of instances in which a reduction method fully retains the MFS, which serves as a proxy metric that requires neither web access nor LLM inference. We validate that coverage strongly correlates with end-to-end success rate, with over 100× speedup in cumulative evaluation time on both benchmarks. Using this framework, we find that extractive HTML reduction methods require either high computation cost or domain-specific optimization to reduce agent latency while maintaining performance. Building on this, we optimize a pruning program on MFS training data, achieving 2.2× faster per-step latency on WorkArena L1 while retaining 84% of the original success rate, and 3.1× faster on WebLinx while retaining 89%.