Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories
Organizations: Techtouch, Inc. · Institute of Science Tokyo
Abstract
Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.
Figures & tables
| 1. Step inputs Task: What are the top three best selling products in Jan. 2023? Guide Text: Open Reports Bestsellers; set Period to Month; enter 01/01/23–01/31/23; run the report and verify three rows. Current state: URL, simplified HTML, action history, previous action, and current memory. | 2. Action prompt and model output Task Instruction: {task} Action History: {history} Agent Memory: {memory} Simplified HTML: {page_text} HINT / workflow data: {guide_text} Output: do(action="Click", element="549") |
| Model and support | Steps | Raw (%) | Corr. (%) | Rec. (pp) |
|---|---|---|---|---|
| Qwen3.5 9B, untrained | 25 | 23 (13.90) | – | – |
| Qwen3.5 9B, untrained + MASM | 25 | 31 (18.80) | – | – |
| GPT 5.5 + MASM | 15 | 49 (29.70) | 58 (35.15) | 9 (+5.45) |
| GPT 5.5 + Guide Text + MASM | 15 | 49 (29.70) | 61 (36.97) | 12 (+7.27) |
| GPT 5.5 + MASM | 25 | 43 (26.06) | 57 (34.55) | 14 (+8.49) |
| GPT 5.5 + Guide Text + MASM | 25 | 51 (30.91) | 63 (38.18) | 12 (+7.27) |
| Subset | Tasks | MASM | Guide + MASM | Gain/loss |
|---|---|---|---|---|
| All tasks | 165 | 58 (35.15%) | 61 (36.97%) | 8 / 5 |
| Guide available | 143 | 53 (37.06%) | 56 (39.16%) | 7 / 4 |
| No guide | 22 | 5 (22.73%) | 5 (22.73%) | 1 / 1 |
| Measure | 15 steps | 25 steps |
|---|---|---|
| Corrected MASM success | 58 (35.15%) | 57 (34.55%) |
| Corrected Guide Text + MASM success | 61 (36.97%) | 63 (38.18%) |
| Guide Text only | 8 (4.85%) | 20 (12.12%) |
| MASM only | 5 (3.03%) | 14 (8.49%) |
| Net change | +3 (+1.82 pp) | +6 (+3.64 pp) |
| Exact paired test |
| Task type | MASM | Guide Text + MASM | Net |
|---|---|---|---|
| Action | 11 (6.67%) | 16 (9.70%) | +5 |
| Mixed | 26 (15.76%) | 29 (17.58%) | +3 |
| Retrieval | 20 (12.12%) | 18 (10.91%) |
| Site | Net change |
|---|---|
| Shopping administration | +4 (+2.42 pp) |
| +2 (+1.21 pp) | |
| GitLab | +1 (+0.61 pp) |
| Shopping | +1 (+0.61 pp) |
| Map | ( pp) |
| Failure mode | Share | |
|---|---|---|
| Scroll oscillation | 32 | 31.4% |
| Exploration unfinished | 13 | 12.7% |
| Premature answer | 13 | 12.7% |
| Invalid action format | 9 | 8.8% |
| Incomplete form or edit | 8 | 7.8% |
| Exact action repetition | 7 | 6.9% |
| Measure | Value |
|---|---|
| Structured progress extractable | 91/165 (55.2%) |
| Failed runs reporting | 18/102 (17.6%) |
| Failed runs reporting | 7/102 (6.9%) |
| Analyzable action limit failures | 40 |
| Mean maximum verified progress | 56.2% |
| Action limit failures at | 8/40 (20.0%) |
| Step 1 | Step 2 | Step 3 | Step 4 | Step 5 | Step 6 | |
| Increment (pp) | 23.6 | 19.0 | 9.9 | 8.6 | 7.4 | 4.2 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Failure mode | Operational definition |
|---|---|
| Scroll oscillation | Repeated up and down scrolling without extracting new evidence or changing strategy. |
| Exploration unfinished | The run exhausts its budget before reaching the page or state needed for the task. |
| Premature answer | A final answer is emitted before required evidence or action confirmation is obtained. |
| Wrong value or entity | The agent selects a similar but incorrect item, person, date, count, or destination. |
| Click without state change | A click leaves the relevant URL, page, and task state unchanged and is not repaired. |
| Exact action repetition | The same executable action is repeated in the same state. |
| Task index | Guide Text effect | Observed mechanism |
|---|---|---|
| 0 | gain: January top three best sellers | Direct route to the report and correct date interval. |
| 18 | gain: two segment travel time | Preserves segment direction and travel modes. |
| 29 | gain: product search for bruxism | Supplies a useful query and explicit completion condition. |
| 130 | gain: add two GitLab maintainers | Provides the correct members page workflow. |
| 132 | gain: subscribe to a trending forum thread | Preserves forum context and verifies subscription. |
| 12 | regression: nearest cafe to Hunt Library | Incomplete guide anchors the model to an unsupported cafe. |