WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
Organizations: Independent Researcher
Abstract
We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same model in all four submissions, the rise of our official hidden-set score from 31.0 to 57.0 reflects changes to the harness, up to run-to-run variance on live sites. We describe the design, the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models. Code is available at https://github.com/jianganghan/WebFovea.
Figures & tables
| Agent | Model | SR (%) |
|---|---|---|
| Claude Computer Use | claude-4-8-opus † | 47 |
| Kimi Agent | Kimi K3 † | 28 |
| Google Computer Use | gemini-2.5-pro | 21 |
| Claude Computer Use | claude-4-5-sonnet | 16 |
| Browser Use | gpt-4o | 11 |
| Agent-E | gpt-4o | 9 |
| Stage | What goes wrong (example from our runs) | Components |
|---|---|---|
| Parsing (§ 4.2 ) | chat-template tokens typed into search boxes (4.9% of episodes); an invented trajectory with a fake final answer | first-action truncation; stripping special tokens and stray trailing markup |
| Execution (§ 4.3 ) | every click at 3/4 of its intended coordinates; native dropdowns that clicks cannot operate; keystrokes lost to the wrong focus; clicks swallowed inside iframes; an Enter the prompt promised but the executor never sent | coordinate-space alignment; select ; fill ; Enter on newline; iframe DOM fallback |
| Feedback (§ 4.4 ) | a successful iframe click reported as “no detectable change”, steering the model away; failed inputs invisible in the screenshot | DOM and child-frame fingerprints; notes naming the element hit; fill read-back; tooltip text hash |
| Observation (§ 4.5 ) | values that appear only on hover; targets far down long pages; digits hidden by custom fonts; verification tools not used when needed | screenshot size cap; hover ; find_text ; read_text ; rule-based triggers |
| Guardrails (§ 4.6 ) | shortcuts through URLs and data APIs; endless repetition; tasks lost to a crashed worker or a failed API call | closed action set; hard constraints; executor budgets; scroll breaker; retries and dispatch |
| Action | Purpose (per-task budget) |
|---|---|
| click , left_double , right_single , drag , scroll , wait | Basic mouse and wait actions (inherited) |
| hover | Reveal tooltips / chart values |
| select(box, option) | Native <select> dropdowns |
| fill(box, content) | Focus, clear, type, read back |
| type(content) | Keystrokes; trailing \n presses Enter |
| hotkey(key) | Page-level keys; browser-level keys blocked |
| Date | Score | Main additions |
|---|---|---|
| Aug 29 | 31.0 | coordinate alignment, select , hover , reload ; closed action set, hard constraints |
| Aug 31 | 41.0 | find_text , back ; shared-queue dispatch |
| Sep 1 | 46.0 | token stripping; fill , Enter on newline; read_text ; scroll-breaker fix |
| Sep 2 | 57.0 | iframe fallback; find_text / read_text search child frames; 80 steps; model-call retries |
| Component | Before after | |
|---|---|---|
| Parsing | ||
| Token stripping (offline replay; = steps) | 7,705 | 551/551 fixed; 0 clean steps changed |
| Execution | ||
| Coordinate alignment ∗ | 24 | 8 20 (+12/ 0) |
| on a real page ( = targets) | 3 | miss by 122–189 px 1–5 px |
| fill + Enter on newline | 10 | 2 4 (+2/ 0); 12% fewer steps |
| Bucket | Tasks |
|---|---|
| Answered correctly (best across runs) | 55 |
| Reference answer outdated or defined differently | 13 |
| Site unreachable from our network | 16 |
| Structural: image-grid CAPTCHA, canvas-only UI, answer needs a typed URL or spans several pages | 6 |
| Capability gaps (iframe, chart, control) | 4 |
| Path not found / wrong decision | 3 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Stage | Component | Origin | Stage | Component | Origin |
|---|---|---|---|---|---|
| Parsing | first-action truncation | new | Observation | 5-screenshot window + full text history | inherited |
| Parsing | special-token / trailing-markup stripping | new | Observation | screenshot size cap (1440 810) | new |
| Parsing | answer extraction from finished | modified | Observation | hover exposed to the model | modified |
| Execution | coordinate-space alignment | new | Observation | find_text (frame-aware, returns no text) | new |
| Execution | select for native dropdowns | new | Observation | read_text with de-obfuscation | new |
| Execution | fill with date handling | new | Observation | rule-based tool triggers | new |