We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).
Figures & tables
Figure 1: One task, one condition, two interfaces. The prompt (translated) and the single success condition shown above the screenshots are identical in both task files; the released variant differs from its donor only in the profile it names ( rail_date_native ). On the left the departure date is picked from the default popup grid, where the month and the day are on screen and one click settles it. On the right the same date goes into a native <input type="date"> , so the agent has to render “4 October 2026” from the prompt in the numeric format the field expects (the screenshot shows the field with an arbitrary date typed in). The figure illustrates the variant mechanism only; the leaderboard submissions do not separate donors from variants, so no per-variant result is reported.
Figure 2: A condition and the event that satisfies it. All four implementations of the guest counter (per-room popup, inline stepper, compact select , pill buttons) emit the same event with the same parameters, so the condition holds for all of them. The widget field is recorded for analysis and never enters the verdict.
Figure 3: One run. The task supplies the prompt and the conditions, the domain configuration supplies the control implementations, and the event log is the only evidence used to decide success.
Harness
Obs.
Notes
browser-use
DOM+ screenshot ‡
baseline executor
ouroboros-cut
DOM+ screenshot
Ouroboros prompts, browser-use
ouroboros-full-isolated
text+ screenshot
Ouroboros, fresh memory
ouroboros-full-evolving
text+ screenshot
Ouroboros, shared memory †
openmanus
DOM+ screenshot ‡
OpenManus prompts, browser-use
openhands
DOM text
OpenHands SDK, browser tools †
Table 1: Harnesses evaluated in this snapshot; all are selected by the AGENT_HARNESS switch: browser-use ( Browser Use, 2024 ) , OpenManus ( FoundationAgents, 2025 ) , OpenHands ( Wang et al., 2025a ) , Qwen3-VL ( Bai et al., 2025 ) , UI-TARS ( Qin et al., 2025 ) , OpenCUA ( Wang et al., 2025b ) , EvoCUA ( Xue et al., 2026 ) and Fara-1.5 ( Awadallah et al., 2026 ) ; Ouroboros is an in-house agent runtime. DOM is an indexed element tree the agent can act on by element handle; text is extracted page text without that tree. A screenshot, when listed, is sent in addition. ‡ The DeepSeek pairs receive DOM text only, without screenshots. † Runs on a single worker; for the evolving runtime because tasks share one memory.
Tab
Activity
Canon.
Var.
Total
Маркет
marketplace
18
14
32
Книги
books, audiobooks
11
10
21
Файлы
document cabinet
11
7
18
Поезда
rail travel
9
17
26
Продукты
grocery delivery
8
2
10
Отели
hotel search
8
37
45
Table 2: Tasks per domain. Variant counts reflect the profiles available in a domain and the donors eligible for them. For example, the hotel domain combines date, city and guest-counter controls and yields 37 variants from 8 canonical tasks.
Model
Harness
EMS
Compl.
Done&pass
Steps
Dur.
Cost $
gemini-3.8-flash
openmanus
0.822
1.000
0.822
7.8
221
17.88
browser-use
0.711
0.993
0.711
9.5
220
23.65
openhands
0.572
0.461
0.454
131.7
132
53.12
gpt-5.6-luna
ouro-full-iso
0.796
0.868
0.796
27.8
177
6.69
browser-use
0.704
0.993
0.704
10.8
356
3.23
ouro-full-env
0.697
0.717
0.678
13.1
297
3.91
Table 3: Public leaderboard submissions (24 model–harness pairs), grouped by model. Models are ordered by their best EMS, rows within a model by EMS. Where a model has several harnesses, bold marks the highest EMS, the shortest duration and the lowest cost in that group; single-harness rows are not bold. ouro-full-iso is ouroboros-full-isolated , ouro-full-env is ouroboros-full-evolving , ouro-cut is ouroboros-cut . EMS is passed_tasks /152. Compl. is agent_completion_rate . Done&pass is the share of tasks the agent reported done and whose conditions all passed, counted from the run logs; it is the same count the public leaderboard shows. Steps and Dur. are per-task means: agent steps and harness time in seconds. Cost is total_cost_usd for the whole 152-task run; — when the submission has no price.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Donor task
Prompt (translated) and default controls
Released variants of this donor
Hotels hotel_search_ scenario
“In Hotels, show options: Zurich, Switzerland, check-in 15.09.2026, check-out 20.09.2026, 2 guests.” Defaults: split popup calendar, autocomplete city field, per-room guest popup.
date → inline calendar, single-field range, typed string; select_city → native select ; counter_guests → inline stepper, compact select , pill buttons; dark theme. (8 variants; four sibling hotel tasks with other cities and guest mixes carry the same profiles.)
Rail rail_book_ to_cart
“In Trains, find a Moscow–St. Petersburg train on 04.10.2026, 1 passenger, Economy tariff, and add it to the cart.” Defaults: grid popup calendar, typeahead station field.
date → native <input type="date"> , typed string; select_station → native select ; dark theme. (4 variants.)
Books digital_books_ named_product_ basket
“In Books, find Quiet Amber in the Last Carriage by Vera Rudneva (text edition) and add it to the cart.” Default: standard search field. Author and title are synthetic (§ 3.1 ).
text_search → filled, outlined, pill, underlined; dark theme. (5 variants; the marketplace search tasks carry the same four search profiles.)
Files files_ download_pdf
“In Files, open Lab reports , choose 2024 and download lab-reports-2024.pdf .” Defaults: card layout, year select , outline buttons.
one three-key profile: collections → list, years → buttons, buttons → icon; dark theme. (2 variants.)
Appendix
Table 4: One donor per domain with its released variants. The prompt and the conditions of every variant are identical to the donor’s; only the named control changes. Prompts are translated from Russian; identifiers are the task file names without the profile suffix.
Class
Control
Solved
BASKET
add / remove / quantity
48/66
COUNTER
stepper, guest count
26/33
FILES
collections, year, download
18/18
FAV
favourites toggle
14/14
CARD
card in a grid or carousel
9/9
DATE
date or range picker
4/4
Appendix
Table 5: Primary classes in the public submissions. Solved is passed/total from ui_classes for gemini-3.8-flash × openmanus (best EMS). Every submission contains these nine classes and no others. Classes with four or fewer tasks are not a ranking. NAV , RADIO , SEARCH and SEAT are absent from ui_classes and are omitted.
Success decided by
Benchmark
Environment
Dom.
Tasks
Verification
Controlled UI variation
Agents
Reference match
Mind2Web ( Deng et al., 2023 )
offline, branded
31
2,350
trajectory match
—
—
AssistantBench ( Yoran et al., 2024 )
live, branded
258
214
answer match
—
BrowserGym
Judge model
WebVoyager ( He et al., 2024 )
live, branded
15
643
GPT-4V judge
—
—
InSTA ( Trabucco et al., 2025 )
live
∼150 K
146,441
LLM judge
—
Gym env
BookingArena ( Logeswaran et al., 2026 )
live, branded
20
120
VLM constraint judge
—
—
GTA ( Huang et al., 2026 )
live, crawled
50+
5,600
answer + path replay
—
—
Appendix
Table 6: Web-agent benchmarks grouped by what decides success (first column): a reference that never inspects the environment, a judge model, programmatic state plus a model, or a program alone. Environment : live, offline, self-hosted, simulated or generated sites and their branding; Dom. : distinct domains or sites; Agents : how agents are connected (descriptive, not a count). Task counts as reported by each paper. Among the systems surveyed, none re-renders a fixed task through an alternative implementation of a control (WorkArena++ changes colours and logos only), and multi-agent support comes through one shared ecosystem.