LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
Organizations: Vera Praxis · Tencent · The Hong Kong University of Science and Technology · Independent Researcher · Department of Computer Science and Engineering, The Chinese University of Hong Kong
Abstract
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.
Figures & tables
| What solving requires | How the agent must act | |||||
|---|---|---|---|---|---|---|
| Benchmark | Environment | Visual puzzle state | Delayed consequences | Dead-end states | Agent-owned recovery | Native GUI actions |
| PuzzleBench ( Zhang et al., 2025 ) | Puzzle VQA | – | – | – | ||
| VisualWebArena ( Koh et al., 2024 ) | Web apps | – | ||||
| AndroidWorld ( Rawles et al., 2025 ) | Mobile apps | – | ||||
| BALROG ( Paglieri et al., 2025 ) | Text/grid games | |||||
| VideoGameBench ( Zhang et al., 2026 ) | Video games | |||||
| Completion | Horizon | Difficulty | Behavior | |||||
|---|---|---|---|---|---|---|---|---|
| Agent | OSR | LC | FGP | LPR | H 50 | LC hard | PER | NGAR |
| Human reference | 100.0 | 100.0 | 100.0 | 100.0 | 10+ | 100.0 | 0.0 | – |
| Native GUI Actions general-purpose, Codex | ||||||||
| GPT-6-Astra | 91.7 | 94.8 | 94.8 | 97.9 | 10+ | 87.5 | 0.0 | – |
| GPT-6-Sol | 82.8 | 91.0 | 92.2 | 95.4 | 10+ | 80.3 | 0.0 | – |
| GPT-6-Luna | 33.9 | 40.0 | 43.8 | 46.7 | 2 | 9.0 | 68.8 | – |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Game | Difficulties | Levels | Level completion |
|---|---|---|---|
| Bolt Unscrew | Easy / Hard | 8 / 8 | Remove every board. |
| Rush Hour | Easy / Medium / Hard | 10 / 10 / 10 | Move the target vehicle through the exit. |
| Nut and Bolt | Easy / Medium / Hard / Extreme / Nightmare | 3 / 3 / 3 / 3 / 1 | Make every nonempty stack full and monochromatic. |
| Truck Escape | Default | 5 | Remove all vehicles. |
| Maze Paint | Easy / Medium / Hard | 10 / 10 / 10 | Paint every traversable cell. |
| Color Connect | Easy / Hard | 10 / 10 | Connect every matching color pair with non-overlapping paths. |
| Objective | Levels | Pointer presses | Restarts | Undos |
|---|---|---|---|---|
| Bolt Unscrew / Easy | 8 | 100 | 1 | 0 |
| Bolt Unscrew / Hard | 8 | 2,500 | 52 | 0 |
| Rush Hour / Easy | 10 | 89 | 0 | 0 |
| Rush Hour / Medium | 10 | 102 | 0 | 0 |
| Rush Hour / Hard | 10 | 442 | 12 | 0 |
| Nut and Bolt / Easy | 3 | 45 | 0 | 0 |
| Native GUI Actions | Code Execution CUA |
|---|---|
| Action space: one GUI-action tool | Action space: code runtime and GUI input |
| click(x,y) primary click | Run JavaScript in a persistent runtime |
| double_click(x,y) double click | over screenshots and derived data |
| drag(x0,y0,x1,y1) drag pointer | Request a screenshot |
| press_key(key) game key | Issue the same coordinate GUI inputs, |
| type_text(text) text entry | directly or from code |
| Agent | Endpoint | Effort | Limit | Price (in / cached / out) |
| General-purpose, Codex scaffold | ||||
| GPT-6-Astra | OpenAI | medium | 258 | 10.00 / 1.00 / 50.00 |
| GPT-6-Sol | OpenAI | medium | 258 | 2.00 / 0.20 / 10.00 |
| GPT-6-Luna | OpenAI | medium | 258 | 0.10 / 0.01 / 0.50 |
| GPT-5.6-Sol | OpenAI | medium | 258 | 4.00 / 0.40 / 20.00 |
| GPT-5.6-Terra | OpenAI | medium | 258 | 2.00 / 0.20 / 12.00 |
| Agent | Requests | Input (M) | Cached (%) | Output (M) | Cost (USD) |
|---|---|---|---|---|---|
| Native GUI Actions | |||||
| GPT-6-Astra | 1,286 | 100.7 | 97.9 | 0.19 | 129.70 |
| GPT-6-Sol | 2,592 | 237.5 | 98.6 | 0.48 | 58.41 |
| GPT-6-Luna | 2,713 | 288.4 | 98.2 | 0.37 | 3.53 |
| GPT-5.6-Sol | 3,077 | 327.5 | 98.3 | 0.55 | 161.87 |
| GPT-5.6-Terra | 2,132 | 218.6 | 98.1 | 0.44 | 56.32 |
| Native GUI Actions | Code Execution CUA | |
| Outcomes | ||
| Episodes solved / timed out / abandoned | 84 / 41 / 35 | 92 / 31 / 37 |
| Timeouts that solve a level in the final 15 min | 15 of 41 | 7 of 31 |
| Level restarts per episode (mean) | 2.4 | 3.6 |
| Interaction | ||
| Tool calls per episode (median) | 100.0 | 80.5 |
| Puzzle | Diagnostic | Human | Astra, Native | Astra, Code |
|---|---|---|---|---|
| Bolt Unscrew Hard L4 | Dead ends with no board removed | 1 | 1 | 5 |
| Attempt that clears the level | 3 | – | 8 | |
| Most boards in one attempt | 14 | 0 | 14 | |
| Nut and Bolt Nightmare | Stack-state changes | 431 | 65 | 67 |
| Revisited states | 115 | 4 | 0 | |
| Solved-stack regressions | 7 | 0 | 0 |
| Agent | Wins/reached | Conditional (%) | Unreached (%) | Exits/failures | Exit time |
|---|---|---|---|---|---|
| GPT-6-Astra | 109/110 | 99.1 | 80.0 | 0/1 | N/A |
| GPT-6-Sol | 104/107 | 97.2 | 70.0 | 0/3 | N/A |
| GPT-6-Luna | 37/49 | 75.5 | 79.7 | 11/12 | 19.5 [5.4, 33.7] |
| GPT-5.6-Sol | 97/102 | 95.1 | 70.4 | 1/5 | 20.6 |
| GPT-5.6-Terra | 68/77 | 88.3 | 74.3 | 9/9 | 17.2 [13.7, 21.0] |
| GPT-5.6-Luna | 58/68 | 85.3 | 76.3 | 10/10 | 7.0 [5.5, 8.6] |
| Agent | Maze LC (%) | Valid/attempted | Productive/valid (%) | Revisited/observed (%) |
|---|---|---|---|---|
| GPT-6-Astra | 100.0 | 557/557 | 441/557 (79.2) | 6/587 (1.0) |
| GPT-6-Sol | 90.0 | 505/508 | 386/505 (76.4) | 7/537 (1.3) |
| GPT-6-Luna | 50.0 | 276/385 | 169/276 (61.2) | 143/408 (35.0) |
| GPT-5.6-Sol | 96.7 | 687/714 | 446/687 (64.9) | 64/745 (8.6) |
| GPT-5.6-Terra | 56.7 | 361/541 | 245/361 (67.9) | 192/566 (33.9) |
| GPT-5.6-Luna | 66.7 | 451/863 | 256/451 (56.8) | 466/890 (52.4) |
| Condition/state | Appended information |
|---|---|
| B/D | State information: The current puzzle state cannot be completed through normal forward play. Recovery is required before progress can continue. |
| B/N | State information: The current puzzle state remains solvable through normal forward play. |
| C/D | B’s D message, followed by: Recovery information: Use the visible button labelled Restart to restart the current level. |
| State | Condition | Timely restart | Verified recovery | Same-level win |
|---|---|---|---|---|
| D | A | 1/24 (4.2%) | 8/24 (33.3%) | 0/24 (0.0%) |
| D | B | 22/24 (91.7%) | 23/24 (95.8%) | 0/24 (0.0%) |
| D | C | 22/24 (91.7%) | 22/24 (91.7%) | 0/24 (0.0%) |
| N | A | 0/12 (0.0%) | — | 10/12 (83.3%) |
| N | B | 0/12 (0.0%) | — | 9/12 (75.0%) |
| Hard level | D checkpoints | Selection | Recovery | Wins (A/B/C) |
|---|---|---|---|---|
| 1 | 10 | +80.0 | +50.0 | 0/0/0 |
| 2 | 2 | +100.0 | +100.0 | 0/0/0 |
| 3 | 2 | +100.0 | +50.0 | 0/0/0 |
| 5 | 2 | +100.0 | +50.0 | 0/0/0 |
| 6 | 4 | +75.0 | +75.0 | 0/0/0 |
| 7 | 2 | +100.0 | +50.0 | 0/0/0 |
| Game | Model | A | B | C |
|---|---|---|---|---|
| Bolt Unscrew | Luna | 0/6 | 5/6 | 5/6 |
| Bolt Unscrew | Sol | 2/6 | 5/6 | 6/6 |
| Bolt Unscrew | Terra | 5/6 | 6/6 | 5/6 |
| Nut and Bolt | All three | 0/9 | 0/9 | 0/9 |
| Contrast or cell | Completion | LC | FGP |
|---|---|---|---|
| K0P0 | 5/18 | – | – |
| K0P1 | 6/18 | – | – |
| K1P0 | 8/18 | – | – |
| K1P1 | 8/18 | – | – |
| K1–K0 under P0 | – | ||
| K1–K0 under P1 | – |
| Configuration | S | PER | SMR | Native | Other | Cells | |
|---|---|---|---|---|---|---|---|
| GPT-6-Astra / Code Execution CUA | 16 | 15 | 0 | 0 | 729 | 381 | 16/16 |
| GPT-6-Sol / Code Execution CUA | 16 | 13 | 3 | 0 | 486 | 696 | 16/16 |
| GPT-6-Luna / Code Execution CUA | 16 | 4 | 12 | 0 | 1620 | 858 | 16/16 |
| GPT-5.6-Sol / Code Execution CUA | 16 | 14 | 0 | 0 | 122 | 1316 | 16/16 |
| GPT-5.6-Terra / Code Execution CUA | 16 | 6 | 10 | 0 | 528 | 1067 | 16/16 |
| GPT-5.6-Luna / Code Execution CUA | 16 | 5 | 11 | 0 | 629 | 568 | 16/16 |
| Objective | S | LC | FGP | P | M | NGAR | |
|---|---|---|---|---|---|---|---|
| GPT-6-Astra / Code Execution CUA | |||||||
| Bolt / easy | 1 | 100.0 | 100.0 | 0 | 0 | 100.0 | 8 |
| Bolt / hard | 0 | 50.0 | 60.7 | 0 | 0 | 65.2 | 8 |
| Color / easy | 1 | 100.0 | 100.0 | 0 | 0 | 38.6 | 10 |
| Color / hard | 1 | 100.0 | 100.0 | 0 | 0 | 25.0 | 10 |
| Maze / easy | 1 | 100.0 | 100.0 | 0 | 0 | 92.0 | 10 |