cs.CLSep 28, 2026

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

Authors: Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang, Chaozheng Wang

Organizations: Vera Praxis · Tencent · The Hong Kong University of Science and Technology · Independent Researcher · Department of Computer Science and Engineering, The Chinese University of Hong Kong

Abstract

GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

    May 11, 2026Maximilian Triebel, Marco Menner, Dominik HelfensteinPuzzleWeak Visual Grounding

  2. EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding

    Jul 19, 2026Yaohan Yang, Minglei Shi, Borui Zhang +2

  3. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

    Jul 9, 2026Zongxia Li, Zhongzhi Li, Yucheng Shi +10Long-Horizon AgentsLong-Horizon Task Planning