PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence
Authors: Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek
Organizations: University of Amsterdam · Eindhoven University of Technology · Eberhard Karls Universität Tübingen
Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and itch.io. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
Figures & tables
Models
Evaluation
Benchmark
Domain
# Games
LLM
VLM
VLA
CUA
Metric
HPC
Dynamic
GameBench ( Costarelli et al., 2024 )
Text
9
✓
✗
✗
✗
Game score
✗
✗
GameArena ( Hu et al., 2025 )
Text
6
✓
✗
✗
✗
Win rate
✗
✗
Balrog ( Paglieri et al., 2025 )
2D
6
✓
✓
✗
✗
Task completion
✗
✗
ING-VP ( Zhang et al., 2024 )
2D
6
✗
✓
✗
✗
Task completion
✗
✗
V-MAGE ( Zheng et al., 2025 )
2D/3D
5
✗
✓
✗
✗
ELO ranking
✗
✗
Table 1: Related work. Comparison of PlaySuite against existing video game benchmarks.
Figure 2 : Examples from the PlaySuite game corpus. PlaySuite contains a broad collection of independently developed games with varied visual styles, mechanics, camera perspectives, and controls. These examples illustrate the range of settings agents must handle, from 2D platform and puzzle games to racing, exploration, shooting, and 3D navigation environments.
Figure 3 : Batched multi-worker execution. Comparison between a sequential one-game-per-model-process evaluation pattern and the PlaySuite batched architecture. Multiple sandboxed game workers run on one node, each with an isolated virtual display, while a shared vLLM-backed inference worker batches ready frame-and-prompt requests on the GPU.
itch.io
PyWeek
Model
History
No History
Partial+
History
No History
Partial+
Vision-Language Models (VLMs)
Qwen3-Omni-30B-A3B
0.62
0.84
19.5%
0.89
1.26
37.0%
Qwen2.5-VL-7B
0.41
0.82
17.7%
0.67
1.18
31.7%
Gemma-4-26B-A4B-it
1.02
0.71
17.3%
1.37
0.96
24.8%
Qwen3-VL-8B
0.64
0.70
15.5%
0.73
1.07
28.3%
Table 2: Main benchmark results across PlaySuite . Mean Progress Score under the history ( k=3 ) and no-history conditions, together with the percentage of no-history runs reaching at least Partial progress, across 5,630 itch.io games and 104 PyWeek games. Progress is measured on the five-level ordinal scale from None (0) to Completed (4). Best results are shown in bold.
Genre
N
Mean
Partial+
Simulation
506
0.893
27.1%
Strategy
223
0.858
25.2%
Shooter
205
0.771
20.4%
Survival
171
0.770
21.7%
Racing
72
0.705
18.9%
Action
958
0.620
14.4%
Table 3 : Progress by genre on itch.io . Mean Progress Score and Partial+ rate in the no-history condition for genres with at least 60 games. PyWeek and history results are reported in Appendix A.9.1 .
Figure 4 : Failure mechanisms for Gemma-3 on Goob. The model correctly perceives the position of the spike and blue slime, but incorrectly places the coin on the right side of the screen, producing a perception error. Although it recognizes that the instructions specify the arrow keys for movement, it repeatedly presses “d” despite receiving no progress, illustrating failures in control selection, instruction following, and recovery from repetitive actions.
Figure 6 : Impact of history length on model performance. Mean Progress Score across history lengths on the 200-game itch subset. Faint lines show individual models and the bold line shows the mean across models.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Harmful Categories
Game Count
Firearms Individual Violence
13
Scheming Manipulation
35
Value Misalignment
91
Inconsistent Behavior
24
Explosives Violence
3
Economic Waste
12
Appendix
Table 4 : Distribution of harmful content categories identified in the itch.io dataset.
Reason
Games
Construct runtime freeze failure
225
Dragging is the only listed control
72
Requires multiple players
25
Not an interactive game
18
Gamepad required
2
Appendix
Table 5: Additional technical exclusions from the evaluation set. Number of games excluded and their identified reason.
Original itch.io genres
PlaySuite genre
Puzzle, Educational
Puzzle
Adventure, Visual Novel, Interactive Fiction, Role Playing
Adventure
Action, Platformer, Rhythm
Action
Simulation, Sports
Simulation
Strategy, Card Game
Strategy
Shooter
Shooter
Appendix
Table 6 : Mapping of original itch.io metadata tags to the 9 canonical PlaySuite genres.
Figure 8 : Genre distribution of the PlaySuite evaluation set. Distribution of the 5,630 evaluated itch.io games and 104 PyWeek games across the nine canonical PlaySuite genres.
Model
Family
Parameters
Hugging Face identifier
Reference
Qwen3-Omni-30B-A3B-Instruct
VLM
30B (3B active)
Qwen/Qwen3-Omni-30B-A3B-Instruct
( Xu et al., 2025b )
Qwen2.5-VL-7B-Instruct
VLM
7B
Qwen/Qwen2.5-VL-7B-Instruct
( Bai et al., 2025b )
Gemma-4-26B-A4B-it
VLM
26B (4B active)
google/gemma-4-26B-A4B-it
( Gemma Team, 2026 )
Qwen3-VL-8B-Instruct
VLM
8B
Qwen/Qwen3-VL-8B-Instruct
( Bai et al., 2025a )
Qwen2.5-Omni-7B
VLM
7B
Qwen/Qwen2.5-Omni-7B
( Xu et al., 2025a )
Gemma-4-E4B-it
VLM
8B (4.5B eff.)
google/gemma-4-E4B-it
( Gemma Team, 2026 )
Appendix
Table 7 : Baseline models. Model family, parameter scale, Hugging Face checkpoint, and reference for each evaluated model.
Model
Parser
Prompt template
Gemma-4-26B-A4B-it
qwen3
standard
EvoCUA-8B
computer_use
computer_use
Qwen3-Omni-30B-A3B
qwen3
standard
Gemma-4-E4B-it
qwen3
standard
UI-TARS-1.5-7B
uitars
uitars
Qwen3-VL-8B
qwen3
standard
Appendix
Table 8 : Model interfaces. Output parser and prompt template used for each evaluated model.
Score
Level
Condition
4
Completed
win_or_completion_screen
3
Substantial
checkpoint_or_objective_completed or ( new_area_or_level_reached and score_or_resource_increased )
2
Partial
new_area_or_level_reached or score_or_resource_increased
1
Minimal
player_control_evident and distinct_interactions ≥2
0
None
No condition above is satisfied.
Appendix
Table 9: Deterministic mapping from milestones to progress. Evaluate rows from top to bottom and return the first matching level. Conditions do not require the lower-level milestones to hold. Field names match the judge output.
Judge
Milestone agreement
Within one
Qwen3.5-9B
84.4%
90.4%
Qwen3.5-27B
83.9%
91.6%
Gemma-4-31B-it
82.5%
77.3%
Molmo2-8B
77.9%
81.2%
Appendix
Table 10: Judge–human agreement on the validation set. Agreement between each Video-LLM judge and the six human annotators, averaged over available comparisons. Milestone agreement is computed over the nine Boolean milestone fields.
Figure 10 : Measured execution efficiency. (a) Fixed overhead per game before the first captured frame, from 50-game launch tests. The naive launcher waits fixed intervals for the virtual display and the game page; readiness polling replaces these waits, and browser reuse keeps the browser and display running between games. (b) Qwen3-VL-8B on one H100, with concurrent game workers sharing a single vLLM server and the same games evaluated at each concurrency level. Throughput increases with worker count. (c) Increasing Qwen3.5-9B judge throughput through video prefetching and a larger batched-token budget ( max_num_batched_tokens =32,768) on 300s videos.
itch.io
PyWeek
Genre
N
k=0
k=3
N
k=0
k=3
Simulation
506
0.893
0.800
13
1.293
1.173
Strategy
223
0.858
0.704
10
1.054
0.702
Shooter
205
0.771
0.682
6
0.763
0.720
Survival
171
0.770
0.644
5
1.174
0.857
Racing
72
0.705
0.654
3
0.976
0.714
Appendix
Table 11: Genre-level progress under both history conditions. Mean Progress Score for history lengths k=0 and k=3 on itch.io and PyWeek . N denotes the number of games.
Figure 11 : Per-model progress by genre on itch.io . Mean Progress Score in the no-history condition at 0.33 Hz for each model and genre. Columns include genres with at least 60 games; n denotes the number of games. Models are ordered by overall no-history performance.
Figure 12 : Per-model progress by genre on PyWeek . Mean Progress Score in the no-history condition at 0.33 Hz for each model and genre. n denotes the number of games. Models are ordered by overall no-history performance. Per-genre counts are small, so these results are primarily descriptive.
Figure 13 : Effect of action frequency on PyWeek . Mean Progress Score at 0.33 Hz and 3 Hz under the no-history and k=3 history conditions. Faint lines show individual model means and darker lines the mean over matched game–model pairs. Error bars denote 95% game-bootstrap intervals.
Figure 14 : Effect of prompt additions. Mean Progress Score with and without an explicit Strategy field (left) or three few-shot examples (right), for eight VLMs at k=3 . Each comparison uses matched game–model cells. Error bars show 95% game-bootstrap intervals for the means.
Figure 15 : Impact of temporal history across both corpora. Mean Progress Score for history lengths k∈{0,3,5,10} on PyWeek and the 200-game itch.io subset. No history achieves the highest mean performance on both corpora, while longer histories provide no consistent improvement.
Figure 16 : Cost–performance trade-off on itch.io . Mean Progress Score against H100 GPU-hours per 1,000 scored games in the no-history setting. Marker color denotes model family and marker size denotes parameter count.
Figure 17 : Example of a passive wait failure mechanism. Qwen2.5-VL successfully starts the game, but since the initial game screen shows no visible enemies or apparent obstacles, it assumes the game has ended and waits, even if no on-screen information confirms that the game has been completed.