PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence
Organizations: University of Amsterdam · Eindhoven University of Technology · Eberhard Karls Universität Tübingen
Abstract
Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and itch.io. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
Figures & tables
| Models | Evaluation | ||||||||
| Benchmark | Domain | # Games | LLM | VLM | VLA | CUA | Metric | HPC | Dynamic |
| GameBench ( Costarelli et al., 2024 ) | Text | 9 | ✓ | ✗ | ✗ | ✗ | Game score | ✗ | ✗ |
| GameArena ( Hu et al., 2025 ) | Text | 6 | ✓ | ✗ | ✗ | ✗ | Win rate | ✗ | ✗ |
| Balrog ( Paglieri et al., 2025 ) | 2D | 6 | ✓ | ✓ | ✗ | ✗ | Task completion | ✗ | ✗ |
| ING-VP ( Zhang et al., 2024 ) | 2D | 6 | ✗ | ✓ | ✗ | ✗ | Task completion | ✗ | ✗ |
| V-MAGE ( Zheng et al., 2025 ) | 2D/3D | 5 | ✗ | ✓ | ✗ | ✗ | ELO ranking | ✗ | ✗ |
| itch.io | PyWeek | |||||
| Model | History | No History | Partial+ | History | No History | Partial+ |
| Vision-Language Models (VLMs) | ||||||
| Qwen3-Omni-30B-A3B | 0.62 | 0.84 | 19.5% | 0.89 | 1.26 | 37.0% |
| Qwen2.5-VL-7B | 0.41 | 0.82 | 17.7% | 0.67 | 1.18 | 31.7% |
| Gemma-4-26B-A4B-it | 1.02 | 0.71 | 17.3% | 1.37 | 0.96 | 24.8% |
| Qwen3-VL-8B | 0.64 | 0.70 | 15.5% | 0.73 | 1.07 | 28.3% |
| Genre | Mean | Partial+ | |
| Simulation | 506 | 0.893 | 27.1% |
| Strategy | 223 | 0.858 | 25.2% |
| Shooter | 205 | 0.771 | 20.4% |
| Survival | 171 | 0.770 | 21.7% |
| Racing | 72 | 0.705 | 18.9% |
| Action | 958 | 0.620 | 14.4% |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Harmful Categories | Game Count |
| Firearms Individual Violence | 13 |
| Scheming Manipulation | 35 |
| Value Misalignment | 91 |
| Inconsistent Behavior | 24 |
| Explosives Violence | 3 |
| Economic Waste | 12 |
| Reason | Games |
| Construct runtime freeze failure | 225 |
| Dragging is the only listed control | 72 |
| Requires multiple players | 25 |
| Not an interactive game | 18 |
| Gamepad required | 2 |
| Original itch.io genres | PlaySuite genre |
| Puzzle, Educational | Puzzle |
| Adventure, Visual Novel, Interactive Fiction, Role Playing | Adventure |
| Action, Platformer, Rhythm | Action |
| Simulation, Sports | Simulation |
| Strategy, Card Game | Strategy |
| Shooter | Shooter |
| Model | Family | Parameters | Hugging Face identifier | Reference |
| Qwen3-Omni-30B-A3B-Instruct | VLM | 30B (3B active) | Qwen/Qwen3-Omni-30B-A3B-Instruct | ( Xu et al., 2025b ) |
| Qwen2.5-VL-7B-Instruct | VLM | 7B | Qwen/Qwen2.5-VL-7B-Instruct | ( Bai et al., 2025b ) |
| Gemma-4-26B-A4B-it | VLM | 26B (4B active) | google/gemma-4-26B-A4B-it | ( Gemma Team, 2026 ) |
| Qwen3-VL-8B-Instruct | VLM | 8B | Qwen/Qwen3-VL-8B-Instruct | ( Bai et al., 2025a ) |
| Qwen2.5-Omni-7B | VLM | 7B | Qwen/Qwen2.5-Omni-7B | ( Xu et al., 2025a ) |
| Gemma-4-E4B-it | VLM | 8B (4.5B eff.) | google/gemma-4-E4B-it | ( Gemma Team, 2026 ) |
| Model | Parser | Prompt template |
| Gemma-4-26B-A4B-it | qwen3 | standard |
| EvoCUA-8B | computer_use | computer_use |
| Qwen3-Omni-30B-A3B | qwen3 | standard |
| Gemma-4-E4B-it | qwen3 | standard |
| UI-TARS-1.5-7B | uitars | uitars |
| Qwen3-VL-8B | qwen3 | standard |
| Score | Level | Condition |
| 4 | Completed | win_or_completion_screen |
| 3 | Substantial | checkpoint_or_objective_completed or ( new_area_or_level_reached and score_or_resource_increased ) |
| 2 | Partial | new_area_or_level_reached or score_or_resource_increased |
| 1 | Minimal | player_control_evident and distinct_interactions |
| 0 | None | No condition above is satisfied. |
| Judge | Milestone agreement | Within one |
| Qwen3.5-9B | 84.4% | 90.4% |
| Qwen3.5-27B | 83.9% | 91.6% |
| Gemma-4-31B-it | 82.5% | 77.3% |
| Molmo2-8B | 77.9% | 81.2% |
| itch.io | PyWeek | |||||
| Genre | ||||||
| Simulation | 506 | 0.893 | 0.800 | 13 | 1.293 | 1.173 |
| Strategy | 223 | 0.858 | 0.704 | 10 | 1.054 | 0.702 |
| Shooter | 205 | 0.771 | 0.682 | 6 | 0.763 | 0.720 |
| Survival | 171 | 0.770 | 0.644 | 5 | 1.174 | 0.857 |
| Racing | 72 | 0.705 | 0.654 | 3 | 0.976 | 0.714 |