cs.ROOct 1, 2026

Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena

Authors: Haojian Huang, Pukun Zhao, Zexi Li, Yehang Zhang, Yangkai Wei, Wenqian Li, Han Yang, Kaiwen Zhou, +2 more

Organizations: HKUST (Guangzhou) · Knowin AI · The Chinese University of Hong Kong

Abstract

Frontier vision-language models (VLMs) combine scene estimation, interaction grounding, and executable actions. Understanding how these abilities support complete robotic tasks is central to evaluating their readiness as robot generalists. We introduce Embodied Agent Arena to examine where local competence supports, or falls short of, complete task success across Geometry, Spatial Reasoning, Affordance, Task Planning, and Manipulation. The arena contains 1,000 cases drawn from 32 established sources and GeoProbe, our new benchmark for geometric estimation on Blender renders and real-scene images. A minimal harness preserves source observations and operations while separating metric precision, functional grounding, and native goal completion. We evaluate seven VLMs, analyze Astra's task-specific advantages, and compare richer-observation execution protocols and multi-round review. Across the arena, Astra's advantage is strongest in precise estimation and usable-contact localization; completing coordinated, goal-directed actions remains the key gap to robot generalism.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

    Jun 28, 2026Taozhao Chen, Ian Manchester, Huaming ChenSemantic RepresentationsPosition

  2. VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

    Dec 27, 2025Borong Zhang, Jiahao Li, Jiachen Shen +7Diffusion-Based Vision-Language-ActionsRobot Policies

  3. Towards Generalizable Visually Grounded Exploration of Household Devices

    Sep 1, 2026Linhao Zheng, Zeming Liu, Wangke Chen +4Embodied AgentsWeak Visual Grounding