Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena
Organizations: HKUST (Guangzhou) · Knowin AI · The Chinese University of Hong Kong
Abstract
Frontier vision-language models (VLMs) combine scene estimation, interaction grounding, and executable actions. Understanding how these abilities support complete robotic tasks is central to evaluating their readiness as robot generalists. We introduce Embodied Agent Arena to examine where local competence supports, or falls short of, complete task success across Geometry, Spatial Reasoning, Affordance, Task Planning, and Manipulation. The arena contains 1,000 cases drawn from 32 established sources and GeoProbe, our new benchmark for geometric estimation on Blender renders and real-scene images. A minimal harness preserves source observations and operations while separating metric precision, functional grounding, and native goal completion. We evaluate seven VLMs, analyze Astra's task-specific advantages, and compare richer-observation execution protocols and multi-round review. Across the arena, Astra's advantage is strongest in precise estimation and usable-contact localization; completing coordinated, goal-directed actions remains the key gap to robot generalism.
Figures & tables
| Evaluation coverage | Diagnostic analysis | ||||||
| Study / benchmark | Fixed probes | Region probes | Online nav. | Robot manip. | State memory | Calib./ pose | Task forms |
| BEAR ( 2025 ) | ✓ | ✓ | ✗ | ✓ ∗ | ✗ | ✗ | ✓ |
| EmbodiedBench ( 2025a ) | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ |
| EmbodiedEval ( 2025 ) | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ |
| Embodied Arena ( 2025 ) | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ |
| Code-based control ( 2026 ) | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ |
| Geometry | Spatial Reasoning | Affordance | Task Planning | Manipulation | ||||
| Agent | Rot. | Trans. | Pass@1 | AbsRel | Mask IoU | Point | Pass@1 | Task SR |
| ( ∘ ) | (cm) | (%) | (%) | (%) | (%) | |||
| Astra | 17.1 | 106.6 | 0.213 | 0.286 | 80.3 | 41.8 | ||
| Sol | 27.0 | 151.9 | 0.273 | 0.264 | 44.6 | 19.8 | ||
| Fable | 25.2 | 121.7 | 0.186 | 0.284 | 63.7 | 24.7 | ||
| Qwen-Max | 20.5 | 138.1 | 0.205 | 0.272 | 35.7 | 9.3 | ||
| (a) | Geometry | Spatial Reasoning | Affordance | ||
|---|---|---|---|---|---|
| Model | Condition | I/T/W | Acc. (%) | Box IoU | SR 50 (%) |
| Astra | Initial | – | 72.2 | 0.2903 | 22.2 |
| Self-review | 3/20/1 | 66.7 | 0.2914 | 22.2 | |
| Crop review | 4/19/1 | 66.7 | 0.2912 | 22.2 | |
| Tool-assisted | 5/15/4 | 72.2 | 0.3174 | 33.3 | |
| Qwen-27B | Initial | – | 55.6 | 0.0336 | 0.0 |
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Cases | Share (%) | Task coverage |
|---|---|---|---|
| Geometry | 370 | 37.0 | Calibration, pose, scale, tracing |
| Spatial Reasoning | 220 | 22.0 | Relations, views, video |
| Affordance | 70 | 7.0 | Functional parts, contact |
| Task Planning | 157 | 15.7 | Household, navigation, science |
| Manipulation | 183 | 18.3 | Robot operations, sequences |
| Total | 1,000 | 100.0 | 660 fixed-input; 340 interactive |
| Diagnostic | Cases | Diagnostic | Cases |
|---|---|---|---|
| Object displacement, fixed camera | 12 | Full camera intrinsics | 24 |
| Object displacement, small view change | 12 | Focal length, centered principal point | 12 |
| Object displacement, large view change | 12 | Relative camera pose | 24 |
| Real tabletop displacement | 4 | Sparse depth | 12 |
| Object dimensions | 12 | Reference-based scale | 44 |
| Benchmark | Cases | Task coverage |
|---|---|---|
| CALVIN ( Mees et al., 2021 ) | 10 | Ten sequences, each with five ordered subgoals |
| CLIPort ( Shridhar et al., 2021 ) | 18 | Eighteen task and color-generalization configurations |
| RLBench ( James et al., 2019 ) | 10 | Ten task configurations |
| RoboCasa ( Nasiriany et al., 2024 ) | 10 | Seven atomic and three composite tasks |
| RoboCasa365 | 10 | Three atomic and seven composite tasks |
| RoboWits | 10 | Ten official task configurations |
| Domain | Observation support | Actions and feedback |
|---|---|---|
| Geometry | Image or image pair; intrinsics for Map-free | One final numerical submission |
| Spatial Reasoning | Images, crops, and video-frame queries | Evidence inspection and answer submission |
| Affordance | Image and interaction instruction | One point or bounding-region submission |
| Task Planning | Images, text, or symbolic state | Navigation, interaction, programs, public progress |
| Manipulation | Protocol-specific scene/state observations and native grounding helpers | Motion or skill calls; execution and task checks |
| Model label | Identifier |
|---|---|
| GPT-6 Astra | gpt-6-astra |
| GPT-5.6 Sol | gpt-5.6-sol |
| Claude Fable 5.1 | claude-fable-5-1 |
| Gemini 3.8 Flash | gemini-3.8-flash |
| Qwen 3.8 Max | qwen3.8-max |
| Qwen 3.5 397B-A17B | qwen3.5-397b-a17b |
| Entry | Code rounds | Tokens | Env. steps | Primitives | Episode seconds | Task seconds |
|---|---|---|---|---|---|---|
| Geometry | 12 | 120k | – | – | – | 960 |
| Spatial: non-VSI | 6 | 96k | – | – | 900 | 960 |
| Spatial: VSI | 8 | 240k | – | – | 1,200 | 1,500 |
| ALFRED, ALFWorld | 8 | 320k | 120 | 540 | 1,200 | 1,500 |
| DiscoveryWorld | 8 | 320k | 100 | 400 | 1,200 | 1,500 |
| HumanCLAW | 101 | 2,000k | 100 | 400 | 10,800 | 11,000 |
| Domain | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B |
|---|---|---|---|---|---|---|---|
| Geometry | 1382/1402 | 1382/1402 | 819/886 | 1382/1402 | 1314/1402 | 1382/1402 | 1376/1402 |
| Spatial Reasoning | 1060/1100 | 1100/1100 | 586/660 | 1097/1100 | 1077/1100 | 1033/1100 | 982/1100 |
| Affordance | 350/350 | 350/350 | 208/210 | 350/350 | 335/350 | 349/350 | 349/350 |
| Task Planning | 157/157 | 157/157 | 157/157 | 157/157 | 157/157 | 157/157 | 157/157 |
| Metric | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B | |
|---|---|---|---|---|---|---|---|---|
| Geometry (370 cases) | ||||||||
| InFlux focal MAPE (%) | 26 | 53.02 | 47.79 | 52.65 | 55.50 | 52.36 | 67.82 | 65.86 |
| InFlux principal error (px) | 26 | 3.82 | 3.82 | 3.82 | 3.82 | 3.88 | 3.82 | 3.82 |
| Angle relative L2 | 12 | 0.760 | 0.915 | 1.32 | 1.41 | 0.962 | 1.30 | 2.04 |
| Distance relative L2 | 10 | 0.294 | 0.418 | 0.785 | 0.560 | 0.419 | 0.411 | 0.331 |
| Vector relative L2 | 8 | 0.425 | 0.591 | 0.625 | 0.454 | 0.647 | 1.02 | 1.03 |
| Source | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B | |
|---|---|---|---|---|---|---|---|---|
| Spatial Reasoning: categorical accuracy (%) | ||||||||
| 3DSRBench | 50 | 73.6 2.6 | 71.2 1.8 | 72.7 3.1 | 74.0 3.7 | 70.0 2.4 | 69.2 3.6 | 57.6 2.6 |
| BOP-ASK | 20 | 96.0 4.2 | 87.0 4.5 | 88.3 2.9 | 92.0 2.7 | 97.0 2.7 | 85.0 3.5 | 84.0 4.2 |
| MindCube | 50 | 90.8 1.1 | 68.0 7.3 | 79.3 8.1 | 90.8 2.7 | 89.2 3.0 | 46.0 2.8 | 50.4 4.1 |
| MMSI-Bench | 20 | 77.0 5.7 | 55.0 11.7 | 66.7 7.6 | 55.0 8.7 | 56.0 9.6 | 35.0 9.4 | 29.0 7.4 |
| RoboSpatial | 20 | 85.0 3.5 | 81.0 8.9 | 88.3 2.9 | 88.0 5.7 | 76.0 4.2 | 90.0 3.5 | 92.0 5.7 |
| Agent | Spatial without VSI | Planning without text/symbolic | Affordance without UMD |
|---|---|---|---|
| Astra | 83.6 | 66.7 | 16.5 |
| Sol | 71.4 | 20.6 | 10.5 |
| Fable | 77.9 | 33.3 | 15.0 |
| Qwen-Max | 80.9 | 12.7 | 10.0 |
| Gemini | 78.4 | 22.2 | 5.5 |
| Qwen-397B | 62.2 | 7.9 | 2.0 |
| Task | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B |
|---|---|---|---|---|---|---|---|
| GP focal-only | |||||||
| InFlux focal | |||||||
| GP depth | |||||||
| GP object size | |||||||
| GP ref. height | |||||||
| GP height ratio |
| Task | Cases | Rounds | Peer | Astra | Peer |
|---|---|---|---|---|---|
| GP focal-only | 12 | 5 | Sol | ||
| InFlux focal | 26 | 5 | Sol | ||
| GP depth | 9 | 3 | Fable | ||
| GP object size | 12 | 3 | Fable | ||
| GP ref. height | 12 | 5 | Sol | ||
| GP height ratio | 12 | 5 | Sol |
| Condition | Cases | Astra | Sol |
|---|---|---|---|
| Blender: fixed camera | 12 | ||
| Blender: fixed, moving-only | 10 | ||
| Blender: small view change | 12 | ||
| Blender: large view change | 12 | ||
| Real photographs | 4 |
| Motion | Cases | Rotation ( ∘ ) | Translation (cm) | Joint success (%) |
|---|---|---|---|---|
| Stationary | 4 | 0.00 | 0.00 | |
| Rotation only | 6 | 0.35 | 1.00 | |
| Translation only | 6 | 0.00 | 1.55 | |
| Rotation + translation | 8 | 3.26 | 14.16 |
| Task | Peer | Rounds | Astra | Peer | Gap SD |
|---|---|---|---|---|---|
| MMSI camera–region | Fable | 3 | 100.0 | 77.8 | |
| VSI medium direction | Qwen-Max | 5 | 100.0 | 83.3 | |
| BOP-ASK left/right | Gemini | 5 | 95.0 | 96.7 | |
| MindCube camera shift | Qwen-Max | 5 | 78.5 | 89.2 | |
| 3DSR object facing | Fable | 3 | 50.0 | 72.2 | |
| RoboSpatial configuration | Q-397B | 5 | 78.0 | 100.0 |
| 3DSRBench 50 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B World height 4 Above / below 7 Camera depth 10 Next to 1 Object distance 2 Parallel / perpendicular 1 Same/different facing 4 Object-facing side 6 In front of actor 4 Actor left/right 7 Object viewpoint 4 |
| BOP-ASK 20 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B Closer to camera 5 Farther from camera 3 Left / right 12 |
| MindCube 50 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B Behind observer 7 Camera translation 13 Imagined observer 2 Object relations 12 Composed rotation 5 Turn then translate 11 |
| MMSI 20 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B Attribute (Appr.) 2 Attribute (Meas.) 2 MSR 2 Motion (Cam.) 1 Motion (Obj.) 1 Camera–camera relation 3 Camera–object relation 1 Camera–region relation 3 Object–object relation 1 Object–region relation 4 |
| RoboSpatial 30 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B Placement compatibility 10 Configuration 10 Free-space point set 10 |
| VSI-Bench 50 cases Task Astra Sol Fable Q-Max Gemini Q-397B Q-27B Appearance order 2 Absolute distance 9 Object count 5 Direction: easy 4 Direction: hard 4 Direction: medium 6 Relative distance 6 Object size 9 Room area 4 Route planning 1 |
| Task | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B |
|---|---|---|---|---|---|---|---|
| Object count | |||||||
| Absolute distance | |||||||
| Object size | |||||||
| Room area |
| Task | Cases | Valid / assigned | Common cases | Common AbsRel |
|---|---|---|---|---|
| Object count | 5 | 25/25 | 5 | |
| Absolute distance | 9 | 25/45 | 2 | |
| Object size | 9 | 42/45 | 8 | |
| Room area | 4 | 20/20 | 4 |
| Action / hands | Cases | Mask IoU | Point (%) | Box (%) |
|---|---|---|---|---|
| RAGNet-3DOI | ||||
| free movement / two hands | 2 | |||
| free / one hand | 1 | |||
| free / two hands | 5 | |||
| pull / one hand | 12 | |||
| ReasonAff-style | ||||
| Task / criterion | Astra | Best peer | Model |
|---|---|---|---|
| UMD cutting: mask IoU | Fable | ||
| UMD cutting: strict box (%) | Fable | ||
| UMD containment: mask IoU | Sol | ||
| UMD containment: strict box (%) | Qwen-Max | ||
| ReasonAff pulling: point (%) | Fable | ||
| RAGNet pulling: point (%) | Qwen-Max |
| Goal family | Astra | Best peer | Model |
|---|---|---|---|
| ALFRED | |||
| Illumination | 1/2 | 1/2 | Fable ∗ |
| Placement | 4/4 | 2/4 | Sol ∗ |
| Container transport | 2/2 | 1/2 | Fable ∗ |
| Cleaning | 2/2 | 1/2 | Sol ∗ |
| Cooling | 2/4 | 2/4 | Fable |
| Model | Process coverage | Target found | Within 20 cm | Near + sit issued | Success |
| Astra | 8/8 | 6 | 5 | 5 | 0 |
| Sol | 8/8 | 3 | 1 | 1 | 0 |
| Fable | 8/8 | 0 | 0 | 0 | 0 |
| Qwen-Max | 0/8 | – | – | – | 0 |
| Gemini | 6/8 | 3 † | 3 † | 3 † | 0 |
| Q-397B | 0/8 | – | – | – | 0 |
| Benchmark | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B | |
|---|---|---|---|---|---|---|---|---|
| Baseline | ||||||||
| CALVIN | 10 | 30.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| CLIPort | 18 | 61.1 | 27.8 | 33.3 | 22.2 | 27.8 | 22.2 | 16.7 |
| RLBench | 10 | 60.0 | 20.0 | 10.0 | 0.0 | 20.0 | 0.0 | 10.0 |
| RoboCasa | 10 | 10.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| RoboCasa365 | 10 | 0.0 | 0.0 | 10.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Benchmark | Astra | Sol | Fable | Q-Max | Gemini | Q-397B | Q-27B | |
|---|---|---|---|---|---|---|---|---|
| Baseline | ||||||||
| CALVIN goals (0–5) | 10 | 3.10 | 1.40 | 0.90 | 0.80 | 1.80 | 0.00 | 0.10 |
| RoboCasa | 4 | – | – | – | – | – | – | – |
| RoboCasa365 | 7 | – | – | – | – | – | – | – |
| ManiSkill conditions | 5 | 60.00 | 0.00 | – | 0.00 | 0.00 | 0.00 | 0.00 |
| LIBERO-PRO | 10 | 20.00 | 20.00 | 30.00 | 5.00 | 20.00 | 0.00 | 5.00 |
| Source | Scoring items | Evidence window | |
|---|---|---|---|
| RoboCasa | 4 | Approach, grasp and lift, stable release | Any verified event during execution |
| RoboCasa365 | 7 | Same three transport milestones | Any verified event during execution |
| ManiSkill | 5 | Two or three native conditions per task | Final verification state |
| LIBERO-PRO | 10 | Two task-specific conditions | Final state and limited process records |
| robosuite | 5 | Two lifting, three stacking, or four wiping items | Native success and terminal reward-derived evidence |
| BEHAVIOR-1K | 5 | Six control-evidence items | Before the first episode-end signal |
| Task | Scored conditions | |
|---|---|---|
| PickCube | 2 | Object placed at its goal; placed and robot static. The suspended goal requires no release. |
| PlaceSphere | 3 | Sphere on its support; on support and released; both with the sphere static. |
| StackCube | 3 | Cube A on B; on B and released; both with cube A static. |
| PlugCharger | 2 | Goal distance m; distance satisfied and angular error rad. |
| PegInsertionSide | 2 | Peg-head m in the hole frame; official insertion success. |
| Item | Positive evidence |
|---|---|
| Localization | Finite coordinates or pose from public RGB-D localization or grasp sampling. |
| Candidate | A nonempty pregrasp or grasp candidate set. |
| Arrival | Measured arrival from approach or recovery feedback for the same arm. |
| Holding | An explicit in-hand check or holding confirmation inside checked placement. |
| Held motion | Associated measured motion of at least 5 cm while holding, or movement with a passed holding check inside checked placement. |
| Release | Holding, active gripper opening, then no longer holding; or confirmed release inside checked placement. |
| Source | Rounds | Tools | Checks | Tokens | Trial | Code | Native |
|---|---|---|---|---|---|---|---|
| CALVIN | 40 | 7,000 | 41 | 1,280 | 120 | 10 | 10 |
| CLIPort | 24 | 256 | 25 | 960 | 60 | 10 | 10 |
| RLBench | 16 | 256 | 17 | 640 | 60 | 10 | 10 |
| RoboCasa | 24 | 256 | 25 | 1,120 | 60 | 6 | 10 |
| RoboCasa365 | 24 | 256 | 25 | 1,120 | 60 | 8 | 10 |
| RoboTwin 2.0 | 16 | 256 | 17 | 640 | 60 | 10 | 20 |
| Source | Views | Resolution |
|---|---|---|
| CALVIN | Static; gripper | ; |
| CLIPort | Front, left, right; no Oracle camera | |
| RLBench | Front, wrist, overhead | |
| RoboCasa/365 | Left/right agent view, eye-in-hand | |
| RoboTwin 2.0 | Head, left, right | |
| RoboWits | Ego, left/right wrist |
| Source | Base effective | Counter and termination semantics |
|---|---|---|
| CLIPort | Pick-and-place actions | |
| RoboCasa/365 | Control steps; field does not truncate episodes in the current wrapper | |
| RoboTwin 2.0 | Native action calls, not all simulator steps | |
| RoboWits | Control steps for the checked case | |
| ManiSkill | Control steps; PickCube termination/boundary checks verify the effective limit | |
| LIBERO-PRO | Wrapper/control counter; field does not truncate episodes |
| Condition | Evidence available during review | Additional response budget |
|---|---|---|
| Initial | Original task observations | None |
| Self-review | Original observations and initial answer | Two responses, up to 1,800 tokens each |
| Crop review | Self-review evidence plus image crops | Two responses, up to 1,800 tokens each |
| Tool-assisted | Original RGB, crops, and specialist measurements | Six responses, up to 4,096 tokens each; up to 16 specialist requests |