Organizations: Ludwig Maximilian University of Munich · Munich Center for Machine Learning · East China University of Science and Technology · National University of Singapore · Amazon
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and State Input, which we formulate as a Conditional Generative Markov Chain. Based on this formulation, we propose MAE (Markov Attention Entropy), a self-evaluation framework that directly converts internal attention signals into architecture-aware reliability scores, and introduce LIBERO-Reflect, a 4,000-episode benchmark combining 2,000 standard episodes and 2,000 challenging episodes across four subsets. Extensive experiments across heterogeneous VLA architectures and diverse scenarios show that MAE consistently outperforms state-of-the-art baselines on AUPR, AUROC, and FPR@95.
Figures & tables
Figure 1 : From multimodal conditioning to action generation in VLAs. We view VLA action generation as a conditional generative Markov chain: visual observations, language instructions, and State Input jointly condition a sequence of latent action states that evolves toward final action. Attention exposes how these latent states route information across modalities, and its entropy serves as a white-box signal for self-evaluating action-generation reliability.
Figure 2 : Attention entropy as an internal self-evaluation signal. Panels A–B show a successful visual query concentrated on the task-relevant object and a failed visual query dispersed across distractor regions. Panels C–D compare oriented episode-level score distributions from visual and text attention entropy on the same Reflect-Goal episodes. Visual entropy yields a much clearer separation between successful and failed episodes than text entropy, supporting visual attention entropy as the action-relevant self-evaluation signal.
Figure 3 : Unified Conditional Generative Markov Chain view of heterogeneous VLAs. Visual observations, language instructions, and State Input form the conditioning context, while latent action states evolve through architecture-specific transition kernels before the executable action is produced. MAE reads attention entropy from this internal transition process and converts it into architecture-aware self-evaluation scores.
Model
Method
Goal Semantics Reflect-Goal
Object Binding Reflect-Object
Spatial Grounding Reflect-Spatial
Composite Reflect-10
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
❶ Latent-Readout VLAs
OpenVLA
Black-box baselines
Random
48.55
39.27
95.36
47.85
35.98
95.87
48.97
39.49
94.82
52.51
28.88
93.92
Verbal. Conf
50.36
40.82
92.91
49.74
37.21
93.64
50.88
41.03
92.47
53.69
30.04
91.86
Self-Cons. †
56.82
41.28
84.90
58.41
42.76
82.35
57.94
44.15
87.80
53.64
27.20
90.64
Table 1: Main self-evaluation results on LIBERO-Reflect . Applicable baselines are grouped as black-box or white-box methods. Purple rules highlight MAE rows, which report the score and relative change against Random ; for FPR@95 , the relative value is the reduction rate. Bold Top-1 values and underlined half-head values improve over Random within the same model, subset, and metric. † Token-statistic and Self-Consistency baselines are reported only for OpenVLA , whose autoregressive discrete action interface exposes action-token probabilities and sampled action-token sequences. Additional applicability details are provided in Appendix F .
Figure 4 : Efficiency–reliability Pareto analysis and layer-band ablation for MAE . The left panel reports added evaluation latency normalized by one rollout wall time. The right panel reports absolute AUROC for MAE-D with Top-1 ; colors are normalized within each column by closeness to the all-layer result, stars mark the best layer band, and the left schematic indicates the layer region.
Figure 5 : Text-head versus visual-head self-evaluation signals. OpenVLA uses MAE-D , and QwenPI-Flow uses MAE-C . Each connector reports Δ AUROC = V - T. Positive gaps across all subsets indicate that text-side entropy is not a stable substitute for visual information.
Figure 6 : Top-m head selection ablation. Panels order selected-head settings from many heads to Top-1 and use independent y-axis scales. Values are percentage-point changes from the Top-1 ranking score, computed as the mean of AUROC and AUPR over the four LIBERO-Reflect subsets. Exact values are reported in Appendix D .
Benchmark subset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
Reflect-Goal Goal Semantics
Random
47.75±0.81
48.12±0.79
95.98±0.63
MAE-C Top-1
80.72±0.79
80.14±0.78
60.15±1.28
Reflect-Object Object Binding
Random
48.91±0.79
48.56±0.79
95.18±0.64
MAE-C Top-1
76.12±0.85
76.64±0.83
81.13±1.27
Reflect-Spatial Spatial Grounding
Random
48.30±0.86
50.31±0.86
94.90±0.62
MAE-C Top-1
84.95±0.77
85.60±0.77
68.46±1.32
Table 2 : Run-to-run stability of QwenPI-Flow with MAE-C Top-1 across all four subsets. Values are percentages reported as sample mean ± sample standard deviation over five independent simulator seeds.
Figure 7 : Pooled standard-only balanced results for OpenVLA . Bar heights report improvement over the balanced random expectation: score minus 50 for AUROC / AUPR , and 95 minus score for FPR@95 . Labels above bars give the original metric values. Purple callouts report the gain of MAE-D Top-1 over Length-normalized Entropy.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Policy
Action-Generation Family
VLM and Perception Stack
Training / Adaptation Signal
Diversity Axis
OpenVLA
Latent-Readout VLAs
OpenVLA-7B VLM with prism-dinosiglip-224px; fused DINOv2 + SigLIP visual encoder; Llama 2 7B language backbone
official openvla/openvla-7b; tokenized action prediction trained on Open X-Embodiment
OpenVLA-7B VLM with the same DINOv2 + SigLIP visual base and Llama 2 7B language backbone
official OpenVLA-OFT checkpoints/code; efficient OFT adaptation with continuous actions, action chunking, and L1 regression
Same backbone family as OpenVLA with a continuous readout recipe; tests adaptation robustness
QwenPI-Flow
Latent-Refinement VLAs
Qwen3-VL-4B-Instruct with Qwen3-VL vision stack plus dinov2_vits14 in the StarVLA config
StarVLA/Qwen3-VL-PI-LIBERO-4in1; QwenPI policy with DiT-B, 7-DoF actions, horizon 8, and data_mix=libero_all
Qwen3-VL family; flow-style refinement; LIBERO-specific training mix; tests cross-family and cross-generation generality
Appendix
Table 3: Detailed configurations for the three VLA policies evaluated in the main experiments. The table highlights architectural, family-level, and training-source diversity.
Component
Goal Semantics Reflect-Goal
Spatial Grounding Reflect-Spatial
Composite Reflect-10
Object Binding Reflect-Object
Overall
Task count
10
10
10
10
40
Standard episodes
500
500
500
500
2,000
Challenging episodes
500
500
500
500
2,000
Total episodes
1,000
1,000
1,000
1,000
4,000
❶ Standard-side policy success rate (%)
QwenPI-Flow
96.8
93.8
94.4
98.0
95.8
Appendix
Table 4: LIBERO-Reflect composition and standard-side policy success rates. The upper block reports the number of tasks and episodes used to form the benchmark. The lower block reports policy success rates on the standard side. Overall success is the mean over four equally sized suites. QwenPI-Flow corresponds to the Qwen3-VL-PI-LIBERO-4in1 checkpoint described in Table 3 .
Policy
Goal Semantics Reflect-Goal
Spatial Grounding Reflect-Spatial
Composite Reflect-10
Object Binding Reflect-Object
Mean
QwenPI-Flow
0.0
7.2
1.8
0.0
2.25
OpenVLA-OFT
4.6
7.2
0.0
1.2
3.25
OpenVLA
0.0
0.4
0.0
0.0
0.10
Appendix
Table 5: Diagnostic success rates for the challenging episodes sampled from LIBERO-PRO and retained in LIBERO-Reflect . Values are percentages and summarize the challenging side used by the benchmark.
Figure 8 : Case studies from the LIBERO-Reflect construction. Each row corresponds to one subset and shows representative standard episodes paired with challenging episodes. The instruction is printed below each panel, yielding 16 data points across the four subsets.
Model
Subset
Text Head
Text AUROC
Visual Head
Visual AUROC
Δ AUROC
❶ Latent-Readout VLA
OpenVLA
Reflect-Goal
MAE-D Top-1
59.26
MAE-D Top-1
63.94
+4.68
Reflect-Object
MAE-D Top-1
69.41
MAE-D Top-1
90.97
+21.56
Reflect-Spatial
MAE-D Top-1
62.31
MAE-D Top-1
66.86
+4.55
Reflect-10
MAE-D Top-1
49.43
MAE-D Top-1
54.74
+5.31
❷ Latent-Refinement VLA
Appendix
Table 6: Exact AUROC values for the text-head versus visual-head comparison in the main paper. Purple cells mark the visual-head signal used by MAE , and purple deltas report visual-head MAE minus text-head MAE under the same Top-1 setting. Subset names use the LIBERO-Reflect split identifiers.
Model
Top-m Setting
Goal Semantics Reflect-Goal
Object Binding Reflect-Object
Spatial Grounding Reflect-Spatial
Composite Reflect-10
AUROC
AUPR
FPR @95
AUROC
AUPR
FPR @95
AUROC
AUPR
FPR @95
AUROC
AUPR
FPR @95
❶ Latent-Readout VLAs
OpenVLA
MAE-D Top-1
63.94
43.23
61.92
90.97
75.88
28.30
66.86
50.74
75.79
54.74
30.45
86.46
MAE-D Top-2
62.41
42.33
68.21
89.14
72.76
30.05
64.72
52.29
77.13
54.01
26.97
85.91
MAE-D Top-4
60.36
41.01
70.03
87.13
69.26
34.18
64.25
52.14
78.80
53.20
26.55
86.88
MAE-D Top-16
59.56
40.67
71.52
79.64
58.98
48.81
63.99
49.67
84.81
53.21
26.56
87.43
Appendix
Table 7: Top-m head-selection ablation values for the main paper. Bold values mark the Top-1 setting used in the main results, while underlined values mark the half-head comparison: Top-16 for 32-head OpenVLA-family models and Top-20 for the 40-head QwenPI-Flow model. Top-1 gives the most stable ranking quality while avoiding noisy aggregation over many heads.
Layer Band
Overall All subsets
Composite Reflect-10
Goal Semantics Reflect-Goal
Object Binding Reflect-Object
Spatial Grounding Reflect-Spatial
All layers
67.88
54.74
63.94
90.97
66.86
Shallow
65.82
51.49
57.00
90.43
64.97
Middle
66.88
54.25
61.93
86.59
65.28
Deep
62.91
47.71
61.49
83.50
65.55
Appendix
Table 8: OpenVLA layer-band ablation values for the main paper. Values are MAE-D with Top-1 AUROC percentages. The all-layer setting is the default configuration because it is strongest overall and avoids task-specific layer tuning.
Table 9: Evaluation settings for MAE across the three VLA policies. The main score uses all layers with Top-1 head selection. Larger Top-m settings are reported only for the head-selection ablation; the half-head setting is Top-16 for 32-head OpenVLA-family policies and Top-20 for the 40-head QwenPI-Flow policy.
Policy
Internal generation step k
Final internal step K used by MAE
K in our implementation
OpenVLA
One autoregressive action-token generation step
The last action-token generation step
Number of generated action tokens
OpenVLA-OFT
One continuous-readout step
The continuous-readout step before the action head produces the action chunk
1
QwenPI-Flow
One flow-style refinement step
The last refinement step before the action trajectory is returned
Number of inference refinement steps
Appendix
Table 10: Model-specific meaning of the internal generation step k and the final internal step K used by MAE .
Figure 9 : Prompt used for the Verbal Confidence baseline. The stitched contact sheet is provided as an image input in the same multimodal API request.
Policy
Input convention
Episode payload
Rendered prompt form
OpenVLA
Pure action-prompt format used by the OpenVLA policy.
One RGB observation image and the lower-cased LIBERO task instruction.
In: What action should the robot take to {instruction.lower()}? Out:
OpenVLA-OFT
OpenVLA-family action prompt with the OFT continuous readout.
Primary image, wrist image, proprioceptive State Input, and the lower-cased task label.
In: What action should the robot take to {task_label.lower()}? Out:
QwenPI-Flow
Qwen3-VL multimodal message format followed by the StarVLA grounding request.
One or more image placeholders followed by the LIBERO instruction and object-localization request.
<|im_start|>user <|vision_start|><|image_pad|>< |vision_end|> Your task is {instruction}. To identify the key objects for your task. Locate their bounding boxes in [x1,y1,x2,y2] format. <|im_end|> <|im_start|>assistant
Appendix
Table 11: Model input templates used for policy conditioning. The table records the prompt forms and non-text policy inputs needed to reproduce the action-generation inputs; benchmark construction and experimental results are reported in Appendices C – D .
Category
Method
Reflect-Goal
Reflect-Object
Reflect-Spatial
Reflect-10
Overall ( N=1,116 )
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
Black-box
Random
49.82
50.11
95.12
50.15
49.85
95.30
49.90
50.05
94.90
50.20
49.90
94.80
50.02
49.98
95.02
Verbal Conf.
50.45
50.62
93.20
50.10
50.30
93.80
51.10
51.40
92.10
52.80
52.40
92.50
51.11
51.18
92.90
Self-Cons.
56.12
55.80
85.30
57.80
57.20
83.10
57.30
56.90
88.20
53.20
53.00
91.20
56.11
55.72
86.95
White-box
MaxProb
54.30
54.10
89.10
55.20
54.90
87.30
55.10
54.70
90.40
52.90
52.60
92.00
54.38
54.08
89.70
Perplexity
54.80
54.55
88.40
55.60
55.20
86.80
55.50
55.10
89.80
53.05
52.80
91.60
54.74
54.41
89.15
Appendix
Table 12 : Standard-only balanced evaluation for OpenVLA . Every positive and negative episode comes from the same standard LIBERO environment, and each subset is balanced 1:1 by matching every natural failure with a randomly sampled success from that subset. Values are percentages. Overall reports global ranking over the pooled 558 successes and 558 failures.
Figure 10 : Standard-only balanced evaluation details. a , All 558 natural failures from 2,000 standard LIBERO episodes are paired with 558 subset-matched successes, yielding 1,116 standard-domain episodes at 50% success prevalence. b , Direction-aware percentage-point gains of MAE-D Top-1 over Length-normalized Entropy across each subset and the pooled set; positive values denote higher AUROC / AUPR or lower FPR@95 .
Component
Detail
GPU Hardware
NVIDIA H100
Model Evaluated
QwenPI-Flow
Avg. Rollout Time
∼ 14.0 s / episode
Avg. MAE Extra Time
0.57 s / episode
MAE Latency Overhead
4.09% ( <0.1× )
Memory Overhead
Negligible (Reuses internal attention)
Appendix
Table 13: Empirical cost analysis of MAE . The overhead strictly satisfies the <0.1× boundary highlighted in the Pareto-optimal zone of the main paper.
Vision-Language-Action (VLA) models are increasingly evaluated across multiple simulation benchmarks, yet adding each benchmark to an evaluation pipeline requires resolving incompatible dependencies, matching underspecified evaluation protocols, and reverse-engineering undocumented preprocessing. This burden scales with the number of models and benchmarks, making comprehensive evaluation impractical for most teams. We present vla-eval, an open-source evaluation harness that eliminates this per-benchmark cost by decoupling model inference from benchmark execution through a WebSocket+msgpack protocol with Docker-based environment isolation. Models integrate once by implementing a single predict() method; benchmarks integrate once via a four-method interface; the full cross-evaluation matrix works automatically. The framework supports 14 simulation benchmarks and six model servers. Parallel evaluation via episode sharding and batch inference achieves up to 47x wall-clock speedup, completing 2,000 LIBERO episodes in ~18 minutes. To validate the framework, we reproduce published scores across six VLA codebases and three benchmarks, documenting previously undocumented pitfalls. We additionally release a VLA leaderboard aggregating 657 published results across 17 benchmarks. Framework, evaluation configs, and all reproduction results are publicly available at https://github.com/allenai/vla-evaluation-harness and https://allenai.github.io/vla-evaluation-harness/leaderboard.
Suhwan Choi, Yunsung Lee, Yubeen Park +4
Allen Institute for AI (AI2) · Seoul National University
While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively understand their limits and failure modes. To address this, we introduce a comprehensive benchmark called VLA-Arena. We propose a novel structured task design framework to quantify difficulty across three orthogonal axes: (1) Task Structure, (2) Language Command, and (3) Visual Observation. This allows us to systematically design tasks with fine-grained difficulty levels, enabling a precise measurement of model capability frontiers. For Task Structure, VLA-Arena's 170 tasks are grouped into four dimensions: Safety, Distractor, Extrapolation, and Long Horizon. Each task is designed with three difficulty levels (L0-L2), with fine-tuning performed exclusively on L0 to assess general capability. Orthogonal to this, language (W0-W4) and visual (V0-V4) perturbations can be applied to any task to enable a decoupled analysis of robustness. Our extensive evaluation of state-of-the-art VLAs reveals several critical limitations, including a strong tendency toward memorization over generalization, asymmetric robustness, a lack of consideration for safety constraints, and an inability to compose learned skills for long-horizon tasks. To foster research addressing these challenges and ensure reproducibility, we provide the complete VLA-Arena framework, including an end-to-end toolchain from task definition to automated evaluation and the VLA-Arena-S/M/L datasets for fine-tuning. Our benchmark, data, models, and leaderboard are available at https://vla-arena.github.io.
Borong Zhang, Jiahao Li, Jiachen Shen +7
Institute for Artificial Intelligence, Peking University. · Zhongguancun Academy. · PKU-PsiBot Joint Lab. +2
Vision-Language-Action (VLA) models are increasingly used as generalist robot policies, yet their evaluation still relies largely on static benchmarks that randomly sample task scenes. In high-dimensional embodied spaces, failures are sparse and clustered, so static benchmarking can underestimate robustness risks. We reframe VLA evaluation as an active failure-discovery problem and propose a failure-aware test-generation approach that combines diversity-driven exploration with surrogate models learned from observed executions. The method steers testing toward high-risk yet diverse scene regions. Across four state-of-the-art VLA models, it uncovers substantially more failures (up to +29.7 % over selected baselines) while revealing more diverse failure modes. This mean that, for instance, in the case of GR00T-N1.6, success rate dropped from 64.4% to 34.7%. More broadly, our findings call for a shift in VLA evaluation: from passive measurement on fixed task suites to adaptive, failure-seeking test generation that exposes the structure of model weaknesses before deployment.
Arusa Kanwal, Pablo Valle, Shaukat Ali +1
Mondragon University Mondragon, Spain · Simula Research Laboratory Oslo, Norway