MAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Organizations: Ludwig Maximilian University of Munich · Munich Center for Machine Learning · East China University of Science and Technology · National University of Singapore · Amazon
Abstract
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and State Input, which we formulate as a Conditional Generative Markov Chain. Based on this formulation, we propose MAE (Markov Attention Entropy), a self-evaluation framework that directly converts internal attention signals into architecture-aware reliability scores, and introduce LIBERO-Reflect, a 4,000-episode benchmark combining 2,000 standard episodes and 2,000 challenging episodes across four subsets. Extensive experiments across heterogeneous VLA architectures and diverse scenarios show that MAE consistently outperforms state-of-the-art baselines on AUPR, AUROC, and FPR@95.
Figures & tables
| Model | Method | Goal Semantics Reflect-Goal | Object Binding Reflect-Object | Spatial Grounding Reflect-Spatial | Composite Reflect-10 | ||||||||
| AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | ||
| ❶ Latent-Readout VLAs | |||||||||||||
| OpenVLA | Black-box baselines | ||||||||||||
| Random | 48.55 | 39.27 | 95.36 | 47.85 | 35.98 | 95.87 | 48.97 | 39.49 | 94.82 | 52.51 | 28.88 | 93.92 | |
| Verbal. Conf | 50.36 | 40.82 | 92.91 | 49.74 | 37.21 | 93.64 | 50.88 | 41.03 | 92.47 | 53.69 | 30.04 | 91.86 | |
| Self-Cons. | 56.82 | 41.28 | 84.90 | 58.41 | 42.76 | 82.35 | 57.94 | 44.15 | 87.80 | 53.64 | 27.20 | 90.64 | |
| Benchmark subset | Method | AUROC | AUPR | FPR@95 |
| Reflect-Goal Goal Semantics | Random | |||
| MAE-C Top-1 | ||||
| Reflect-Object Object Binding | Random | |||
| MAE-C Top-1 | ||||
| Reflect-Spatial Spatial Grounding | Random | |||
| MAE-C Top-1 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Policy | Action-Generation Family | VLM and Perception Stack | Training / Adaptation Signal | Diversity Axis |
| OpenVLA | Latent-Readout VLAs | OpenVLA-7B VLM with prism-dinosiglip-224px; fused DINOv2 + SigLIP visual encoder; Llama 2 7B language backbone | official openvla/openvla-7b; tokenized action prediction trained on Open X-Embodiment | Llama/OpenVLA family; large-scale action-token training; tests autoregressive discrete readout behavior |
| OpenVLA-OFT | Latent-Readout VLAs | OpenVLA-7B VLM with the same DINOv2 + SigLIP visual base and Llama 2 7B language backbone | official OpenVLA-OFT checkpoints/code; efficient OFT adaptation with continuous actions, action chunking, and L1 regression | Same backbone family as OpenVLA with a continuous readout recipe; tests adaptation robustness |
| QwenPI-Flow | Latent-Refinement VLAs | Qwen3-VL-4B-Instruct with Qwen3-VL vision stack plus dinov2_vits14 in the StarVLA config | StarVLA/Qwen3-VL-PI-LIBERO-4in1; QwenPI policy with DiT-B, 7-DoF actions, horizon 8, and data_mix=libero_all | Qwen3-VL family; flow-style refinement; LIBERO-specific training mix; tests cross-family and cross-generation generality |
| Component | Goal Semantics Reflect-Goal | Spatial Grounding Reflect-Spatial | Composite Reflect-10 | Object Binding Reflect-Object | Overall |
| Task count | 10 | 10 | 10 | 10 | 40 |
| Standard episodes | 500 | 500 | 500 | 500 | 2,000 |
| Challenging episodes | 500 | 500 | 500 | 500 | 2,000 |
| Total episodes | 1,000 | 1,000 | 1,000 | 1,000 | 4,000 |
| ❶ Standard-side policy success rate (%) | |||||
| QwenPI-Flow | 96.8 | 93.8 | 94.4 | 98.0 | 95.8 |
| Policy | Goal Semantics Reflect-Goal | Spatial Grounding Reflect-Spatial | Composite Reflect-10 | Object Binding Reflect-Object | Mean |
| QwenPI-Flow | 0.0 | 7.2 | 1.8 | 0.0 | 2.25 |
| OpenVLA-OFT | 4.6 | 7.2 | 0.0 | 1.2 | 3.25 |
| OpenVLA | 0.0 | 0.4 | 0.0 | 0.0 | 0.10 |
| Model | Subset | Text Head | Text AUROC | Visual Head | Visual AUROC | AUROC |
| ❶ Latent-Readout VLA | ||||||
| OpenVLA | Reflect-Goal | MAE-D Top-1 | 59.26 | MAE-D Top-1 | 63.94 | +4.68 |
| Reflect-Object | MAE-D Top-1 | 69.41 | MAE-D Top-1 | 90.97 | +21.56 | |
| Reflect-Spatial | MAE-D Top-1 | 62.31 | MAE-D Top-1 | 66.86 | +4.55 | |
| Reflect-10 | MAE-D Top-1 | 49.43 | MAE-D Top-1 | 54.74 | +5.31 | |
| ❷ Latent-Refinement VLA | ||||||
| Model | Top-m Setting | Goal Semantics Reflect-Goal | Object Binding Reflect-Object | Spatial Grounding Reflect-Spatial | Composite Reflect-10 | ||||||||
| AUROC | AUPR | FPR @95 | AUROC | AUPR | FPR @95 | AUROC | AUPR | FPR @95 | AUROC | AUPR | FPR @95 | ||
| ❶ Latent-Readout VLAs | |||||||||||||
| OpenVLA | MAE-D Top-1 | 63.94 | 43.23 | 61.92 | 90.97 | 75.88 | 28.30 | 66.86 | 50.74 | 75.79 | 54.74 | 30.45 | 86.46 |
| MAE-D Top-2 | 62.41 | 42.33 | 68.21 | 89.14 | 72.76 | 30.05 | 64.72 | 52.29 | 77.13 | 54.01 | 26.97 | 85.91 | |
| MAE-D Top-4 | 60.36 | 41.01 | 70.03 | 87.13 | 69.26 | 34.18 | 64.25 | 52.14 | 78.80 | 53.20 | 26.55 | 86.88 | |
| MAE-D Top-16 | 59.56 | 40.67 | 71.52 | 79.64 | 58.98 | 48.81 | 63.99 | 49.67 | 84.81 | 53.21 | 26.56 | 87.43 | |
| Layer Band | Overall All subsets | Composite Reflect-10 | Goal Semantics Reflect-Goal | Object Binding Reflect-Object | Spatial Grounding Reflect-Spatial |
| All layers | 67.88 | 54.74 | 63.94 | 90.97 | 66.86 |
| Shallow | 65.82 | 51.49 | 57.00 | 90.43 | 64.97 |
| Middle | 66.88 | 54.25 | 61.93 | 86.59 | 65.28 |
| Deep | 62.91 | 47.71 | 61.49 | 83.50 | 65.55 |
| Policy | Action-generation family | Attention depth | Main score | Head-selection settings | Evaluation role |
| OpenVLA | Latent-Readout VLAs | 32 layers / 32 heads | MAE-D Top-1 | Top-1 , Top-2 , Top-4 , Top-16 | Autoregressive discrete readout; supports token-statistic baselines |
| OpenVLA-OFT | Latent-Readout VLAs | 32 layers / 32 heads | MAE-D Top-1 | Top-1 , Top-2 , Top-4 , Top-16 | Same VLA family with continuous readout |
| QwenPI-Flow | Latent-Refinement VLAs | 36 layers / 40 heads | MAE-C Top-1 | Top-1 , Top-2 , Top-4 , Top-20 | Cross-family flow-style refinement policy |
| Policy | Internal generation step | Final internal step used by MAE | in our implementation |
| OpenVLA | One autoregressive action-token generation step | The last action-token generation step | Number of generated action tokens |
| OpenVLA-OFT | One continuous-readout step | The continuous-readout step before the action head produces the action chunk | |
| QwenPI-Flow | One flow-style refinement step | The last refinement step before the action trajectory is returned | Number of inference refinement steps |
| Policy | Input convention | Episode payload | Rendered prompt form |
| OpenVLA | Pure action-prompt format used by the OpenVLA policy. | One RGB observation image and the lower-cased LIBERO task instruction. | In: What action should the robot take to {instruction.lower()}? Out: |
| OpenVLA-OFT | OpenVLA-family action prompt with the OFT continuous readout. | Primary image, wrist image, proprioceptive State Input, and the lower-cased task label. | In: What action should the robot take to {task_label.lower()}? Out: |
| QwenPI-Flow | Qwen3-VL multimodal message format followed by the StarVLA grounding request. | One or more image placeholders followed by the LIBERO instruction and object-localization request. | <|im_start|>user <|vision_start|><|image_pad|>< |vision_end|> Your task is {instruction}. To identify the key objects for your task. Locate their bounding boxes in [x1,y1,x2,y2] format. <|im_end|> <|im_start|>assistant |
| Category | Method | Reflect-Goal | Reflect-Object | Reflect-Spatial | Reflect-10 | Overall ( ) | ||||||||||
| AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | AUROC | AUPR | FPR@95 | ||
| Black-box | Random | 49.82 | 50.11 | 95.12 | 50.15 | 49.85 | 95.30 | 49.90 | 50.05 | 94.90 | 50.20 | 49.90 | 94.80 | 50.02 | 49.98 | 95.02 |
| Verbal Conf. | 50.45 | 50.62 | 93.20 | 50.10 | 50.30 | 93.80 | 51.10 | 51.40 | 92.10 | 52.80 | 52.40 | 92.50 | 51.11 | 51.18 | 92.90 | |
| Self-Cons. | 56.12 | 55.80 | 85.30 | 57.80 | 57.20 | 83.10 | 57.30 | 56.90 | 88.20 | 53.20 | 53.00 | 91.20 | 56.11 | 55.72 | 86.95 | |
| White-box | MaxProb | 54.30 | 54.10 | 89.10 | 55.20 | 54.90 | 87.30 | 55.10 | 54.70 | 90.40 | 52.90 | 52.60 | 92.00 | 54.38 | 54.08 | 89.70 |
| Perplexity | 54.80 | 54.55 | 88.40 | 55.60 | 55.20 | 86.80 | 55.50 | 55.10 | 89.80 | 53.05 | 52.80 | 91.60 | 54.74 | 54.41 | 89.15 | |
| Component | Detail |
| GPU Hardware | NVIDIA H100 |
| Model Evaluated | QwenPI-Flow |
| Avg. Rollout Time | 14.0 s / episode |
| Avg. MAE Extra Time | 0.57 s / episode |
| MAE Latency Overhead | 4.09% ( ) |
| Memory Overhead | Negligible (Reuses internal attention) |