Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.
Figures & tables
Figure 1: Evaluation framework. (a) Target state and robot context co-vary in natural observations. (b) Crossed comparisons vary state at fixed context (orange row) and context at fixed state (blue column). (c) The same fixed representation–readout pair is evaluated for accuracy, responsiveness, and stability. A separate fixed policy receives observations, and its outcomes are related to the diagnostics (dashed connection). Readout predictions are not policy inputs. Scenes, grids, and small plots are schematic illustrations.
Evaluation
Data and independent units
Purpose
Primary confirmation
24 train / 8 development / 12 confirmation initializations per task; 4 contexts × 6 coordinates per scene
Natural accuracy, controlled response, matched training
Representation/model extension
32 new evaluation initializations per task; radial/lateral context axes
Fixed-candidate replication and comparison
Free object
32 mug-position initializations; 64 support follow-up initializations
Physical scope and estimability
Failure prediction
96 development / 128 evaluation initializations; 2 normal action-noise rollouts each
Additional pre-rollout outcome information
Table 1: Evaluation design. Behavioral experiments use a task-adapted, frozen SmolVLA policy in LIBERO ( Liu et al., 2023 ) , separately from the Meta-World representation analysis. Appearance interventions, stove evaluations, and development controls use the distinct sample counts reported with their results.
Task
Model
Nat. MAE
Ctrl. MAE
G [95% CI]
EG
R
Drawer
SmolVLA
5.554
74.184
0.1806 [0.1653, 0.1978]
0.8194
16.499
OpenVLA
3.892
75.884
0.1525 [0.1387, 0.1662]
0.8475
9.797
Faucet
SmolVLA
0.0763
0.6698
0.0211 [0.0035, 0.0388]
0.9789
0.1285
OpenVLA
0.0570
0.6475
0.0512 [0.0292, 0.0717]
0.9488
0.1008
Table 2: Primary confirmation with full-natural readouts. Values are equal-weight means over 12 scenes per task. MAE and R use mm for Drawer and rad for Faucet; G and EG are dimensionless. Brackets are individual 95% scene-bootstrap intervals. Unit mean gain does not imply accuracy at every intermediate state. Appendix A.7 reports intervals for all metrics.
Figure 2: Confirmatory predictions and responses. Left: all 12 scenes × 4 context curves per combination (faint), mean predictions (bold), and identity lines (dashed). Right: scene-level endpoint gains and post-hoc all-coordinate slopes; large markers and bars show means and individual 95% intervals. Both summaries are below unit response. Mean gain does not characterize intermediate non-monotonicity.
Model
Drawer G : radial / lateral
Faucet G : radial / lateral
SmolVLA
0.2058 / 0.2060
0.0097 / 0.0122
OpenVLA
0.1692 / 0.1661
0.0429 / 0.0394
π0 base
0.1094 / 0.1086
−0.0106 / −0.0011
GR00T N1.6
0.2413 / 0.2426
0.0428 / 0.0435
Table 3: Mean endpoint gain of fixed final-mean-ridge candidates on 32 new initializations per task. Radial and lateral axes share initializations and use local ±10 mm hand-position changes relative to the robot. Model-native input paths differ, so this is not a controlled ranking of model quality. Appendix B provides MAE and candidate details.
Evaluation
Initializations
Result
6 cm position variation
32
Natural MAE 128.314 mm; controlled MAE 39.765 mm; G=0.140 . Natural error is already large.
Support-conditioned follow-up
64
43 physically valid; 5 support-valid; intersection 0. Below the minimum of 32 for estimation.
Table 4: Free-object results and estimability. The first evaluation passes its physical and visual criteria. The support follow-up is non-estimable, not a zero-gain or model-response failure. The target is one position component, not full 6-DoF state.
Figure 3: Matched controlled minus matched natural training. Points and bars show means and individual 95% paired scene-bootstrap intervals over 12 confirmation scenes. Negative values indicate lower error or sensitivity. For display only, MAE and R are divided by S=0.16 m for Drawer and S=1.4 rad for Faucet; metric definitions are unchanged.
Model
Information added to M0
Brier
Log loss
M0
None
0.181515
0.543533
MC
Controlled MAE
0.175654
0.528602
MGR
EG,R
0.171381
0.519075
MCGR
Controlled MAE, EG,R
0.169677
0.514777
Table 5: Failure-prediction losses on the same 128 evaluation initializations; lower is better. Two normal-noise rollouts are grouped within each equally weighted initialization. M0/MGR form the original fixed comparison; MC/MCGR are post-hoc additions. This evaluates failure prediction, not an improvement to policy success.
Figure 4: Mean loss differences and individual 95% intervals from 20,000 paired initialization-bootstrap draws. Negative differences favor the first model. Original and post-hoc comparisons are distinguished. Intervals describe evaluation uncertainty conditional on the development fits.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Condition
Artifact name
Training observations
Full natural
natural_full
All 24 natural training trajectories
Matched natural
natural36
Six scenes, six observations each; 36 total
Matched controlled
exact_control36
The same six scenes and 36 actual coordinates
Appendix
Table S1: Primary training conditions. Implementation names are included solely to map stored artifacts to scientific conditions.
Mean or adaptive 2×2 ; no numerical robot state in this feature path
GR00T N1.6
Final native visual tokens
Native 9×9 grid; mean or adaptive 2×2
SigLIP2
Vision pooler output
768; revision-pinned processor
VC-1
Normalized CLS
768; official resize/crop and ImageNet normalization
Appendix
Table S3: Frozen feature paths. Model-native token grids and preprocessing are retained.
Task
Metric
SmolVLA
OpenVLA
Drawer
Natural MAE
5.554 [4.921, 6.328]
3.892 [3.357, 4.477]
Controlled MAE
74.184 [72.277, 75.993]
75.884 [74.810, 76.974]
G
0.1806 [0.1653, 0.1978]
0.1525 [0.1387, 0.1662]
EG
0.8194 [0.8022, 0.8347]
0.8475 [0.8338, 0.8613]
R
16.499 [13.189, 20.415]
9.797 [7.576, 12.035]
Faucet
Natural MAE
0.0763 [0.0596, 0.0940]
0.0570 [0.0446, 0.0707]
Appendix
Table S4: Complete primary metrics: mean [individual 95% scene-bootstrap interval], n=12 per task. MAE/ R are mm for Drawer and rad for Faucet; G,EG are dimensionless.
Figure S1: All primary scene-level context sensitivities R , with means and individual 95% intervals. Each point is one initialization; contexts and observations are not independent sample units.
Task
Model
Matched natural MAE
Matched controlled MAE
G
Drawer
SmolVLA
4.155 mm
49.369 mm
0.1217
OpenVLA
3.242 mm
49.317 mm
0.1451
Faucet
SmolVLA
0.1025 rad
0.5976 rad
0.0193
OpenVLA
0.0755 rad
0.5873 rad
0.0442
Appendix
Table S5: Actual-coordinate-matched development evaluation.
Task
Model
Readout
Nat. MAE
Ctrl. MAE
G
EG
R
Drawer
SmolVLA
Linear
4.846
76.295
0.1785
0.8215
15.514
RBF
6.068
75.934
0.1696
0.8304
15.316
MLP
3.650
78.502
0.1506
0.8494
13.176
OpenVLA
Linear
3.606
76.217
0.1487
0.8513
10.534
RBF
4.590
77.307
0.1262
0.8738
9.752
MLP
2.408
78.310
0.1337
0.8663
8.822
Appendix
Table S6: Representative full-natural development results. MAE/ R are mm for Drawer and rad for Faucet. RBF uses fixed σ=d~ , λ=0.001 ; MLP averages three seed-specific metrics.
Task
Model
Natural MAE
Controlled MAE
G
Drawer
SmolVLA
7.303
72.420 / 74.559
0.2058 / 0.2060
OpenVLA
4.317
76.258 / 76.555
0.1692 / 0.1661
π0 base
6.689
81.547 / 82.161
0.1094 / 0.1086
GR00T N1.6
3.711
69.011 / 68.812
0.2413 / 0.2426
Faucet
SmolVLA
0.0698
0.6867 / 0.6792
0.0097 / 0.0122
OpenVLA
0.0484
0.6789 / 0.6778
0.0429 / 0.0394
Appendix
Table S7: Fixed final-mean-ridge results on 32 new initializations per task. MAEs are mm for Drawer and rad for Faucet. Paired entries denote radial/lateral axes.
Task
Model
Endpoint G
All-coordinate B
Drawer
SmolVLA
0.1806 [0.1653, 0.1978]
0.1618 [0.1477, 0.1762]
OpenVLA
0.1525 [0.1387, 0.1662]
0.1561 [0.1427, 0.1691]
Faucet
SmolVLA
0.0211 [0.0035, 0.0388]
0.0262 [0.0141, 0.0373]
OpenVLA
0.0512 [0.0292, 0.0717]
0.0550 [0.0394, 0.0702]
Appendix
Table S8: Endpoint gains and all-coordinate slopes: means [individual 95% intervals].
Task
Model
R increases
R decreases
Drawer
SmolVLA
12
0
OpenVLA
10
2
Faucet
SmolVLA
11
1
OpenVLA
9
3
Appendix
Table S9: Scene-level context-sensitivity changes under matched controlled replacement ( n=12 per task).
Design
Composition
Purpose
Joint-A
36 matched natural + 36 matched controlled = 72
Combine observations of the same scenes/coordinates
Joint-B
Full natural + 36 controlled; 688 Drawer / 515 Faucet
Retain all original natural training data
Fixed budget
36 total; controlled fractions 1/3,1/2,2/3 , each with two complementary assignments
Compare mixtures at fixed sample count
Appendix
Table S10: Joint-training designs. All fit a single readout to both observation types.
Task
Model
Pooling
Natural MAE
Controlled MAE
G
Drawer
SmolVLA
Full mean
7.501
72.656
0.200
Object ROI
10.038
42.553
0.521
Equal-area control
6.736
85.426
0.108
OpenVLA
Full mean
3.606
76.217
0.149
Object ROI
5.708
41.143
0.514
Equal-area control
3.586
70.235
0.243
Appendix
Table S11: Representative natural-full region-pooling controls. MAE units are mm for Drawer and rad for Faucet. Blank task/model cells inherit the preceding entry.
Evaluation
Roots
Natural MAE (mm)
Controlled MAE (mm)
G
Development
16
121.767
42.293
0.167
New roots
32
128.314
39.765
0.140
Appendix
Table S12: Free-object results with 6 cm state variation. Both evaluations pass their physical and visual criteria. Natural error is already large.
Gate
Roots
Physically valid
43
Prespecified support-valid
5
Physical and support-valid
0
Required minimum for response estimation
32
Appendix
Table S13: Free-object follow-up eligibility. No root satisfies both required gates.
Condition
Mean ∣Δq^∣ (mm)
First translation-command shift
Robot-low
4.482
0.03613
Robot-high
15.506
0.04550
Appendix
Table S14: Initial changes relative to sham under identical physical state. Translation shifts are normalized command L2 differences, not distances in physical coordinate units.
Input
Definition
Scaling
E0
Absolute prediction error at the normal initial observation
Divide by 0.14 m
Dq
Initial coordinate distance to success boundary −0.14 m
Divide by 0.14 m
Deef
Euclidean distance from end effector to cabinet origin
m
Drobot
Standardized RMS difference of seven arm joints from policy-training state mean
Policy-training standard deviation; dimensionless
EG
Mean context-wise endpoint gain error from six controlled snapshots
Dimensionless
R
Prediction range over three contexts, averaged over two coordinates
Divide by 0.14 m
Appendix
Table S15: Failure-predictor inputs, all measured before rollout.
Loss
M0
MGR
Difference
Individual 95% interval
Brier
0.181515
0.171381
−0.010134
[ −0.017898 , −0.002478 ]
Log loss
0.543533
0.519075
−0.024458
[ −0.043453 , −0.005586 ]
Appendix
Table S16: Original fixed failure-prediction comparison. Intervals condition on development fitting.
Successes among two rollouts
Roots
Mean whole-trajectory MAE (mm)
0/2
28
55.075
1/2
25
35.295
2/2
75
6.701
Appendix
Table S17: Post-hoc relationship between whole-trajectory error and success, grouping two normal rollouts per evaluation root.
Task
Original
Sham
Robot-low
Robot-high
BG-low
BG-high
Drawer
41
41
41
36
42
41
Stove
60
60
61
60
60
60
Appendix
Table S18: Appearance-intervention successes out of 64 rollouts per task/condition (32 roots, two noises).
Task
Central
−0.12 rad
+0.12 rad
Drawer
8
3
1
Stove
8
8
8
Appendix
Table S19: Initial development-pilot successes out of eight under different robot contexts.
Model
Added to M0
Brier
Log loss
M0
None
0.181515
0.543533
MC
Controlled MAE
0.175654
0.528602
MGR
EG,R
0.171381
0.519075
MCGR
Controlled MAE, EG,R
0.169677
0.514777
Appendix
Table S20: All failure-prediction models on 128 evaluation roots.
Models may share sensitivity to distribution shift
Novel state–context combinations are jointly OOD
Limited support-aware follow-ups
Joint-distribution OOD remains possible
Appendix
Table S28: Alternative explanations and the scope of the available controls.
Figure S2: Controlled observations at fixed robot contexts. Each row contains six coordinate settings from one development scene: Drawer seed 108300, robot-context frame 6; Faucet seed 110300, robot-context frame 12. Coordinates increase from left to right, follow the simulator’s native sign convention, and are rounded for display. A fixed 180∘ orientation correction presents the rendered images upright. These are illustrative examples; quantitative validation is reported in Appendices A and D.
Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected into the vision-language prefix, or fed directly to the action expert -- and almost always as a single current frame. Three questions remain open: (1) whether, and on which tasks, current state actually improves closed-loop control; (2) how much state history helps, and whether its benefit reflects genuine temporal variation rather than added conditioning capacity; and (3) where state should enter the model -- the vision-language backbone or the action-generation module. We answer these questions through controlled experiments on a flow-matching VLA, fixing the backbone, training data, action representation, and evaluation protocol throughout. We implement five representative interfaces -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluate them on 45 atomic tasks spanning three task families plus 20 composite tasks; we then sweep the state-history length from 1 to 96 frames to examine how historical state information affects model performance. The experiments yield systematic answers to all three questions, distilled into testable design principles for state-aware VLAs.
Yiren Zhao, Ziyang Chen, Ziyang Rao +5
1The Hong Kong University of Science and Technology (Guangzhou) · 3AI2 Robotics X-Lab · 2Multimedia Laboratory (MMLab), The Chinese University of Hong Kong
Vision-language-action (VLA) policies and World-Action Models (WAM) represent two increasingly important paradigms for robotic manipulation. However, it remains unclear whether future prediction in WAMs leads to behaviorally meaningful improvements beyond final task success. In this paper, we ask whether WAMs merely add future prediction, or whether they change robot behavior and internal representations in ways that are actionable for control. We introduce a model-agnostic diagnostic framework that compares WAMs and VLAs through two complementary lenses: behavioral rollout analysis and sparse-autoencoder-based feature analysis. The behavioral protocol measures action dynamics consistency, target-object progress, distractor disturbance, and runtime cost. The feature-space protocol characterizes internal representations as memorized, reactive, or predictive, revealing whether models encode future-oriented structure. Across LIBERO and RoboTwin2.0, we evaluate 7 policies spanning direct VLAs and joint, sequential, and auxiliary WAMs. Our results show that success alone hides key differences: WAMs often improve object-level behavior and target selectivity, but their gains depend on architecture and incur higher inference cost. Sequential WAMs show the clearest predictive structure, while auxiliary and joint WAMs respectively compress or entangle future information. These findings suggest future directions for WAMs design to preserve behaviorally actionable future representations for efficient manipulation.
Hung Mai, Bin Zhu, Tuan Do
National Economics University, Vietnam · N2TP Technology · Singapore Management University +1
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we probe the residual stream of π0.5 and find that task progress, the normalized time remaining in a trajectory, is linearly readable from the activations. We find that this signal is present in the pretrained PaliGemma backbone prior to training on any robot-specific data. A single linear probe generalizes to unseen tasks and varies under language counterfactuals when trained on multi-prompt data, but does not enable meaningful steering of the policy. These properties make the signal directly useful for instrumenting deployed VLAs. We use the probe as a simple label-free OOD detector, which detects stalled task progress, and find it competitive with state-of-the-art methods. Our results suggest that VLAs have rich, linearly readable internal representations of semantic quantities like task progress, and that learning to read these signals offers a lightweight, interpretable path toward monitoring deployed visuomotor policies.
Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan +2