Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
Organizations: Changwon National University · Chung-Ang University
Abstract
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.
Figures & tables
| Evaluation | Data and independent units | Purpose |
|---|---|---|
| Primary confirmation | 24 train / 8 development / 12 confirmation initializations per task; 4 contexts 6 coordinates per scene | Natural accuracy, controlled response, matched training |
| Representation/model extension | 32 new evaluation initializations per task; radial/lateral context axes | Fixed-candidate replication and comparison |
| Free object | 32 mug-position initializations; 64 support follow-up initializations | Physical scope and estimability |
| Failure prediction | 96 development / 128 evaluation initializations; 2 normal action-noise rollouts each | Additional pre-rollout outcome information |
| Task | Model | Nat. MAE | Ctrl. MAE | [95% CI] | ||
|---|---|---|---|---|---|---|
| Drawer | SmolVLA | 5.554 | 74.184 | 0.1806 [0.1653, 0.1978] | 0.8194 | 16.499 |
| OpenVLA | 3.892 | 75.884 | 0.1525 [0.1387, 0.1662] | 0.8475 | 9.797 | |
| Faucet | SmolVLA | 0.0763 | 0.6698 | 0.0211 [0.0035, 0.0388] | 0.9789 | 0.1285 |
| OpenVLA | 0.0570 | 0.6475 | 0.0512 [0.0292, 0.0717] | 0.9488 | 0.1008 |
| Model | Drawer : radial / lateral | Faucet : radial / lateral |
|---|---|---|
| SmolVLA | 0.2058 / 0.2060 | 0.0097 / 0.0122 |
| OpenVLA | 0.1692 / 0.1661 | 0.0429 / 0.0394 |
| base | 0.1094 / 0.1086 | / |
| GR00T N1.6 | 0.2413 / 0.2426 | 0.0428 / 0.0435 |
| Evaluation | Initializations | Result |
|---|---|---|
| 6 cm position variation | 32 | Natural MAE 128.314 mm; controlled MAE 39.765 mm; . Natural error is already large. |
| Support-conditioned follow-up | 64 | 43 physically valid; 5 support-valid; intersection 0. Below the minimum of 32 for estimation. |
| Model | Information added to | Brier | Log loss |
|---|---|---|---|
| None | 0.181515 | 0.543533 | |
| Controlled MAE | 0.175654 | 0.528602 | |
| 0.171381 | 0.519075 | ||
| Controlled MAE, | 0.169677 | 0.514777 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Condition | Artifact name | Training observations |
|---|---|---|
| Full natural | natural_full | All 24 natural training trajectories |
| Matched natural | natural36 | Six scenes, six observations each; 36 total |
| Matched controlled | exact_control36 | The same six scenes and 36 actual coordinates |
| Task | Training | Development | Confirmation |
|---|---|---|---|
| Drawer | 108200–108223 | 108300–108307 | 108400–108411 |
| Faucet | 110200–110223 | 110300–110307 | 110400–110411 |
| Model | Feature | Dimension/pooling |
|---|---|---|
| SmolVLA | Final prefix image tokens | Four spatial means concatenated; 3,840 primary |
| OpenVLA | Final-layer 256 causal image states | Mean; 4,096 primary |
| base | Final native image/prefix features | Mean or adaptive ; no numerical robot state in this feature path |
| GR00T N1.6 | Final native visual tokens | Native grid; mean or adaptive |
| SigLIP2 | Vision pooler output | 768; revision-pinned processor |
| VC-1 | Normalized CLS | 768; official resize/crop and ImageNet normalization |
| Task | Metric | SmolVLA | OpenVLA |
|---|---|---|---|
| Drawer | Natural MAE | 5.554 [4.921, 6.328] | 3.892 [3.357, 4.477] |
| Controlled MAE | 74.184 [72.277, 75.993] | 75.884 [74.810, 76.974] | |
| 0.1806 [0.1653, 0.1978] | 0.1525 [0.1387, 0.1662] | ||
| 0.8194 [0.8022, 0.8347] | 0.8475 [0.8338, 0.8613] | ||
| 16.499 [13.189, 20.415] | 9.797 [7.576, 12.035] | ||
| Faucet | Natural MAE | 0.0763 [0.0596, 0.0940] | 0.0570 [0.0446, 0.0707] |
| Task | Model | Matched natural MAE | Matched controlled MAE | |
|---|---|---|---|---|
| Drawer | SmolVLA | 4.155 mm | 49.369 mm | 0.1217 |
| OpenVLA | 3.242 mm | 49.317 mm | 0.1451 | |
| Faucet | SmolVLA | 0.1025 rad | 0.5976 rad | 0.0193 |
| OpenVLA | 0.0755 rad | 0.5873 rad | 0.0442 |
| Task | Model | Readout | Nat. MAE | Ctrl. MAE | |||
|---|---|---|---|---|---|---|---|
| Drawer | SmolVLA | Linear | 4.846 | 76.295 | 0.1785 | 0.8215 | 15.514 |
| RBF | 6.068 | 75.934 | 0.1696 | 0.8304 | 15.316 | ||
| MLP | 3.650 | 78.502 | 0.1506 | 0.8494 | 13.176 | ||
| OpenVLA | Linear | 3.606 | 76.217 | 0.1487 | 0.8513 | 10.534 | |
| RBF | 4.590 | 77.307 | 0.1262 | 0.8738 | 9.752 | ||
| MLP | 2.408 | 78.310 | 0.1337 | 0.8663 | 8.822 |
| Task | Model | Natural MAE | Controlled MAE | |
|---|---|---|---|---|
| Drawer | SmolVLA | 7.303 | 72.420 / 74.559 | 0.2058 / 0.2060 |
| OpenVLA | 4.317 | 76.258 / 76.555 | 0.1692 / 0.1661 | |
| base | 6.689 | 81.547 / 82.161 | 0.1094 / 0.1086 | |
| GR00T N1.6 | 3.711 | 69.011 / 68.812 | 0.2413 / 0.2426 | |
| Faucet | SmolVLA | 0.0698 | 0.6867 / 0.6792 | 0.0097 / 0.0122 |
| OpenVLA | 0.0484 | 0.6789 / 0.6778 | 0.0429 / 0.0394 |
| Task | Model | Endpoint | All-coordinate |
|---|---|---|---|
| Drawer | SmolVLA | 0.1806 [0.1653, 0.1978] | 0.1618 [0.1477, 0.1762] |
| OpenVLA | 0.1525 [0.1387, 0.1662] | 0.1561 [0.1427, 0.1691] | |
| Faucet | SmolVLA | 0.0211 [0.0035, 0.0388] | 0.0262 [0.0141, 0.0373] |
| OpenVLA | 0.0512 [0.0292, 0.0717] | 0.0550 [0.0394, 0.0702] |
| Task | Model | increases | decreases |
|---|---|---|---|
| Drawer | SmolVLA | 12 | 0 |
| OpenVLA | 10 | 2 | |
| Faucet | SmolVLA | 11 | 1 |
| OpenVLA | 9 | 3 |
| Design | Composition | Purpose |
|---|---|---|
| Joint-A | 36 matched natural + 36 matched controlled = 72 | Combine observations of the same scenes/coordinates |
| Joint-B | Full natural + 36 controlled; 688 Drawer / 515 Faucet | Retain all original natural training data |
| Fixed budget | 36 total; controlled fractions , each with two complementary assignments | Compare mixtures at fixed sample count |
| Task | Model | Pooling | Natural MAE | Controlled MAE | |
|---|---|---|---|---|---|
| Drawer | SmolVLA | Full mean | 7.501 | 72.656 | 0.200 |
| Object ROI | 10.038 | 42.553 | 0.521 | ||
| Equal-area control | 6.736 | 85.426 | 0.108 | ||
| OpenVLA | Full mean | 3.606 | 76.217 | 0.149 | |
| Object ROI | 5.708 | 41.143 | 0.514 | ||
| Equal-area control | 3.586 | 70.235 | 0.243 |
| Evaluation | Roots | Natural MAE (mm) | Controlled MAE (mm) | |
|---|---|---|---|---|
| Development | 16 | 121.767 | 42.293 | 0.167 |
| New roots | 32 | 128.314 | 39.765 | 0.140 |
| Gate | Roots |
|---|---|
| Physically valid | 43 |
| Prespecified support-valid | 5 |
| Physical and support-valid | 0 |
| Required minimum for response estimation | 32 |
| Condition | Mean (mm) | First translation-command shift |
|---|---|---|
| Robot-low | 4.482 | 0.03613 |
| Robot-high | 15.506 | 0.04550 |
| Input | Definition | Scaling |
|---|---|---|
| Absolute prediction error at the normal initial observation | Divide by 0.14 m | |
| Initial coordinate distance to success boundary m | Divide by 0.14 m | |
| Euclidean distance from end effector to cabinet origin | m | |
| Standardized RMS difference of seven arm joints from policy-training state mean | Policy-training standard deviation; dimensionless | |
| Mean context-wise endpoint gain error from six controlled snapshots | Dimensionless | |
| Prediction range over three contexts, averaged over two coordinates | Divide by 0.14 m |
| Loss | Difference | Individual 95% interval | ||
|---|---|---|---|---|
| Brier | 0.181515 | 0.171381 | [ , ] | |
| Log loss | 0.543533 | 0.519075 | [ , ] |
| Successes among two rollouts | Roots | Mean whole-trajectory MAE (mm) |
|---|---|---|
| 0/2 | 28 | 55.075 |
| 1/2 | 25 | 35.295 |
| 2/2 | 75 | 6.701 |
| Task | Original | Sham | Robot-low | Robot-high | BG-low | BG-high |
|---|---|---|---|---|---|---|
| Drawer | 41 | 41 | 41 | 36 | 42 | 41 |
| Stove | 60 | 60 | 61 | 60 | 60 | 60 |
| Task | Central | rad | rad |
|---|---|---|---|
| Drawer | 8 | 3 | 1 |
| Stove | 8 | 8 | 8 |
| Model | Added to | Brier | Log loss |
|---|---|---|---|
| None | 0.181515 | 0.543533 | |
| Controlled MAE | 0.175654 | 0.528602 | |
| 0.171381 | 0.519075 | ||
| Controlled MAE, | 0.169677 | 0.514777 |
| Comparison | Brier difference | Log-loss difference |
|---|---|---|
| [ ,0.000381] | [ ,0.002338] | |
| [ ,0.001869] | [ ,0.006211] | |
| [ , ] | [ , ] | |
| [ , ] | [ , ] |
| Task | Roots | Planned | Success | Timeout | Unknown |
|---|---|---|---|---|---|
| Mug to plate | 32 | 64 | 13 | 51 | 0 |
| Book to compartment | 32 | 64 | 2 | 57 | 5 |
| Middle book to shelf | 32 | 64 | 1 | 48 | 15 |
| Right book to shelf | 32 | 64 | 11 | 44 | 9 |
| Descriptive total | 128 | 256 | 27 | 200 | 29 |
| Subset | Drawer | Faucet |
|---|---|---|
| Main controlled grid | 32/32 | 32/32 |
| Natural- -matched grid | 0/32 | 18/32 |
| Target- -support-restricted main grid | 0/32 | 32/32 |
| Support-restricted natural-matched grid | 0/32 | 0/32 |
| Task | Model | Sham | Robot-high |
|---|---|---|---|
| Drawer | SmolVLA | 0.2124 | 0.1730 |
| OpenVLA | 0.1608 | 0.1126 | |
| 0.0992 | |||
| GR00T | 0.2388 | 0.2326 | |
| Faucet | SmolVLA | 0.0135 | 0.0042 |
| OpenVLA | 0.0345 |
| Encoder | Drawer | Faucet |
|---|---|---|
| SigLIP2 | 0.3357 | 0.0794 |
| VC-1 | 0.1979 | 0.0503 |
| R3M | 0.3445 | 0.1391 |
| DINOv3 | 0.1837 | 0.1611 |
| Task | Natural MAE | Controlled MAE | |||
|---|---|---|---|---|---|
| Drawer | 3.947 | 77.460 | 0.1837 | 0.8163 | 7.158 |
| Faucet | 0.0414 | 0.5532 | 0.1611 | 0.8389 | 0.0937 |
| Task | Appearance | Natural MAE | Controlled MAE | ||
|---|---|---|---|---|---|
| Drawer | Original/sham | 3.947 | 77.460 | 0.1837 | 7.158 |
| Robot-low | 12.127 | 69.297 | 0.2036 | 10.336 | |
| Robot-high | 17.906 | 55.802 | 0.2587 | 14.443 | |
| Background-low | 6.396 | 79.937 | 0.1957 | 8.143 | |
| Background-high | 8.431 | 80.598 | 0.2106 | 6.352 | |
| Faucet | Original/sham | 0.0414 | 0.5532 | 0.1611 | 0.0937 |
| Explanation | Available evidence | Remaining question |
|---|---|---|
| Physical/rendering errors | Actual coordinate, robot, visibility, repeated-render and target-pixel checks | Does not establish natural joint-distribution membership |
| Different target-coordinate ranges | Actual-coordinate matching and separate faucet state-value matching | Does not match the entire context distribution |
| Variation confined to artificial images | Prediction differences in natural-only near-equal- pairs | Observational matching does not isolate a causal context factor |
| One VLA or readout fails | Multiple VLAs, feature paths, pooling, nonlinear readouts, non-VLA controls | Models may share sensitivity to distribution shift |
| Novel state–context combinations are jointly OOD | Limited support-aware follow-ups | Joint-distribution OOD remains possible |