cs.ROSep 29, 2026

Where Predictive Supervision Goes Shapes What VLA Policies Learn

Authors: Hanseul Kim, Jewon Yeom, Youngjoon Jeong, Minsoo Jo, Taesup Kim

Organizations: Graduate School of Data Science Seoul National University

Abstract

Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

    Sep 29, 2026Junghyun Kim, Ngseo Kim, ChungWoo Lee +7Latent PredictionSpurious Correlations

  2. FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

    Jul 16, 2026Wei Li, Peijin Jia, Yuan Ma +9Photometric SupervisionForesight