Where Predictive Supervision Goes Shapes What VLA Policies Learn
Organizations: Graduate School of Data Science Seoul National University
Abstract
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
Figures & tables
| LIBERO | LIBERO-PRO | |||||||
|---|---|---|---|---|---|---|---|---|
| Variant | mean | Lang. | Position | Object | Task | Env. | mean | Ret. |
| Baseline (no forecast) | ||||||||
| Special-token ( ) | ||||||||
| Special-token ( ) | ||||||||
| Special-token ( ) ∗ | ||||||||
| Vision-token | ||||||||
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Prediction readout | Representative methods | How supervision reaches policy vision tokens |
| Autoregressive output tokens | CoT-VLA ( Zhao et al., 2025 ) | Sequence-mediated, no explicit same-patch attachment |
| Learned/query tokens | FLARE, DreamVLA, WoG, HiF-VLA ( Zheng et al., 2025 ; Zhang et al., 2025 ; Su et al., 2026 ; Lin et al., 2026 ) | Attention-mediated, no explicit same-patch attachment |
| Predictive latent alignment | VLA-JEPA, FutureVLA ( Sun et al., 2026 ; Xu et al., 2026 ) | Latent-state or intermediate alignment, no explicit same-patch attachment |
| Future readout + spatial auxiliary | FoMoVLA ( Li et al., 2026 ) | Indirect forecast readout with a direct spatial auxiliary |
| Separate world model | AHEAD ( Syed et al., 2026 ) | Patch-aligned outside the frozen policy |
| Special tokens | This study | Attention-mediated, no guaranteed same-position path |
| Component | Setting |
|---|---|
| Data | 500 LIBERO-Object episodes from agentview_rgb , with every tenth episode held out, giving 450 training and 50 validation episodes |
| Training clips | Every valid clip start frame in the training episodes |
| Input | images, patches, and an grid of 64 visual tokens |
| Encoder | 12-block ViT with width 192, three attention heads, and learned absolute positional embeddings |
| Future target | Stop-gradient patch residual at from a momentum encoder with |
| Loss | Per-patch squared error weighted by the normalized norm of each target residual |
| Component | Setting |
|---|---|
| Decoder data | 35 held-out episodes for fitting and 15 disjoint episodes for evaluation |
| Decoder input | Features from real current and future frames only |
| Optimization | Adam for 12,000 steps, batch size 48, learning rate with cosine decay |
| Objective | |
| Evaluation set | 617 samples per seed, comprising 418 near-center and 199 spatial-tail samples |
| Moving pixels | Mean absolute RGB change above for images scaled to , excluding samples with at most 20 moving pixels |
| All pixels | Moving pixels | |||
|---|---|---|---|---|
| Split | Vision | Special | Vision | Special |
| All | ||||
| Near-center | ||||
| Spatial tail | ||||
| Target | Split | Vision-token visual | Special-token visual | Special-token stream |
|---|---|---|---|---|
| Position | Near | |||
| Tail | ||||
| Identity | Near | |||
| Tail | ||||
| Future displacement | Near | |||
| Tail |
| Interface | Same-region gradient share | Centered gradient rank |
|---|---|---|
| Controlled, initialization | ||
| Vision-token | ||
| Shuffled route | ||
| Special-token | ||
| , trained | ||
| Vision-token | ||
| Condition | Init. share | Eff. rank | Position (spatial tail) | Gap closed (path on, %) | Gap closed (path off, %) |
|---|---|---|---|---|---|
| Special-token ( ) | – | ||||
| Vision-token | – |
| Condition | Identity | Future disp. | Action chunk | Geometry |
|---|---|---|---|---|
| Special-token ( ) | ||||
| Vision-token |
| Spatial-tail | ||||
| Condition | Gap closed (%) | Eff. rank | Position | Action |
| Untied special-token stream | ||||
| Native credit | ||||
| Credit blocked | ||||
| Credit shuffled | ||||
| Shared block weights | ||||
| Condition | Gap closed (%) | Eff. rank | Position (spatial tail) | Action (spatial tail) |
|---|---|---|---|---|
| Vision-token | ||||
| Pool-16 | ||||
| Pool-4 | ||||
| Pool-1 | ||||
| Shuffle | ||||
| Block |
| LIBERO | LIBERO-PRO | |||||
| Variant / suite | none | lang. | pos. | obj. | task | env. |
| Baseline (no forecast) | ||||||
| Spatial | ||||||
| Object | ||||||
| Goal | ||||||
| LIBERO-10 | ||||||
| LIBERO | LIBERO-PRO | |||||||
|---|---|---|---|---|---|---|---|---|
| Variant | mean | lang. | pos. | obj. | task | env. | mean | Ret. |
| Baseline | ||||||||
| Special-token | ||||||||
| Vision-token | ||||||||
| Interface | Prefix | L6 | L11 | L14 | L17 |
|---|---|---|---|---|---|
| Special-token, | |||||
| Special-token, | |||||
| Special-token, , anchored | |||||
| Vision-token |
| Target | Baseline | Special-token | Vision-token | Shuffled future |
|---|---|---|---|---|
| Object displacement | ||||
| EEF displacement | ||||
| Object position | ||||
| Object-to-gripper geometry |
| Future displacement | Position | |||
|---|---|---|---|---|
| Variant | L10 | L12 | L14 | L12 |
| Variable index | ||||
| Anchored index | ||||
| Task | Training | Held out | Median length |
|---|---|---|---|
| Pick from clutter | 219 | 24 | s |
| Pick two in order | 89 | 10 | s |
| Open pot and place | 90 | 10 | s |
| Total | 398 | 44 |
| Condition | Task | Blocks | Changed factor |
|---|---|---|---|
| In distribution | bowl tasks | 7 | none |
| Blurred camera | bowl tasks | 10 | main-camera image |
| Unseen pot layout | pot task | 11 | layout and lid grasp |
| Task | Eval. rounds | Eval. median | Training median |
|---|---|---|---|
| Pot | 33 | mm | mm |
| Bowl pick | 19 | mm | mm |
| Policy | ID (7) | Blur (10) | Pot (11) | OOD (21) |
|---|---|---|---|---|
| Baseline | ||||
| Special, | ||||
| Special, | ||||
| Special, , anchored | ||||
| Vision-token |
| Readout | Baseline | Special 16 | Special 256 | Anchored 256 | Vision |
|---|---|---|---|---|---|
| Future, s ahead ( ) | |||||
| Gripper width | 0.435 | 0.606 | 0.561 | 0.566 | 0.688 |
| Joint motion | 0.647 | 0.735 | 0.701 | 0.713 | 0.749 |
| Present state | |||||
| Container position (mm error) | 57.0 | 53.5 | 49.5 | 48.7 | 49.2 |
| Object position (mm error) | 58.0 | 56.1 | 60.5 | 55.2 | 54.1 |