Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address this issue, we propose MotionWeave, a motion-centric future-dynamics framework for action-chunk prediction with two modules: the Action-Induced Motion Grounder (AIMG) and the Horizon Residual Composer (HRC). Specifically, AIMG conditions on action and proprioceptive representations to construct horizon-specific queries that localize interaction regions associated with each future action timestep from current visual tokens. HRC extracts differences between interaction representations at adjacent horizons, encodes them as temporal motion cues, and injects them into action tokens through a gated residual. During training, robot-arm masks rendered from future frames are used to construct KL-based motion-grounding supervision, while inference uses only the current observation. On six MetaWorld tasks, MotionWeave achieves a 75.3% average success rate, an absolute gain of 8.6% over π0 (66.7%), especially on sustained-interaction tasks. Our code is available at https://github.com/autu-mn/MotionWeave.
Figures & tables
Figure 1: Overview of MotionWeave. The current observation It and instruction l are tokenized into visual tokens and text tokens, concatenated with a learnable action query placeholder A0 , and jointly processed by LLM to produce action tokens A ; a separate proprioceptive encoder maps St to P . The Action-Induced Motion Grounder (AIMG) combines A and P into horizon-specific queries that attend over visual tokens V to produce interaction tokens M with spatial attention α . The Horizon Residual Composer (HRC) injects adjacent-horizon differences of M into A via gate g to obtain A′ , decoded by Action DiT into a 4-step chunk. Robot-arm masks [ 8 ] supervise α via KL loss in training only.
Figure 2: Qualitative rollouts on six tasks. Each row shows key frames from observation to successful completion.
Method
Source
Pick Place
Disassemble
Stick Pull
Assembly
Shelf Place
Hand Insert
Avg.
π0 [ 1 ]
RSS’25
72.0
52.0
68.0
76.0
64.0
68.0
66.7
DreamVLA [ 22 ]
NeurIPS’25
60.0
68.0
72.0
70.0
56.0
64.0
65.0
WoG [ 16 ]
ICML’26
28.0
60.0
68.0
24.0
64.0
60.0
50.7
Fast-WAM [ 21 ]
arXiv’26
44.0
56.0
16.0
64.0
56.0
60.0
49.3
MotionWeave (ours)
—
60.0
92.0
76.0
80.0
76.0
68.0
75.3
Table 1: Success rates (%) on six MetaWorld tasks. All methods share the evaluation random seeds, and bold denotes the best result for each task. The “Source” column reports the publication status of each baseline.
Variant
Avg.
Baseline
58.0
+ AIMG w/o Lmot
62.0
+ AIMG
70.0
+ AIMG + HRC
75.3
Table 2: Component ablation results for MotionWeave. We report macro-average success (%) over the six tasks. Each row adds one component relative to the preceding configuration.
Figure 3: Hyperparameter curves in average success (%). (a) Horizon query count peaks at 4 with 75.3%, aligned with the action chunk length H=4 , falling back to 73.3%/70.0% at 6/8 queries. (b) Motion loss weight λmotion peaks at 0.05 with 75.3%, falling back to 72.7%/68.7% at 0.10/0.20.
Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.
Wei Li, Peijin Jia, Yuan Ma +9
LiAuto · University of Chinese Academy of Sciences · Tsinghua University
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task π0.5 policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
Vision-Language-Action (VLA) models generalize across static manipulation but fail when objects move during task execution. They map the current observation to an action and assume the scene is stationary between observation and execution, so at any non-trivial object speed the resulting latency exceeds the time available to grasp. We close this gap with AHEAD (Anticipatory Horizon Extrapolation with Adaptive Dynamics), a predict-then-act wrapper that augments a frozen VLA with a motion-aware latent world model. A small world model trained on manipulation video forecasts future patch tokens in the VLA's feature space, conditioned on per-token velocity and acceleration from optical flow. A language-and-motion saliency mask concentrates prediction on task-relevant patches, and the model rolls forward for an adaptive horizon, halting when prediction uncertainty crosses a threshold. The frozen action decoder then receives the predicted future tokens in place of the current ones. AHEAD adds 4.9M parameters to a frozen 7B OpenVLA and reaches 79 to 97% success across 20 dynamic simulation scenarios where the strongest baseline reaches 31 to 58%. On a physical UFactory xArm 7, AHEAD succeeds on 29/30 to 30/30 on three conveyor and rolling-ball tasks, 23/30 on paddle interception, and 19/30 on projectile catching where every baseline scores 0/30.
Shahram Najam Syed, Arthur Jakobsson, Haoran Hao +1
Robotics Institute, Carnegie Mellon University, Pittsburgh, USA