Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .
Figures & tables
Figure 1: Overview of Vela. At inference, Vela jointly predicts a fixed set of spline control points and a motion-dependent horizon, defining a continuous action trajectory that can be sampled at arbitrary control rates. During training, motion-dependent horizon adaptation constructs supervision targets that adequately represent the demonstrated motion. The same control point budget thus covers longer spans for smooth motion and shorter spans when finer temporal resolution is needed.
Model
Level 1
Level 2
Level 3
Level 4
Level 5
Avg.
OpenVLA-OFT ( Kim et al., 2025 )
29.0
17.6
8.8
6.4
4.2
13.2
π0 ( Black et al., 2024 )
29.4
21.9
11.0
7.6
5.1
15.0
X-VLA ( Zheng et al., 2026 )
30.1
22.6
10.3
6.0
4.1
14.6
OpenWAM- α ( Wang et al., 2026d )
35.9
28.2
21.6
13.0
10.0
21.7
GR00T N1.5 ( Bjorck et al., 2025 )
43.3
32.9
18.7
13.3
9.7
23.6
τ0 -VLA ( Cai et al., 2026c )
46.5
34.2
20.5
13.0
10.5
24.9
Table 1: Evaluation on LIBERO-X ( Wang et al., 2026a ) across five progressively challenging levels. We report success rate (%), with green values indicating absolute gains over baseline π0.5 . π0.5 -Vela denotes variant initialized from π0.5 and post-trained in trajectory space.
Model
Table Top
Simple PnP
Long Horizon
Overall
SR
Score
SR
Score
SR
Score
SR
Score
X-VLA ( Zheng et al., 2026 )
8.6
24
50.0
54
6.2
25
23.7
36
InternVLA-A1 ( Cai et al., 2026a )
4.3
11
43.0
47
17.9
46
23.9
36
FastWAM ( Yuan et al., 2026b )
2.9
13
49.5
53
16.9
38
25.6
37
Cosmos3-Edge ( Agarwal et al., 2026 )
18.6
32
41.5
44
24.1
49
29.3
42
GigaBrain-0.7 ( Team et al., 2026 )
37.9
59
42.0
45
20.0
37
33.3
46
Table 2: Evaluation on EBench ( Gao et al., 2026 ) across its three manipulation task families. We report success rate (SR, %) and the official task-progress Score (Score, %) with green values indicating the absolute gains over baseline π0.5 .
Figure 2: Vela executing cooking egg cake , which spans four subtasks (top to bottom).
Figure 3: Real-world results on the cooking egg cake task across four subtasks.
Configuration
Success Rate (%)
Method
Post-hoc Traj. Fit.
Max. Steps
Level 1
Level 2
Level 3
Level 4
Level 5
Avg.
✗
10
65.2
53.2
36.0
24.1
18.0
39.3
π0.5
✓
10
61.7
50.5
37.0
24.7
20.4
38.9
✓
25
62.6
52.2
38.8
24.7
21.9
40.0
π0.5 -Vela
N/A
25
66.1
53.5
40.7
27.3
22.5
42.0
Vela
N/A
25
67.2
55.9
44.9
31.5
27.1
45.3
Table 3: Ablation of post-hoc trajectory fitting (Post-hoc Traj. Fit.) on LIBERO-X. Post-hoc Traj. Fit. spline fits and resamples the predicted point-wise action chunk, whereas Vela predicts trajectory parameters directly. Max. Steps denotes the maximum number of actions executed before replanning.
Method
Level 1
Level 2
Level 3
Level 4
Level 5
Avg.
π0.5
65.2
53.2
36.0
24.1
18.0
39.3
Vela w/ Fixed H.
66.4
54.3
42.8
28.1
24.4
43.2
Vela w/ Adap. H.
67.2
55.9
44.9
31.5
27.1
45.3
Table 4: Effect of adaptive temporal support on LIBERO-X. We compare Vela with globally fixed temporal support (Fixed H.) against the full model with adaptive temporal support (Adap. H.).
Objectives
Success Rate (%)
CP-FM
DT-FM
Level 1
Level 2
Level 3
Level 4
Level 5
Avg.
✓
65.2
54.7
43.1
31.3
25.8
44.0
✓
✓
65.7
56.0
44.0
31.3
27.6
44.9
✓
67.2
55.9
44.9
31.5
27.1
45.3
Table 5: Ablation of trajectory geometry objectives on LIBERO-X.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Ours
Qwen- RobotManip
π0.5
InternVLA- A1.5
π0
GigaBrain- 0.7
Cosmos3- Edge
FastWAM
X-VLA
TableTop
collect coffee beans
15
5
30
0
20
50
0
0
0
flip cup collect cookies
60
90
25
30
30
40
30
0
5
frame against pen holder
65
95
45
0
25
70
20
0
30
install gear
25
50
10
15
30
20
15
15
5
peg in hole
15
10
10
0
15
15
30
0
5
Appendix
Table 6: Task-level results on EBench. Success rate (SR, %) for each of the 26 tasks, grouped into dexterous-and-precise tabletop, mobile pick-and-place (PnP), and mobile long-horizon tasks. The final row reports benchmark-wide performance.
Figure 4: Detailed capability profile of Vela and π0.5 on EBench. We report the official task progress Score (%) across (a) the four generalization conditions, (b) atomic skills, and (c) operating mode, temporal horizon, and precision. Vela consistently improves over π0.5 across the displayed categories
Figure 5: Fine-grained capability breakdown on EBench. We compare Vela with representative baselines using the official task progress Score (%) across operating mode (Mobile / Fixed-base), temporal horizon (Short / Long), precision (Low / Medium / High), and four generalization conditions (Background / Instruction / Object / Mix).
Figure 6: Robustness to replanning interval on LIBERO-X. We vary the maximum number of executed actions before replanning for both π0.5 and Vela.
Figure 7: Vela executing shredding potato , which spans four subtasks (top to bottom).
Figure 8: Additional real-world potato-shredding evaluation. Each bar reports the success rate over ten trials for one of the four ordered subtasks.
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task π0.5 policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address this issue, we propose MotionWeave, a motion-centric future-dynamics framework for action-chunk prediction with two modules: the Action-Induced Motion Grounder (AIMG) and the Horizon Residual Composer (HRC). Specifically, AIMG conditions on action and proprioceptive representations to construct horizon-specific queries that localize interaction regions associated with each future action timestep from current visual tokens. HRC extracts differences between interaction representations at adjacent horizons, encodes them as temporal motion cues, and injects them into action tokens through a gated residual. During training, robot-arm masks rendered from future frames are used to construct KL-based motion-grounding supervision, while inference uses only the current observation. On six MetaWorld tasks, MotionWeave achieves a 75.3% average success rate, an absolute gain of 8.6% over π0 (66.7%), especially on sustained-interaction tasks. Our code is available at https://github.com/autu-mn/MotionWeave.
Jingqiu Wang, Yan Wang
School of Data Science and Engineering East China Normal University, Shanghai, China