Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.
Figures & tables
Figure 1: Comparison of VLA models, joint WAMs, and Devol-ONE. ( Left ) Standard VLA models map vision-language features directly to actions. ( Middle ) Joint WAMs use a shared transformer to predict future video and actions. ( Right ) Devol-ONE combines a VL stream, a JEPA dynamics stream initialized from a V-JEPA pretrained backbone, and an action expert through layerwise joint attention.
Figure 2: Overview of Devol-ONE pipeline. A VLM, World Model (WM), and Action Expert are trained as unified layerwise Mixture-of-Transformers blocks with masked joint attention, optimized with a world-model loss Ldyn and an action loss Lact ( left ). At inference ( right ), a JEPA encoder embeds the current observation, and the predictor is autoregressively rolled out through the shared MoT block to forecast future latents conditioned on the VLM. The resulting latent trajectory conditions the Action Expert, which uses a DiT head to generate the action chunk.
Figure 3: Example rollouts on LIBERO.
Model
Spatial
Object
Goal
Long
Avg
Diffusion Policy ( Chi et al., 2023 )
78.3
92.5
68.3
50.5
72.4
Octo ( Team et al., 2024 )
78.9
85.7
84.6
51.1
75.1
OpenVLA ( Kim et al., 2024 )
84.7
88.4
79.2
53.7
76.5
SpatialVLA ( Qu et al., 2025 )
88.2
89.9
78.6
55.5
78.1
CoT-VLA ( Zhao et al., 2025 )
87.5
91.6
87.6
69.0
81.1
WorldVLA (512) ( Cen et al., 2025b )
87.6
96.2
83.4
60.0
81.8
Table 1: Success rates (%) on the LIBERO benchmark.
Method
Camera
Robot
Language
Light
Background
Noise
Layout
Total
OpenVLA ( Kim et al., 2024 )
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
WorldVLA ( Cen et al., 2025b )
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
NORA
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
UniVLA ( Bu et al., 2025a )
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
Fast-WAM ( Yuan et al., 2026 )
44.5
68.9
60.7
53.7
37.7
16.4
78.2
51.5
π0 ( Black et al., 2024 )
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
Table 2: Zero-shot performance on LIBERO-Plus. All methods are trained only on the standard LIBERO dataset without fine-tuning on LIBERO-Plus dataset.
Click Alarmclock
Dump Bin Bigbin
Place Bread Basket
Place Can Basket
RoboTwin
Clean
Randomized
Clean
Randomized
Clean
Randomized
Clean
Randomized
ACT ( Zhao et al., 2023 )
32.0
4.0
68.0
1.0
6.0
0.0
1.0
0.0
Diffusion Policy ( Chi et al., 2023 )
61.0
5.0
49.0
0.0
14.0
0.0
18.0
0.0
RDT ( Liu et al., 2024 )
61.0
12.0
64.0
32.0
10.0
2.0
19.0
6.0
π0 ( Black et al., 2024 )
63.0
11.0
83.0
24.0
17.0
4.0
41.0
5.0
DP3 ( Ze et al., 2024 )
77.0
14.0
85.0
53.0
26.0
1.0
67.0
2.0
Table 3: Simulation results on the RoboTwin 2.0 leaderboard. We report success rates (%) under Clean and Randomized settings for four representative tasks.
Abbreviation
Task prompt (verbatim)
Episodes
Ethernet
Insert the Ethernet connector
118
2arm box
Stack realsense boxes with both arms
200
1arm box
Stack realsense boxes with only the left arm
200
Basket
Place the snacks into a basket
200
4cups
Stack four different-colored cups
200
Table 4: Flexiv manipulation tasks and demonstration counts. Abbreviations are used in Table 7 . Details on the task and task criteria are listed in Appendix A.5.1
Table 8
Task
Ethernet
2arm box
1arm box
Basket
4cups
Avg.
GigaBrain-0.7 Team et al. (2026)
0.0
80.0
90.0
20.0
70.0
52.0
DM0.5 Yu et al. (2026)
0.0
40.0
45.0
25.0
20.0
26.0
Fast-WAM Yuan et al. (2026)
0.0
55.0
60.0
0.0
0.0
23.0
π0 Black et al. (2024)
40.0
90.0
100.0
35.0
10.0
55.0
π0.5 Intelligence et al. (2025)
45.0
100.0
100.0
40.0
35.0
64.0
Devol-ONE (Ours)
55.0
100.0
100.0
50.0
75.0
76.0
Table 7: Real-world rollout success rates (%). Full task prompts are listed in Table 4 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Task
π0.5
X-VLA
Fast-WAM
GigaBrain-0.7
X-WAM
Devol-ONE (Ours)
Adjust Bottle
99 / 77
99 /33
99 /0
98 / 98
92/13
98 / 98
Beat Block Hammer
93 /23
87/9
82/0
94 / 74
92/21
86/ 93
Blocks Ranking RGB
72/44
3/4
81 /0
68/ 76
85 /36
55/ 64
Blocks Ranking Size
45/21
56 /25
52 /0
42/ 28
33/0
26/ 39
Click Alarmclock
65/50
62/43
99 /60
98 / 96
95/61
98 / 99
Click Bell
31/35
100 /63
95/7
98 / 84
100 /52
100 / 100
Appendix
Table 8: Per-task success rates on all 50 RoboTwin 2.0 tasks (Clean/Randomized, %). Baselines are co-trained on all 50 tasks and taken from the official RoboTwin 2.0 leaderboard; Devol-ONE is trained with single-task fine-tuning. In each column, the best result is in bold and the second best is underlined . Overall is the mean of the Clean and Randomized averages.
Figure 4: Example rollouts on RoboTwin 2.0.
Factor
#Ep.
Zero-shot
SFT 50k
SFT 150k
Δ
Camera
1,599
47.3
92.1
93.9
+ 1.8
Robot
1,550
45.7
43.1
48.2
+ 5.1
Language
1,537
85.9
81.7
84.8
+ 3.1
Light
1,142
94.5
97.2
97.9
+ 0.7
Background
1,076
91.8
96.6
96.8
+ 0.2
Noise
1,601
67.5
92.9
94.3
+ 1.4
Appendix
Table 9: In-distribution results on LIBERO-Plus (%). Δ : change from 50k to 150k steps. #Ep.: evaluation episodes per factor. Robot denotes robot initial-state perturbations.
Conditioning
Camera
Robot
Language
Light
Background
Noise
Layout
Total
VL only
48.7
43.1
82.8
90.6
90.1
62.1
79.1
69.0
JEPA only
37.5
55.4
72.3
92.7
89.9
46.8
77.6
65.1
Layerwise (Ours)
47.3
45.7
85.9
94.5
91.8
67.5
80.5
71.4
Appendix
Table 10: Action-expert conditioning on LIBERO-Plus (zero-shot, %). Best result per column in bold.
Hyperparameter
Value
Vision-language stream:
Backbone
Qwen3-VL-2B-Instruct
Hidden size
2048
Attention implementation
SDPA
Dynamics stream:
Encoder
V-JEPA2 ViT-L/16 ( Assran et al., 2025 )
Appendix
Table 11: Backbone and world-model hyperparameters.
Hyperparameter
Value
Action dimensions
14
State dimensions
16
Future action window size K
6
Past action window size
0
Action horizon H
7
Repeated diffusion steps
8
Appendix
Table 12: Action expert hyperparameters.
Property
Value
Robot-action stream:
AgiBotWorld-RL archives (Home + Industry)
78
AgiBotWorld-RL sample weight (per archive)
1.0
DROID sample weight
150
Composition (DROID : AgiBotWorld-RL)
≈ 65.8% : 34.2%
Action representation
delta-position, delta-rotation-vector
Appendix
Table 13: Pretraining data mixture.
Hyperparameter
Value
Optimizer (AdamW):
β1
0.9
β2
0.95
ϵ
10−8
Weight decay
10−8
Learning rate:
Appendix
Table 14: Optimization hyperparameters.
Figure 5: Three sample episodes of the Ethernet task (rows, top to bottom), each showing the egocentric view and the left-wrist view at the first and final frame.
Figure 6: Three sample episodes of the 2arm box task variant (dual-arm; rows, top to bottom), each showing the egocentric view and both wrist views at the first and final frame.
Figure 7: Three sample episodes of the 1arm box task variant (single, left arm only; rows, top to bottom).
Figure 8: Three sample episodes of the Basket task (rows, top to bottom).
Figure 9: Three sample episodes of the 4cups task (rows, top to bottom).