Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling
Organizations: Devol Robots, San Francisco, CA, United States
Abstract
Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.
Figures & tables
| Model | Spatial | Object | Goal | Long | Avg |
| Diffusion Policy ( Chi et al., 2023 ) | 78.3 | 92.5 | 68.3 | 50.5 | 72.4 |
| Octo ( Team et al., 2024 ) | 78.9 | 85.7 | 84.6 | 51.1 | 75.1 |
| OpenVLA ( Kim et al., 2024 ) | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 |
| SpatialVLA ( Qu et al., 2025 ) | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 |
| CoT-VLA ( Zhao et al., 2025 ) | 87.5 | 91.6 | 87.6 | 69.0 | 81.1 |
| WorldVLA (512) ( Cen et al., 2025b ) | 87.6 | 96.2 | 83.4 | 60.0 | 81.8 |
| Method | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| OpenVLA ( Kim et al., 2024 ) | 0.8 | 3.5 | 23.0 | 8.1 | 34.8 | 15.2 | 28.5 | 15.6 |
| WorldVLA ( Cen et al., 2025b ) | 0.1 | 27.9 | 41.6 | 43.7 | 17.1 | 10.9 | 38.0 | 25.0 |
| NORA | 2.2 | 37.0 | 65.1 | 45.7 | 58.6 | 12.8 | 62.1 | 39.0 |
| UniVLA ( Bu et al., 2025a ) | 1.8 | 46.2 | 69.6 | 69.0 | 81.0 | 21.2 | 31.9 | 42.9 |
| Fast-WAM ( Yuan et al., 2026 ) | 44.5 | 68.9 | 60.7 | 53.7 | 37.7 | 16.4 | 78.2 | 51.5 |
| ( Black et al., 2024 ) | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 |
| Click Alarmclock | Dump Bin Bigbin | Place Bread Basket | Place Can Basket | |||||
| RoboTwin | Clean | Randomized | Clean | Randomized | Clean | Randomized | Clean | Randomized |
| ACT ( Zhao et al., 2023 ) | 32.0 | 4.0 | 68.0 | 1.0 | 6.0 | 0.0 | 1.0 | 0.0 |
| Diffusion Policy ( Chi et al., 2023 ) | 61.0 | 5.0 | 49.0 | 0.0 | 14.0 | 0.0 | 18.0 | 0.0 |
| RDT ( Liu et al., 2024 ) | 61.0 | 12.0 | 64.0 | 32.0 | 10.0 | 2.0 | 19.0 | 6.0 |
| ( Black et al., 2024 ) | 63.0 | 11.0 | 83.0 | 24.0 | 17.0 | 4.0 | 41.0 | 5.0 |
| DP3 ( Ze et al., 2024 ) | 77.0 | 14.0 | 85.0 | 53.0 | 26.0 | 1.0 | 67.0 | 2.0 |
| Abbreviation | Task prompt (verbatim) | Episodes |
| Ethernet | Insert the Ethernet connector | 118 |
| 2arm box | Stack realsense boxes with both arms | 200 |
| 1arm box | Stack realsense boxes with only the left arm | 200 |
| Basket | Place the snacks into a basket | 200 |
| 4cups | Stack four different-colored cups | 200 |
| Task | Ethernet | 2arm box | 1arm box | Basket | 4cups | Avg. |
| GigaBrain-0.7 Team et al. (2026) | 0.0 | 80.0 | 90.0 | 20.0 | 70.0 | 52.0 |
| DM0.5 Yu et al. (2026) | 0.0 | 40.0 | 45.0 | 25.0 | 20.0 | 26.0 |
| Fast-WAM Yuan et al. (2026) | 0.0 | 55.0 | 60.0 | 0.0 | 0.0 | 23.0 |
| Black et al. (2024) | 40.0 | 90.0 | 100.0 | 35.0 | 10.0 | 55.0 |
| Intelligence et al. (2025) | 45.0 | 100.0 | 100.0 | 40.0 | 35.0 | 64.0 |
| Devol-ONE (Ours) | 55.0 | 100.0 | 100.0 | 50.0 | 75.0 | 76.0 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | X-VLA | Fast-WAM | GigaBrain-0.7 | X-WAM | Devol-ONE (Ours) | |
| Adjust Bottle | 99 / 77 | 99 /33 | 99 /0 | 98 / 98 | 92/13 | 98 / 98 |
| Beat Block Hammer | 93 /23 | 87/9 | 82/0 | 94 / 74 | 92/21 | 86/ 93 |
| Blocks Ranking RGB | 72/44 | 3/4 | 81 /0 | 68/ 76 | 85 /36 | 55/ 64 |
| Blocks Ranking Size | 45/21 | 56 /25 | 52 /0 | 42/ 28 | 33/0 | 26/ 39 |
| Click Alarmclock | 65/50 | 62/43 | 99 /60 | 98 / 96 | 95/61 | 98 / 99 |
| Click Bell | 31/35 | 100 /63 | 95/7 | 98 / 84 | 100 /52 | 100 / 100 |
| Factor | #Ep. | Zero-shot | SFT 50k | SFT 150k | |
| Camera | 1,599 | 47.3 | 92.1 | 93.9 | 1.8 |
| Robot | 1,550 | 45.7 | 43.1 | 48.2 | 5.1 |
| Language | 1,537 | 85.9 | 81.7 | 84.8 | 3.1 |
| Light | 1,142 | 94.5 | 97.2 | 97.9 | 0.7 |
| Background | 1,076 | 91.8 | 96.6 | 96.8 | 0.2 |
| Noise | 1,601 | 67.5 | 92.9 | 94.3 | 1.4 |
| Conditioning | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| VL only | 48.7 | 43.1 | 82.8 | 90.6 | 90.1 | 62.1 | 79.1 | 69.0 |
| JEPA only | 37.5 | 55.4 | 72.3 | 92.7 | 89.9 | 46.8 | 77.6 | 65.1 |
| Layerwise (Ours) | 47.3 | 45.7 | 85.9 | 94.5 | 91.8 | 67.5 | 80.5 | 71.4 |
| Hyperparameter | Value |
| Vision-language stream: | |
| Backbone | Qwen3-VL-2B-Instruct |
| Hidden size | 2048 |
| Attention implementation | SDPA |
| Dynamics stream: | |
| Encoder | V-JEPA2 ViT-L/16 ( Assran et al., 2025 ) |
| Hyperparameter | Value |
| Action dimensions | 14 |
| State dimensions | 16 |
| Future action window size | 6 |
| Past action window size | 0 |
| Action horizon | 7 |
| Repeated diffusion steps | 8 |
| Property | Value |
| Robot-action stream: | |
| AgiBotWorld-RL archives (Home + Industry) | 78 |
| AgiBotWorld-RL sample weight (per archive) | 1.0 |
| DROID sample weight | 150 |
| Composition (DROID AgiBotWorld-RL) | 65.8% : 34.2% |
| Action representation | delta-position, delta-rotation-vector |
| Hyperparameter | Value |
| Optimizer (AdamW): | |
| 0.9 | |
| 0.95 | |
| Weight decay | |
| Learning rate: |