cs.ROOct 8, 2026

Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer

Authors: Shuang Luo, Yilun Kong, Yunpeng Qing, Yihang Jiao, Zhi Hou, Shunyu Liu, Xiaogang Wang, Dacheng Tao

Organizations: Nanyang Technological University · ACE Robotics

Abstract

Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs). However, such VLM backbones offer insufficient physical dynamics priors, which limits the generalization capabilities of robot policies. Recent efforts therefore integrate video-generation World Models (WMs) into robot policies through various strategies, using predictive dynamics to facilitate action generation. Despite these advances, harnessing semantic understanding and dynamics prediction as complementary guidance for action generation remains challenging. In this paper, we introduce ACT3\mathrm{ACT}^3, a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control actions while preserving the distinct roles of context streams. Specifically, ACT3\mathrm{ACT}^3 enables the dedicated action expert to access VLM and WM representations through layerwise attention, with each backbone attending only within its own stream. This straightforward interaction design maintains independent forward propagation in the context streams while allowing both backbones to be updated through control supervision. Experiments on both simulated and real-world robotic manipulation benchmarks show that the proposed ACT3\mathrm{ACT}^3 yields results superior to its counterparts.

Explore similar work

CardsList
  1. PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

    May 20, 2026Shizhe Chen, Paul Pacaud, Cordelia SchmidLanguage-Conditioned Robot ManipulationPoint Cloud Learning

  2. TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation

    May 7, 2026Hanyu Zhou, Chuanhao Ma, Gim Hee LeeLanguage-Conditioned Robot ManipulationRobotic Manipulation

  3. Veo-Act: Enhancing VLA Policies with Frontier Video Models

    Apr 6, 2026Zhongru Zhang, Chenghan Yang, Qingzhou Lu +4Robotic ControlLanguage-Conditioned Robot Manipulation