cs.ROOct 1, 2026

eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing

Authors: Dehao Huang, Jianbang Liu, Jianpan Gao, Chao Tang, Zilang Cen, Zedong Dan, Jiaheng Wang, Tingguang Li, +2 more

Organizations: Southern University of Science and Technology, Shenzhen, China. · Beijing Zhongguancun Academy, Beijing, China. · Samsung Robotics eXperience. · Wuhan University, Wuhan, China. · Sun Yat-sen University, Guangzhou, China.

Abstract

Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RL Token: Bootstrapping Online RL with Vision-Language-Action Models

    Apr 24, 2026Charles Xu, Jost Tobias Springenberg, Michael Equi +4Diffusion-Based Vision-Language-ActionsReinforcement Fine-Tuning

  2. LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

    May 11, 2026Boyang Shen, Kaixiang Yang, Hao Wang +4Action PredictionPhotometric Supervision

  3. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Diffusion-Based Vision-Language-ActionsRobotic Manipulation