Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Figures & tables
Figure 1: Motivation for task-specific action-relevant RL token construction. The task-specific action-relevant information needed for action refinement and action-value estimation, which is represented by yellow highlight, is distributed across the tokens and layers of a frozen VLA. Existing methods that use fixed layers and tokens can miss features useful for the downstream task, whereas eRLT aggregates them across both dimensions to improve online RL sample efficiency.
Figure 2: Layer and token choices affect online RL. Controlled RLT comparisons on Drawer Opening and Bowl-to-Plate. (a) With all tokens retained, the best source layer differs across tasks. (b) At the final layer, selected token subsets can outperform all tokens, while the best subset also varies across tasks. Values report mean normalized success-rate AUC over three seeds; annotations show absolute improvements over the Final/All baselines.
Figure 3: Overview of eRLT. (a) Task-specific action-relevant multi-layer token routing constructs an RL token using learnable routing tokens and layer-wise routing. (b) Expert action prediction teaches the RL token to preserve differences needed for action refinement. (c) The critic uses successful and failed transitions to emphasize cues needed for action-value estimation.
Method
LIBERO
RoboTwin
All tasks
Drawer Opening
Bowl-to- Plate
Cabinet-Bowl to-Plate
Two-Moka Pots
Adjust Bottle
Move Can– Pot
Place Container Plate
Mean
DSRL ∗
0.856
0.801
0.130
0.244
0.804
0.608
0.649
0.585
RLT ∗
0.709
0.707
0.171
0.207
0.797
0.461
0.490
0.506
eRLT (ours)
0.881
0.818
0.247
0.264
0.897
0.523
0.752
0.626
Table 1: Main simulation results. Mean normalized success-rate AUC on four LIBERO and three RoboTwin tasks. The final column averages all seven tasks. Best and second-best results are bold and underlined, respectively.
Figure 4: Online learning curves in simulation. Evaluation success rate over online RL epochs on three representative LIBERO tasks (top) and three RoboTwin tasks (bottom). Dark curves show the smoothed trends, and light curves show the corresponding raw evaluations.
Table 6
Method
USB Connector
Ribbon Cable
AUC ↑
SR (%) ↑
AUC ↑
SR (%) ↑
Frozen VLA
–
30.00
–
70.00
DSRL ∗
0.266
40.00
0.272
50.00
RLT
0.269
50.00
0.484
100.00
eRLT (ours)
0.562
90.00
0.710
100.00
Table 4: Real-world results. Normalized AUC and final-window success rate (SR) under a fixed task-specific online interaction budget.
Figure 5: Real-world experimental setup and online learning curves. The RB-Y1 robot setup (left), representative execution sequences for USB connector insertion (top center) and motherboard ribbon-cable insertion (bottom center), and the 10-trajectory rolling success rate versus the number of training trajectories for both tasks (right). The shaded region marks the warm-up phase.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Lightweight online RL with a frozen VLA. A frozen VLA and a lightweight actor jointly generate executable action chunks. Online interactions train the lightweight critic and actor, and the RL token supplies task-specific action-relevant information to both modules.
Quantity
Shape
Construction
Frozen multimodal prefix
B×968×2048
3×256 image slots +200 language positions
Routing input
B×(968+K)×2048
Prefix followed by K routing tokens
Per-layer routing-token states
B×K×2048
Last K positions at one selected index
Per-layer summaries
B×m×2048
Mean over routing tokens, then stack m indices
Layer weights
m
Softmax of task-level logits with temperature τ
RL token zt
B×2048
Weighted sum over selected indices
Appendix
Table 5: Tensor shapes for the π0.5 eRLT routing module. B denotes batch size.
Figure 7: Simulation task overview. The four LIBERO tasks (top) and three RoboTwin tasks (bottom) used in our simulation experiments.
Task (suite)
Language instruction
Initial states
Horizon
Frozen SR
LIBERO
Drawer Opening (Goal)
Open the middle drawer of the cabinet.
0–49
240
0.74
Bowl-to-Plate (Spatial)
Pick up the black bowl next to the cookie box and place it on the plate.
300–349
240
0.78
Cabinet-Bowl-to-Plate (Spatial)
Pick up the black bowl on the wooden cabinet and place it on the plate.
450–499
240
0.12
Two-Moka-Pots (LIBERO-10)
Put both moka pots on the stove.
400–449
240
0.14
RoboTwin
Appendix
Table 6: Simulation tasks, horizons, and frozen-VLA success rates.
Configuration
LIBERO
RoboTwin
Demonstration data
40 episodes
2,500 trajectories from 50 tasks
Validation split
5% of frames
5 of 50 trajectories per task
Train/validation frames
6,300 / 332
62,799 / 7,008
Action target
10×7=70
32×14=448
Probe
2048 – 256 – 256 – 70
2048 – 256 – 256 – 448
Optimizer
AdamW
AdamW
Appendix
Table 7: Offline action-relevant initialization used in simulation.
Setting
Layers
D
K
τ
Router interval U
Router target rate
LIBERO
{0,3,6,9,12,15,18}
2048
1
1.0
200
0.005
RoboTwin
{0,3,6,9,12,15,18}
2048
1
1.0
200
0.005
Real world
{0,3,6,9,12,15,18}
2048
1
1.0
/
/
Appendix
Table 8: eRLT routing hyperparameters across experimental settings.
Task
Method
Total
Autonomous
Assisted
Autonomous successes
USB connector
DSRL ∗
80
57
23
20
USB connector
RLT
80
44
36
18
USB connector
eRLT
80
66
14
46
Ribbon cable
DSRL ∗
90
53
37
19
Ribbon cable
RLT
90
65
25
44
Ribbon cable
eRLT
90
70
20
63
Appendix
Table 9: Trajectory composition for each real-world experimental condition. Successes are counted over autonomous trajectories only.
Figure 8: Online optimization dynamics for USB connector insertion. Rows correspond to DSRL ∗ , RLT, and eRLT. Within each row, the panels show (a) actor and critic losses, (b) critic Q estimates for policy and replay actions, and (c) the weighted behavior-cloning and Q terms in the actor objective. Faint curves denote raw logged values and solid curves denote trailing-five-trajectory means.
Figure 9: Online optimization dynamics for ribbon-cable insertion. Rows correspond to DSRL ∗ , RLT, and eRLT, using the same metrics and smoothing convention as Figure 8 .
Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires further fine-tuning -- for example, via reinforcement learning (RL). We introduce a lightweight method that enables sample-efficient online RL fine-tuning of pretrained VLAs using just a few hours of real-world practice. We (1) adapt the VLA to expose an "RL token," a compact readout representation that preserves task-relevant pretrained knowledge while serving as an efficient interface for online RL, and (2) train a small actor-critic head on this RL token to refine the actions, while anchoring the learned policy to the VLA. Online RL with the RL token (RLT) makes it possible to fine-tune even large VLAs with RL quickly and efficiently. Across four real-robot tasks (screw installation, zip tie fastening, charger insertion, and Ethernet insertion), RLT improves the speed on the hardest part of the task by up to 3x and raises success rates significantly within minutes to a few hours of practice. It can even surpass the speed of human teleoperation on some of the tasks.
Charles Xu, Jost Tobias Springenberg, Michael Equi +4
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic manipulation is composed of many frequent closed-loop spatial adjustments, for which excessive abstraction may waste computation and weaken low-level geometric cues essential for precise control. Existing early-exit strategies attempt to reduce computation by stopping at predefined layers or applying heuristic rules such as action consistency, but they do not directly answer when a representation is actually sufficient for action. In this paper, we present LoopVLA, a recurrent VLA architecture that jointly learns representation refinement, action prediction, and sufficiency estimation. LoopVLA iteratively applies a shared Transformer block to refine multimodal tokens, and at each iteration produces both a candidate action and a sufficiency score that estimates whether further refinement is necessary. By sharing parameters across iterations, LoopVLA decouples refinement from absolute layer indices and grounds sufficiency estimation in the evolving representation itself. Since sufficiency has no direct supervision, we introduce a self-supervised distribution alignment objective, where intermediate confidence scores are trained to match the relative action quality across refinement steps, thereby linking sufficiency learning to policy optimization signals. Experiments on LIBERO, LIBERO-Plus, and VLA-Arena show that LoopVLA pushes the efficiency-performance frontier of VLA policies, reducing parameters by 45% and improving inference throughput by up to 1.7 times while matching or outperforming strong baselines in task success.
Boyang Shen, Kaixiang Yang, Hao Wang +4
1Huazhong University of Science and Technology · 2Wuhan United Imaging Surgical Co.,Ltd. (UIS)
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading to high inference latency and increased training cost, and (ii) insufficient utilization of generated actions resulting in potential performance loss. To address these issues, we develop a training framework to finetune VLA models for generating significantly fewer action tokens with high parallelism, effectively reducing inference latency and training cost. Furthermore, we introduce an inference optimization technique with a novel voting-based ensemble strategy to combine current and previous action predictions, improving the utilization of generated actions and overall performance. Our results demonstrate that we achieve superior performance compared with state-of-the-art VLA models, achieving significantly higher success rates and 39× faster inference than OpenVLA with 46 Hz throughput on edge platforms, demonstrating practical deployability. The code is available at https://github.com/LukeLIN-web/VOTE.
Juyi Lin, Amir Taherin, Arash Akbari +11
Northeastern University, Boston, USA · EmbodyX,San Mateo,USA