Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Figures & tables
Figure 1: Motivation for task-specific action-relevant RL token construction. The task-specific action-relevant information needed for action refinement and action-value estimation, which is represented by yellow highlight, is distributed across the tokens and layers of a frozen VLA. Existing methods that use fixed layers and tokens can miss features useful for the downstream task, whereas eRLT aggregates them across both dimensions to improve online RL sample efficiency.
Figure 2: Layer and token choices affect online RL. Controlled RLT comparisons on Drawer Opening and Bowl-to-Plate. (a) With all tokens retained, the best source layer differs across tasks. (b) At the final layer, selected token subsets can outperform all tokens, while the best subset also varies across tasks. Values report mean normalized success-rate AUC over three seeds; annotations show absolute improvements over the Final/All baselines.
Figure 3: Overview of eRLT. (a) Task-specific action-relevant multi-layer token routing constructs an RL token using learnable routing tokens and layer-wise routing. (b) Expert action prediction teaches the RL token to preserve differences needed for action refinement. (c) The critic uses successful and failed transitions to emphasize cues needed for action-value estimation.
Method
LIBERO
RoboTwin
All tasks
Drawer Opening
Bowl-to- Plate
Cabinet-Bowl to-Plate
Two-Moka Pots
Adjust Bottle
Move Can– Pot
Place Container Plate
Mean
DSRL ∗
0.856
0.801
0.130
0.244
0.804
0.608
0.649
0.585
RLT ∗
0.709
0.707
0.171
0.207
0.797
0.461
0.490
0.506
eRLT (ours)
0.881
0.818
0.247
0.264
0.897
0.523
0.752
0.626
Table 1: Main simulation results. Mean normalized success-rate AUC on four LIBERO and three RoboTwin tasks. The final column averages all seven tasks. Best and second-best results are bold and underlined, respectively.
Figure 4: Online learning curves in simulation. Evaluation success rate over online RL epochs on three representative LIBERO tasks (top) and three RoboTwin tasks (bottom). Dark curves show the smoothed trends, and light curves show the corresponding raw evaluations.
Table 6
Method
USB Connector
Ribbon Cable
AUC ↑
SR (%) ↑
AUC ↑
SR (%) ↑
Frozen VLA
–
30.00
–
70.00
DSRL ∗
0.266
40.00
0.272
50.00
RLT
0.269
50.00
0.484
100.00
eRLT (ours)
0.562
90.00
0.710
100.00
Table 4: Real-world results. Normalized AUC and final-window success rate (SR) under a fixed task-specific online interaction budget.
Figure 5: Real-world experimental setup and online learning curves. The RB-Y1 robot setup (left), representative execution sequences for USB connector insertion (top center) and motherboard ribbon-cable insertion (bottom center), and the 10-trajectory rolling success rate versus the number of training trajectories for both tasks (right). The shaded region marks the warm-up phase.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Lightweight online RL with a frozen VLA. A frozen VLA and a lightweight actor jointly generate executable action chunks. Online interactions train the lightweight critic and actor, and the RL token supplies task-specific action-relevant information to both modules.
Quantity
Shape
Construction
Frozen multimodal prefix
B×968×2048
3×256 image slots +200 language positions
Routing input
B×(968+K)×2048
Prefix followed by K routing tokens
Per-layer routing-token states
B×K×2048
Last K positions at one selected index
Per-layer summaries
B×m×2048
Mean over routing tokens, then stack m indices
Layer weights
m
Softmax of task-level logits with temperature τ
RL token zt
B×2048
Weighted sum over selected indices
Appendix
Table 5: Tensor shapes for the π0.5 eRLT routing module. B denotes batch size.
Figure 7: Simulation task overview. The four LIBERO tasks (top) and three RoboTwin tasks (bottom) used in our simulation experiments.
Task (suite)
Language instruction
Initial states
Horizon
Frozen SR
LIBERO
Drawer Opening (Goal)
Open the middle drawer of the cabinet.
0–49
240
0.74
Bowl-to-Plate (Spatial)
Pick up the black bowl next to the cookie box and place it on the plate.
300–349
240
0.78
Cabinet-Bowl-to-Plate (Spatial)
Pick up the black bowl on the wooden cabinet and place it on the plate.
450–499
240
0.12
Two-Moka-Pots (LIBERO-10)
Put both moka pots on the stove.
400–449
240
0.14
RoboTwin
Appendix
Table 6: Simulation tasks, horizons, and frozen-VLA success rates.
Configuration
LIBERO
RoboTwin
Demonstration data
40 episodes
2,500 trajectories from 50 tasks
Validation split
5% of frames
5 of 50 trajectories per task
Train/validation frames
6,300 / 332
62,799 / 7,008
Action target
10×7=70
32×14=448
Probe
2048 – 256 – 256 – 70
2048 – 256 – 256 – 448
Optimizer
AdamW
AdamW
Appendix
Table 7: Offline action-relevant initialization used in simulation.
Setting
Layers
D
K
τ
Router interval U
Router target rate
LIBERO
{0,3,6,9,12,15,18}
2048
1
1.0
200
0.005
RoboTwin
{0,3,6,9,12,15,18}
2048
1
1.0
200
0.005
Real world
{0,3,6,9,12,15,18}
2048
1
1.0
/
/
Appendix
Table 8: eRLT routing hyperparameters across experimental settings.
Task
Method
Total
Autonomous
Assisted
Autonomous successes
USB connector
DSRL ∗
80
57
23
20
USB connector
RLT
80
44
36
18
USB connector
eRLT
80
66
14
46
Ribbon cable
DSRL ∗
90
53
37
19
Ribbon cable
RLT
90
65
25
44
Ribbon cable
eRLT
90
70
20
63
Appendix
Table 9: Trajectory composition for each real-world experimental condition. Successes are counted over autonomous trajectories only.
Figure 8: Online optimization dynamics for USB connector insertion. Rows correspond to DSRL ∗ , RLT, and eRLT. Within each row, the panels show (a) actor and critic losses, (b) critic Q estimates for policy and replay actions, and (c) the weighted behavior-cloning and Q terms in the actor objective. Faint curves denote raw logged values and solid curves denote trailing-five-trajectory means.
Figure 9: Online optimization dynamics for ribbon-cable insertion. Rows correspond to DSRL ∗ , RLT, and eRLT, using the same metrics and smoothing convention as Figure 8 .