Organizations: J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA · Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA · Zachry Department of Civil and Environmental Engineering, Texas A&M University, College Station, TX 77843, USA · Department of Electrical and Computer Engineering, Texas A&M University, College Station, TX 77843, USA
Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen π0.5 policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same π0.5 policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time. Project website is available at https://rollingoat.github.io/roboirs/.
Figures & tables
Fig. 1 : Real-robot setup under visual distribution shift (top) and baseline comparison (bottom).
Fig. 2 : Overview of the proposed RoboIRS framework. (a) RoboIRS intervenes between transformer blocks by adding a scaled steering vector to the selected residual-stream representation without updating the policy parameters. (b) To construct the steering vector, residual-stream representations are collected from successful and failed rollouts and used to train a linear classifier. The classifier weight provides the steering direction that is injected into the residual stream during inference.
Fig. 3 : Steering strength vs. SR from a few tasks to show the typical patterns: monotonic increase/decrease; increase at moderate strengths before degrading at larger strengths.
ID
Suite
Task
Unsteered
Steered
Sign
L1
goal_env
Open top drawer; put bowl inside
12.0±7.2
34.0±2.0
−1
L2
goal_env
Put bowl on plate
32.7±1.2
96.7±3.1
−1
L3
goal_env
Put bowl on cabinet
72.0±7.2
89.3±2.3
−1
L4
goal_env
Put wine bottle on cabinet
17.3±3.1
68.0±2.0
+1
L5
goal_swap
Put bowl on plate
70.0±4.0
66.7±9.9
+1
L6
object_env
Alphabet soup to basket
59.3±1.2
71.3±1.2
−1
TABLE I : Success rates (%) of unsteered and steered π0.5 . Bold indicates improvements larger than one unsteered standard deviation for LIBERO-PRO. The Sign column denotes the steering direction s . Parentheses in the real-robot rows denote successes/total trials.
Steering Strategy
Avg. SR (%)
One Random Direction
44.0
Unsteered
44.4
Worst Token per Layer
52.7
Highest-AUC Layer
54.3
Random Token per Layer CAA
56.5
Random Token per Layer
57.3
TABLE II : Ablation study over steering denoise steps, vectors, and locations, ordered by average success rate (SR).
ID
Un- steered
Grad.
Re-rank ( K=16 )
CAA
RoboIRS
L1
12.0
12.7
44.7
34.0
34.0
L2
32.7
30.7
74.0
68.7
96.7
L3
72.0
58.7
79.3
75.3
89.3
L4
17.3
19.3
51.3
94.0
68.0
L5
70.0
63.3
70.7
85.3
66.7
L6
59.3
56.7
79.3
66.0
71.3
TABLE III : Comparison of task-level success rates (%) and average inference time among baselines and RoboIRS.
Fig. 4 : t-SNE visualization of the last-layer’s residuals. Each panel compares residual representations from unsteered and steered policies separately for successful and failed rollouts.