Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.
Figures & tables
Fig. 1: Principal steering subspaces. A one-time finite-difference probe estimates a response Gram matrix and partitions its eigenbasis into controlled directions Vk and complement V⊥ . During online learning, SAC emits zt , a fresh ϵt∼N(0,ID−k) fills the complement, and the frozen generator decodes wt=Vkzt+V⊥ϵt into an action chunk.
Fig. 2: Query schedules for training and deployment. (a) During online training, inference blocks execution so that st , the complete five-waypoint chunk, and st+1 define one unambiguous replay transition. (b) During deployment, the next query runs while the current chunk executes. The first two returned waypoints are no longer used under the measured query latency and are discarded; the aligned remainder is temporally aggregated with previous predictions before publication.
Fig. 3: Real-robot online RL interaction loop. The infer – execute cycle repeats until a terminal outcome is labeled. Scene reset and SAC optimization then run concurrently, and a join barrier admits the next ready state only after both branches finish.
Task
Transitions
Mean length
State dim.
Action dim.
Hidden MLP
Time dim.
Diff. train steps
Lift
31,127
104
19
7
512×3
16
20
Can
62,756
209
23
7
512×3
16
20
Square
80,731
269
23
7
1024×3 + cond.
32
100
Table I: Simulation datasets and configurations. All tasks use 300 multi-human demonstrations and action-chunk horizon H=4 . “cond.” denotes the Square observation-conditioning MLP with widths [512,64] ; time dim. is the time-embedding width.
Fig. 4: Decoder-response spectra. (a) Cumulative finite-difference response energy for the diffusion policy on Lift, Can, and Square. (b) The corresponding spectra for the flow-matching policy. In both panels, the dotted diagonal is the isotropic reference k/D , and the vertical line marks the common operating dimension k=8 .
Fig. 5: Online SAC with frozen diffusion policies. (a) Lift, (b) Can, and (c) Square compare full-latent DSRL with PSS at k=8 . Each point is a 200-episode evaluation of a checkpoint from one training seed; lines connect unsmoothed evaluations.
Fig. 6: Online SAC with frozen flow-matching policies. (a) Lift uses a 2×106 -step reporting horizon within longer runs. Each point is a 200-episode evaluation of a checkpoint from one training seed; lines connect unsmoothed evaluations.
Generator
Task
Full latent
PSS
Δ (p.p.)
Diffusion
Lift
0.981
0.981
+0.0
Diffusion
Can
0.831
0.884
+5.3
Diffusion
Square
0.716
0.878
+16.2
Flow
Lift
0.773
0.999
+22.6
Flow
Can
0.302
0.914
+61.2
Flow
Square
0.362
0.459
+9.7
Table II: Final-window simulation results. Entries are mean success rates over the final five evaluations at the reporting horizons defined above; Δ is PSS minus full latent in percentage points.
Fig. 7: Diffusion-Square design ablations. (a) The controlled dimension is fixed while the response-basis ranking is changed. (b) The leading response basis is retained while the controlled dimension is varied. The k=8 trace is shared by both panels. All curves show unsmoothed 200-episode evaluations from one matched seed through 4×106 environment steps.
Study
Setting
Final-five mean
Basis
Top- 8
0.838
Basis
Random- 8
0.346
Basis
Least- 8
0.498
Dimension
k=2
0.519
Dimension
k=4
0.714
Dimension
k=8
0.838
Table III: Final-window ablation summary. Each entry averages the final five evaluations at or before 4×106 steps in Fig. 7 ; the top- 8 run is common to both studies.
Fig. 8: Wall-clock profile of one 141-episode robot-training session. The upper bar partitions total session time into online interaction and the concurrent reset-and-learning window; the lower bar decomposes cumulative online-interaction time across the 141 episodes into motion, settling, inference, and environment/I/O work. Durations are measured from runtime logs and rounded, so the displayed components differ slightly in their totals.
Fig. 9: Representative screwdriver-placement trajectories. (a) The frozen VLA selects an ineffective approach/grasp and does not complete placement. (b) After online adaptation, SAC+ PSS steers the frozen decoder to grasp and place the screwdriver successfully. These sequences illustrate behavior; aggregate outcomes are reported in Table IV .
Policy
Successes/trials
Rate [95% CI]
Frozen VLA
2/10
0.20 [0.057, 0.510]
SAC+ PSS
6/10
0.60 [0.313, 0.832]
Table IV: Exploratory real-robot outcomes. Intervals are two-sided 95% Wilson intervals for evaluation-trial success. The small study is reported descriptively; it is not a statistically powered benchmark.