Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.
Figures & tables
Fig. 1: Principal steering subspaces. A one-time finite-difference probe estimates a response Gram matrix and partitions its eigenbasis into controlled directions Vk and complement V⊥ . During online learning, SAC emits zt , a fresh ϵt∼N(0,ID−k) fills the complement, and the frozen generator decodes wt=Vkzt+V⊥ϵt into an action chunk.
Fig. 2: Query schedules for training and deployment. (a) During online training, inference blocks execution so that st , the complete five-waypoint chunk, and st+1 define one unambiguous replay transition. (b) During deployment, the next query runs while the current chunk executes. The first two returned waypoints are no longer used under the measured query latency and are discarded; the aligned remainder is temporally aggregated with previous predictions before publication.
Fig. 3: Real-robot online RL interaction loop. The infer – execute cycle repeats until a terminal outcome is labeled. Scene reset and SAC optimization then run concurrently, and a join barrier admits the next ready state only after both branches finish.
Task
Transitions
Mean length
State dim.
Action dim.
Hidden MLP
Time dim.
Diff. train steps
Lift
31,127
104
19
7
512×3
16
20
Can
62,756
209
23
7
512×3
16
20
Square
80,731
269
23
7
1024×3 + cond.
32
100
Table I: Simulation datasets and configurations. All tasks use 300 multi-human demonstrations and action-chunk horizon H=4 . “cond.” denotes the Square observation-conditioning MLP with widths [512,64] ; time dim. is the time-embedding width.
Fig. 4: Decoder-response spectra. (a) Cumulative finite-difference response energy for the diffusion policy on Lift, Can, and Square. (b) The corresponding spectra for the flow-matching policy. In both panels, the dotted diagonal is the isotropic reference k/D , and the vertical line marks the common operating dimension k=8 .
Fig. 5: Online SAC with frozen diffusion policies. (a) Lift, (b) Can, and (c) Square compare full-latent DSRL with PSS at k=8 . Each point is a 200-episode evaluation of a checkpoint from one training seed; lines connect unsmoothed evaluations.
Fig. 6: Online SAC with frozen flow-matching policies. (a) Lift uses a 2×106 -step reporting horizon within longer runs. Each point is a 200-episode evaluation of a checkpoint from one training seed; lines connect unsmoothed evaluations.
Generator
Task
Full latent
PSS
Δ (p.p.)
Diffusion
Lift
0.981
0.981
+0.0
Diffusion
Can
0.831
0.884
+5.3
Diffusion
Square
0.716
0.878
+16.2
Flow
Lift
0.773
0.999
+22.6
Flow
Can
0.302
0.914
+61.2
Flow
Square
0.362
0.459
+9.7
Table II: Final-window simulation results. Entries are mean success rates over the final five evaluations at the reporting horizons defined above; Δ is PSS minus full latent in percentage points.
Fig. 7: Diffusion-Square design ablations. (a) The controlled dimension is fixed while the response-basis ranking is changed. (b) The leading response basis is retained while the controlled dimension is varied. The k=8 trace is shared by both panels. All curves show unsmoothed 200-episode evaluations from one matched seed through 4×106 environment steps.
Study
Setting
Final-five mean
Basis
Top- 8
0.838
Basis
Random- 8
0.346
Basis
Least- 8
0.498
Dimension
k=2
0.519
Dimension
k=4
0.714
Dimension
k=8
0.838
Table III: Final-window ablation summary. Each entry averages the final five evaluations at or before 4×106 steps in Fig. 7 ; the top- 8 run is common to both studies.
Fig. 8: Wall-clock profile of one 141-episode robot-training session. The upper bar partitions total session time into online interaction and the concurrent reset-and-learning window; the lower bar decomposes cumulative online-interaction time across the 141 episodes into motion, settling, inference, and environment/I/O work. Durations are measured from runtime logs and rounded, so the displayed components differ slightly in their totals.
Fig. 9: Representative screwdriver-placement trajectories. (a) The frozen VLA selects an ineffective approach/grasp and does not complete placement. (b) After online adaptation, SAC+ PSS steers the frozen decoder to grasp and place the screwdriver successfully. These sequences illustrate behavior; aggregate outcomes are reported in Table IV .
Policy
Successes/trials
Rate [95% CI]
Frozen VLA
2/10
0.20 [0.057, 0.510]
SAC+ PSS
6/10
0.60 [0.313, 0.832]
Table IV: Exploratory real-robot outcomes. Intervals are two-sided 95% Wilson intervals for evaluation-trial success. The small study is reported descriptively; it is not a statistically powered benchmark.
Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.
Pengfei Zhang, Teng Sun, Xianchao Xiu
School of Mechatronic Engineering and Automation, Shanghai University, Shanghai 200444, China
Behavior cloning with high-capacity generative policies achieves strong imitation performance, but is often limited by demonstration coverage and distribution shift. Direct reinforcement learning fine-tuning can improve performance, but updating large action decoders is frequently unstable and sample inefficient. We propose Lagrangian Perturbation Diffusion Steering (LP-DS), a lightweight adaptation method that improves a frozen generative policy by learning a compact noise-space perturbation before decoding. LP-DS optimizes this perturbation with a Lagrangian trust-region objective, improving downstream value while constraining deviation from the latent prior. Across RoboMimic manipulation, OpenAI Gym locomotion, and Adroit dexterous manipulation benchmarks, LP-DS improves sample efficiency, success, and return while maintaining higher action-space entropy than unconstrained noise-space steering, with return improvements of up to 25% over prior baselines. Additional evaluations with flow-matching backbones, a large vision-language-action model, and physical Franka deployment show that LP-DS is not limited to compact diffusion policies or simulated benchmarks. Project page: https://sites.google.com/view/lp-ds/home.
Hikmet Simsir, Ozgur S. Oguz
Department of Computer Engineering, Bilkent University, Ankara, Türkiye.
Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation. However, the stochasticity of diffusion policies is not suitable for stable and precise control in high-dimensional robotic systems, where small action variations can accumulate into inconsistent motion and reduced robustness. To address this issue, we propose SteerGenPO, a latent-space reinforcement learning framework that steers a trained generative policy into a robust deterministic robotic controller. The key idea is to replace stochastic latent sampling of the trained generative policy with a learned latent actor that predicts a state-dependent latent input for the generative policies. This separates exploration and control: stochastic generative sampling provides diverse action proposals during policy learning, while deterministic latent steering provides stable and adaptive control at deployment. We evaluate SteerGenPO on six Isaac Lab benchmarks and a Unitree G1 locomotion task. The results show SteerGenPO improves over both classical RL and generative RL baselines, while its deterministic latent steering produces more stable inference-time behaviors and more reliable command responses.