cs.ROSep 27, 2026

Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies

Authors: Jialeng Ni, Nathan Zhao, Kunpeng Song

Organizations: University of Michigan, Ann Arbor, MI, USA. · XPENG Robotics, Santa Clara, CA, USA.

Abstract

Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.

Figures & tables

Explore similar work

CardsList
  1. Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy

    Sep 11, 2026Pengfei Zhang, Teng Sun, Xianchao XiuRobot PoliciesLatent Action Models

  2. Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies

    May 31, 2026Hikmet Simsir, Ozgur S. OguzOffline Reinforcement LearningBehavior Cloning

  3. Steering Generative Reinforcement Learning into Stable Robotic Controller

    Jun 15, 2026Yixuan Wang, Shutong Ding, Ke Hu +3