cs.ROSep 28, 2026

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

Authors: Yisheng Zhang, Tao Wang, Sicun Gao

Organizations: University of California San Diego, La Jolla, CA, USA · Tsinghua University, Beijing, China

Abstract

Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Model-Free Output Feedback Stabilization via Policy Gradient Methods

    Jan 27, 2026Ankang Zhang, Ming Chi, Xiaoling Wang +1Robust ControlStabilization

  2. Stability of Control Lyapunov Function Guided Reinforcement Learning

    May 3, 2026Zachary Olkin, William D. Compton, Aaron D. AmesReinforcement Learning ControlLyapunov Function