Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
Figures & tables
Figure 1: A velocity penalty has task-dependent effects in direct trajectory optimization, shown by the mean stable dwell for IPOPT-optimized Acrobot and Pendubot trajectories. Moderate velocity regularization helps in different ranges for the two plants, while an overly large weight suppresses successful motion in both.
Task
Zeroth-order term r0(q)
Velocity error ev(v)
Acrobot
1−3[ec(θ1)+ec(θ2,a)]
θ˙12+θ˙22
Pendubot
1−3[ec(θ1)+ec(θ2,a)]
θ˙12+θ˙22
Cartpole
1−3ec(θ)
vc2+θ˙2
Double cart pendulum
2−3[ec(θ1)+ec(θ2,a)]
vc2+θ˙12+θ˙22
Quadcopter
1−3tanh(dp/0.8)−eR , eR=1−(q⊤q∗)2
∥vbody∥22
Franka reach
1−3tanh(dp/0.8)−3eq , eq=meanj(qj−qj∗)2
∥q˙∥22
Table 1: Reward family used in the Isaac Lab sweep. Every row has the form rβ=r0−βev . The β=0 condition removes velocity only from the reward; the policy still receives the full task-defined configuration and velocity observation.
Figure 2: Velocity-reward sweeps with full configuration and velocity observations. Blue solid curves show stable dwell and green dashed curves show success rate; bands are mean ± one population standard deviation over five seeds. Every task learns reaching, braking, and sustained stabilization at β=0 , while sufficiently large velocity penalties generally reduce both metrics.
Figure 3: No-velocity-observation ablation at β=0 . Each task keeps exactly the same zeroth-order reward and velocity-aware success predicate used in the full-observation experiment. Bars show the five-seed mean and population standard deviation of best dwell (left) and corresponding success rate (right).
Figure 4: Capping only the squared-velocity error prevents the large- β performance collapse. Five-seed beta sweeps for Pendubot (left) and Franka (right) compare the uncapped reward with the velocity cap c=4 in Equation 10 . Curves and bands are the mean ± one population standard deviation. The cap leaves the successful low- β regime intact and restores substantial performance where the raw reward fails.
Task
β
Reward
Dwell (s) ↑
Gt
∣mintGt∣
Value MSE ↓
Pendubot
0
Raw
6.87
31.9
9.10×102
3.20×102
1
Raw
0.018
−928
3.01×106
2.65×107
1
Clipped ( c=4 )
5.72
−245
1.15×103
1.04×103
Franka
0
Raw
14.58
94.5
69.1
0.495
1
Raw
0.008
−819
1.91×103
1.09×103
1
Clipped ( c=4 )
9.46
−99.2
1.87×102
1.49×102
Table 2: Reward clipping contracts return tails and critic error. Dwell is the best evaluation duration in each seed-42 run. The remaining columns are means over the final five matched one-iteration probes; Gt denotes the scalar return target at time step t .
Figure 5: Velocity capping contracts critic scale and restores target-centered structure. From left to right: failed raw Pendubot, successful velocity-capped Pendubot, failed raw Franka, and successful velocity-capped Franka, all at β=1 and with c=4 in the capped conditions using seed 42. Pendubot fixes angular velocities at zero, while Franka holds all remaining joints at a common target-conditioned reference. Red stars mark the upright or zero-error target.
Figure 6: Configuration completeness determines whether a zeroth-order reward is sufficient. From left to right: best stable dwell (s), success-ever rate, position–orientation pass rate PpR , and joint-velocity pass rate Pv . Full uses a target joint posture, whereas Partial specifies only the end-effector position and orientation. Curves and bands show the mean ± one population standard deviation over five seeds.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Environments.
Figure 8: Rendered views of the seven environments.
dp<0.10 m, ∥vbody∥2<0.20 m/s, ∥ωbody∥2<0.30 rad/s, tilt angle <0.15 rad.
Franka Reach (Full / Partial)
dp<0.05 m, dR<0.20 rad, ∥q˙∥2<0.25 rad/s.
Humanoid
dxy<0.35 m, ∥vroot∥2<0.35 m/s, ∥ωroot∥2<0.75 rad/s.
Appendix
Table 3: Instantaneous success predicates for all Isaac Lab evaluations. All conditions in a row must hold simultaneously.
Parameter
Six-task shared setting
Humanoid setting
Parallel environments
4096
4096
Rollout steps per environment
600
600
Transitions per PPO update
2,457,600
2,457,600
Training iterations
2300
1500
Checkpoint interval / final checkpoint
150 / 2299
150 / 1499
Training seeds
0,1,2,3,42
0,1,2,3,42
Appendix
Table 4: Rollout, runner, and actor–critic settings.
Parameter
Six-task shared setting
Humanoid setting
Optimizer
Adam
Adam
Initial learning rate
10−3
10−4
Learning-rate schedule
adaptive KL
adaptive KL
Desired KL
0.01
0.008
Policy ratio clip ϵ
0.2
0.2
Learning epochs per rollout
5
5
Appendix
Table 5: PPO objective and optimization settings.
Figure 9: Training return ( Equation 19 ) across the seven Isaac Lab tasks. Each seed is first smoothed with a centered 51-iteration moving average; curves and shaded regions then show the across-seed mean and one standard deviation, respectively.
Figure 10: Training value loss ( Equation 20 ) across the seven Isaac Lab tasks. Each seed is first smoothed with a centered 51-iteration moving average; curves and shaded regions then show the across-seed mean and one standard deviation, respectively.
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China · College of Automation, Nanjing University of Posts and Telecommunications, Nanjing 210023, China