PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Organizations: The University of Manchester, United Kingdom · BAE Systems, United Kingdom · University of Warwick, United Kingdom
Abstract
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
Figures & tables
| Method | Tracking | Stability | Efficiency | Fall | cvar tilt |
| tar | |||||
| rma | |||||
| promo - sorl | |||||
| promo (ours) |
| Method | Tracking | Stability | Efficiency | Fall | Scalarized return | 4-anchor Ctrl. |
| dpmorl | ||||||
| moppo | ||||||
| Specialists | ||||||
| promo |
| Preference | Vel. RMSE [m/s] | Pos. MAE [m] | Energy [J/m] | cot [-] | Peak att. [rad] | Act. rate [rad/s] | Trq. rate [N m/s] |
| Replay-round | |||||||
| Balanced | |||||||
| Efficiency | |||||||
| Stability | |||||||
| Tracking | |||||||
| Move-fast | |||||||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Term | Group | Coefficient | Physical interpretation |
| track_lin_vel_xy_exp | Tracking | Rewards planar velocity tracking with kernel standard deviation | |
| track_ang_vel_z_exp | Tracking | Rewards yaw-rate tracking with kernel standard deviation | |
| lin_vel_z_l2 | Stability | Penalizes vertical base motion | |
| ang_vel_xy_l2 | Stability | Penalizes roll and pitch angular velocity | |
| flat_orientation_l2 | Stability | Penalizes body tilt away from upright posture | |
| joint_torques_l2 | Efficiency | Penalizes squared applied torque |
| Protocol | Preference specification | Purpose |
| Training distribution | , | Uniform truncated-simplex coverage without zeroed objectives |
| Balanced | Nominal semantic trade-off | |
| Tracking-heavy | Tracking-focused anchor | |
| Stability-heavy | Stability-focused anchor | |
| Efficiency-heavy | Efficiency-focused anchor | |
| Two-objective mixtures | , , | Representative simplex edges |
| Item | Value |
| Robot | Unitree Go2, - dof quadruped |
| Simulator | Isaac Sim / IsaacLab |
| Actor hidden sizes | |
| Critic shared trunk | |
| Semantic critic heads | scalar heads |
| Privileged encoder | MLP over privileged state and preference |
| Parameter | Range / setting |
| Static and dynamic friction | |
| Restitution coefficient | |
| Base mass offset | kg |
| Force perturbation | N |
| Torque perturbation | N m |
| Velocity perturbation | m/s every – s |
| Parameter | Value |
| Parallel environments | |
| Rollout horizon | policy steps per env |
| Discount factor | |
| gae parameter | |
| ppo clipping | |
| Entropy coefficient |
| Metric | Value |
| Mean utility per step | |
| cvar 10 utility | |
| Fall rate | |
| Preference–objective correlation | |
| tracking | |
| stability |
| threshold | |||||
| Non-dominated solutions | |||||
| Non-dominated set retained (%) |
| Regime | Command protocol | Evaluation focus |
| Replay | Recorded, identical | Command-matched comparison |
| Fast | Manual joystick | Dynamic locomotion |
| Slow | Manual joystick | Near-standstill locomotion |
| Variant | Reward | Ep. len. | Ctrl. | Fall | hv |
| Fixed priors (ours) | |||||
| Prior removed | |||||
| Priors inside objectives | |||||
| Prior as 4th objective |
| Type | Reward | Ep. len. | Ctrl. | cvar 10 | Fall | hv | |
| None | – | 26.8 | 997.3 | 0.80 | 0.012 | 0.10 | 0.56 |
| Action | 26.8 | 996.9 | 0.88 | 0.012 | 0.07 | 0.55 | |
| 26.6 | 994.9 | 0.78 | 0.012 | 0.07 | 0.74 | ||
| 26.9 | 997.8 | 0.79 | 0.012 | 0.08 | 0.66 | ||
| 26.8 | 996.4 | 0.86 | 0.011 | 0.11 | 0.59 | ||
| 26.5 | 996.7 | 0.48 | 0.009 | 0.13 | 0.85 |