cs.ROOct 1, 2026

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

Authors: Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger

Organizations: The University of Manchester, United Kingdom · BAE Systems, United Kingdom · University of Warwick, United Kingdom

Abstract

Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MAPL: Multi-Objective Preference Learning for Robot Locomotion

    Jun 24, 2026Xiyue Chen, Muhan Lin, Shuyang Shi +1Multi-Objective Reinforcement LearningPreference Learning

  2. Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications

    Jul 1, 2026Merve Atasever, Cagan Bakirci, Alfredo Reina Corona +2LocomotionGait Dynamics

  3. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    Jul 31, 2026Manith Adikari, Bei Peng, Samuele Vinanzi +1Multi-Objective Reinforcement LearningReinforcement Learning From Human Feedback