Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.
Figures & tables
Fig. 1: Unsecured payload transportation over uneven terrain. The training task uses two unsecured stacked boxes on an edgeless carrier board. (A) Simulated training environment. (B, C) Real-world examples with the training payload configuration. (D–F) Real-world examples with payload configurations not used during training: a three-box stack in (D, E) and two side-by-side boxes in (F) .
Fig. 2: Overview of PAMORT’s two-stage training framework. In Phase 1, a preference-conditioned MO policy is trained with grouped rewards and sampled preferences. In Phase 2, the MO policy is frozen and the weight adjuster is trained to adapt the preference online. TAR [ 1 ] ’s representation and velocity-estimation details are omitted for clarity.
PAMORT
Reward term
Weight
TAR*
MO group
WA
Lin. velocity tracking
1.5
✓
move
✓
Ang. velocity tracking
0.75
✓
move
✓
Lin. velocity ( z )
−2.0
✓
move
–
Ang. velocity ( xy )
−0.05
✓
move
–
Joint velocity
−1.0×10−3
✓
move
–
TABLE I: Reward configurations
Parameter
Range
Unit
Foot friction (static, dynamic)
[0.3,1.2]2
–
Foot restitution
[0.0,0.15]
–
Base mass offset
[−1.0,3.0]
kg
Actuator & PD gain factors
[0.9,1.1]12×3
–
Action delay
[0.0,0.01]
s
Bottom box size (x,y,z)
[15,39]2×[7.5,45]
cm
TABLE II: Domain randomization settings
Fig. 3: Empirical Pareto front of undiscounted episodic returns for the move and payload reward groups (upper right is better). PAMORT’s MO policy yields non-dominated solutions for 0.4≤wpayload≤1.0 .
Terrain Type
Parameter
RandomRough
hmax=0.15dkm
CoarseBlocks
lcell=0.9m,hmax=0.15dkm
FineBlocks
lcell=0.45m,hmax=0.15dkm
SlopeUp/Down
s=0.6dk
WideStairsUp/Down
ltread=0.9m,hstep=0.225dkm
NarrowStairsUp/Down
ltread=0.3m,hstep=0.15dkm
TABLE III: Terrain families and parameters for simulation evaluation
Fig. 4: Representative level 6 simulation terrains. The ascending and descending variants share the same geometry and differ only in traversal direction.
Fig. 5: Success rates of TAR* and PAMORT across terrain difficulty levels under the two-payload training configuration. Each panel shows one of the nine one-way terrain types, with success rate over 100 trials at each level.
No payload
One payload
Two payloads (training setup)
Three payloads
Terrain
TAR*
PAMORT
TAR*
PAMORT
TAR*
PAMORT
TAR*
PAMORT
RandomRough
0.699
0.670
0.678
0.663
0.625
0.616
0.528
0.505
CoarseBlocks
0.942
0.969
0.907
0.947
0.867
0.922
0.718
0.763
FineBlocks
0.821
0.834
0.771
0.806
0.688
0.749
0.550
0.602
SlopeUp
0.496
0.548
0.418
0.471
0.367
0.403
0.289
0.314
WideStairsUp
0.695
0.737
0.648
0.694
0.585
0.660
0.461
0.492
TABLE IV: Mean simulation success rates of TAR* (extension of TAR [ 1 ] ) and PAMORT over levels 0–9 for different payload counts
Payload
Box
Mass [kg]
Static Friction
Size [cm]
Two
Top
1.24
0.39 (vs. bottom)
22×22×22
Bottom
2.33
0.35 (vs. board)
27×27×27
Three
Top
1.25
0.41 (vs. mid)
22×22×16
Mid
1.43
0.40 (vs. bottom)
34×26×18
Bottom
0.86
0.36 (vs. board)
28×27×20
TABLE V: Physical parameters of real-world payloads
Two payloads (training setup)
Three payloads
Terrain
TAR*
PAMORT
TAR*
PAMORT
38% Slope
0.6
1.0
0.7
0.9
44% Slope
0.5
0.9
0.6
0.9
16cm Step
0.9
1.0
0.7
0.8
20cm Step
0.6
0.6
0.8
0.7
Mean
0.650
0.875
0.700
0.825
TABLE VI: Real-world success rates over 10 trials per task
Fig. 6: Payload preference wpayload and torso pitch during a representative two-payload traversal on a 44% slope, which is slightly beyond the maximum training difficulty. Positive pitch denotes a nose-down posture, and t1 – t6 correspond to the snapshots.
Quadruped robots are increasingly expected to carry objects while moving through human environments. But what happens when a person interacts directly with the payload rather than with the robot? If the payload is unrestrained, the robot must distinguish intentional external interactions from ordinary payload motion, while still keeping the load balanced and maintaining stable locomotion. How can a quadruped infer and compliantly respond to such interactions using only onboard measurements? In this work, we develop a force-aware locomotion framework that treats payload interactions as commands that shape the motion of the combined robot-payload system. Our approach separates the learning of force-aware locomotion and force estimation on an unrestrained payload. We combine a compliant load-carrying policy with a causal force estimator, trained through estimator-in-the-loop data aggregation and finetuning, to predict interactions from onboard robot measurements. Our simulations and real-world experiments show that the resulting controller can maintain stable payload-carrying locomotion, yield compliantly to external interactions, and use the inferred force to support human-guided changes in the robot's trajectory.
Shaunak A. Mehta, Mayank Mishra, Prajit KrisshnaKumar +2
Fujitsu Research of America, Pittsburgh, PA, 15213.
In this paper, we present a robust nonprehensile object transportation framework for quadruped robots. An uncertainty-aware trajectory optimization method generates object motions with minimal closed-loop sensitivity to uncertain parameters. The resulting reference trajectory is tracked using a coupled convex model predictive controller that jointly predicts the CoM dynamics of the quadruped and the payload followed by a whole-body QP that enforces ground reaction constraints. The approach is evaluated through extensive simulations and real-world experiments under variations in the object's inertial parameters. Its performance is compared with fixed-orientation and straight-line trajectories as baseline. The results show that the optimized object motion reduces the sliding by approximately 50% compared with the fixed-orientation baseline and 30% compared with the straight-line baseline, while also achieving lower robot CoM tracking errors.
Ainoor Teimoorzadeh, Riccardo Pretto, Mario Selvaggio +2
Munich Institute of Robotics & Machine Intelligence, Technical University of Munich (TUM), Munich, Germany · Tampere University, Finland · PRISMA Lab, Department of Electrical Engineering and Information Technology, University of Naples Federico II, Via Claudio 21, 80125, Naples, Italy +1
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
Amr Mousa, Rifny Rachman, Neil Karavis +2
The University of Manchester, United Kingdom · BAE Systems, United Kingdom · University of Warwick, United Kingdom