Transferring heavy payloads in maritime settings relies on efficient crane operation, limited by hazardous double-pendulum payload sway. This sway motion is further exacerbated in offshore environments by external perturbations from wind and ocean waves. Manual suppression of these oscillations on an underactuated crane system by human operators is challenging. Existing control methods struggle in such settings, often relying on simplified analytical models, while deep reinforcement learning (RL) approaches tend to generalise poorly to unseen conditions. Deploying a predictive controller onto compute-constrained, highly non-linear physical systems without relying on extensive offline training or complex analytical models remains a significant challenge. Here we show a complete real-time control pipeline centered on the MuJoCo MPC framework that leverages a cross-entropy method planner to evaluate candidate action sequences directly within a physics simulator. By using simulated rollouts, this sampling-based approach successfully reconciles the conflicting objectives of dynamic target tracking and sway damping without relying on complex analytical models. We demonstrate that the controller can run effectively on a resource-constrained embedded hardware, while outperforming traditional PID and RL baselines in counteracting external base perturbations. Furthermore, our system demonstrates robustness even when subjected to unmodeled physical discrepancies like the introduction of a second payload.
Figures & tables
Fig. 1 : Experimental setup of a boom-type crane mounted on a motion platform.
Fig. 2 : MJPC crane controller architecture. Right: the physical crane with its frames of reference (global, base, crane, boom tip, payload) and joint positions ( qbase , qcrane , qpayload ). Left: at each step, the state estimates of qcrane and qpayload and the predicted qbase sequence are fed to the CEM planner in MJPC, which samples candidate policies (white), rolls them forward through simulation, and selects the best (purple). The immediate action of the best policy is issued as the crane control command; an operator can override the planner via joystick ( ⊗ ).
Parameter
Value
Unit
Boom length
2.384
m
Slew range
[ − 1.5, 1.5]
rad
Luff range
[0, 0.95]
rad
Hoist cable range
[0.07, 2.0]
m
Slew velocity range
± 0.92
rad/s
Luff velocity range
± 0.48
rad/s
TABLE I : Physical measurements of the crane and motion platform
Joint
Kv
Iarm
Damping
RMSE
Slew
7800
1000
0.01
1.5×10−3 rad
Luff
13000
2200
0.01
9.5×10−5 rad
Hoist †
25000
3200
0.0
1.1×10−2 m
† Add frictionloss = 30 to prevent payload slip
TABLE II : Identified actuator and joint parameters, with RMSE against measured crane trajectories.
Parameter
Value
Horizon length ( H )
80 steps (0.8 s)
Model time-step ( Δt )
0.01 s
Planner iterations ( I )
5
Trajectory sample size ( N )
20
Elite count ( M )
5
Sampling noise ( σ )
0.2
TABLE III : CEM hyperparameter values
Term
Formulation
Weight
rtarget : target tracking
∥ppayload−ptarget∥2+ϵ2−ϵ
Aα(d)
with ϵ=0.05
rsway : sway damping
∣θsway∣2+δ2
2Aα(d)
with δ=2.0
rvel : relative velocity
∥p˙payload−p˙platform∥2
Bβ(d)
rctrl : control effort
∥uˉslew∥2,∥uˉluff∥2
1
TABLE IV : Cost function terms
Method
Static
Slow
Medium
Fast
Pos (m)
Ang (deg)
Pos (m)
Ang (deg)
Pos (m)
Ang (deg)
Pos (m)
Ang (deg)
RL
0.27 (0.40)
7.85 (3.89)
0.39 (0.19)
9.92 (2.33)
0.47 (0.12)
13.82 (2.80)
0.43 (0.12)
13.25 (2.01)
PID
0.08 (0.03)
1.42 (0.11)
0.22 (0.07)
1.75 (0.78)
0.16 (0.04)
3.66 (0.87)
0.20 (0.18)
3.61 (26.65)
MJPC (ours)
0.03 (0.01)
1.86 (1.32)
0.09 (0.01)
1.51 (0.34)
0.11 (0.02)
1.75 (0.78)
0.11 (0.02)
2.21 (1.24)
MJPC [edge] (ours)
0.03 (0.01)
1.44 (0.77)
0.06 (0.02)
2.15 (0.75)
0.07 (0.02)
2.57 (0.38)
0.10 (0.02)
3.47 (1.02)
RL Δ
0.42
3.04
0.35
3.86
-0.04
0.27
0.48
2.53
TABLE V : Summary statistics of mean position error and payload tilt over the 10–20 s station-keeping interval, over 10 replications. In the top four rows, bold marks results significantly better than both baselines; two bold values are jointly best, not significantly different from each other. The bottom three rows ( Δ ) report each controller’s error increase from its single-payload baseline under an unmodeled double payload, over four replications.
Fig. 3 : Payload position error from target (top row) and angular tilt (bottom row) over various motion platform settings. We plot the median (solid line) and IQR (shaded area) over 10 replications.
Fig. 4 : Mean nominal return over 50 replications for various CEM planner iterations I in simulation. We also plot the planning time per cycle for each iteration parameter.
Fig. 5 : Mean payload position error from target, on the physical crane, over 60 s interval for various CEM planning iterations. The range is capped at 10 iterations, beyond which planning time exceeds the 20 Hz minimum control period.
The 4th "AI Olympics with RealAIGym" competition, to be held at IJCAI-ECAI 2026 in Bremen, challenges participants to develop a global control policy for swinging up and stabilizing an underactuated two-link system in its upright position. In contrast to previous editions, participants develop and evaluate their control strategies directly on remotely accessible CloudPendulum hardware, with limited interaction time and without prior knowledge of the system's model parameters. This paper presents an optimal-control-based approach employing real-time nonlinear model predictive control implemented using sequential quadratic programming. The results demonstrate that the proposed SQP-based MPC controller achieves reliable swing-up and stabilization performance, while maintaining robustness against disturbances.
Nick Karydakis, Konstantinos Chatzilygeroudis
Laboratory of Automation and Robotics (LAR) in the Department of Electrical & Computer Engineering, University of Patras, GR-26504 Patras, Greece
This paper presents an efficient model predictive path integral (MPPI) control framework for systems with complex nonlinear dynamics. To improve the computational efficiency of classic MPPI while preserving control performance, we replace the nonlinear dynamics used for trajectory propagation with a learned linear deep Koopman operator (DKO) model, enabling faster rollout and more efficient trajectory sampling. The DKO dynamics are learned directly from interaction data, eliminating the need for analytical system models. The resulting controller, termed MPPI-DK, is evaluated in simulation on pendulum balancing and surface vehicle navigation tasks, and validated on hardware through reference-tracking experiments on a quadruped robot. Experimental results demonstrate that MPPI-DK achieves control performance close to MPPI with true dynamics while substantially reducing computational cost, enabling efficient real-time control on robotic platforms.
Real-world reinforcement learning (RL) offers a promising route to dexterous manipulation policies that can adapt directly from physical interaction, but learning is hindered by inefficient early exploration and costly failures. We propose a framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions. A small set of MPC trajectories is first used to populate an offline replay buffer and to pretrain the actor and critic. During online learning, MPC intermittently guides data collection while an off-policy Soft Actor-Critic learner trains from both prior MPC experience and newly collected physical interaction, with control gradually transitioning to the learned policy. On continuous in-hand rotation with a 16-DoF Allegro hand, initialized from 20 MPC trajectories collected in 12 minutes on hardware, the policy reaches 100% success after 7 minutes of online RL, with about three object drops on average during training. After 20 minutes of online learning, the policy achieves more than five times the rotation speed of the MPC controller and completes 1000 consecutive rotations without a drop. Ablations show complementary benefits from MPC-based pretraining, retained MPC experience, and online MPC guidance, while additional experiments demonstrate rapid adaptation to new object geometries and successful goal-conditioned reorientation.
Emek Barış Küçüktabak, Karankumar Patel, Zhaodong Yang +3
Honda Research Institute USA, San Jose, CA, USA · Skild AI, San Mateo, CA, USA · Georgia Institute of Technology, Atlanta, GA, USA