cs.ROOct 6, 2026

Pareto-Optimal Entropy-Regularized Trajectory Optimization

Authors: Dimitrios S. Georgiou, Augustinos D. Saravanos, Evangelos A. Theodorou

Organizations: Daniel Guggenheim School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA. · Department of Aeronautics and Astronautics, Massachusetts Institute of Technology, Cambridge, MA, USA.

Abstract

Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature. For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins. Sampling-augmented variants mitigate this susceptibility through stochastic exploration, but often sample only around the few trajectories they retain for reoptimization, based solely on their cost which restricts exploration breadth. We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality. Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations. This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective. Across multiple systems and hundreds of environments, PER-DDP achieves higher success rates than state-of-the-art sampling-augmented TO methods and finds reliable solutions in environments beyond the reach of all baselines.

Figures & tables

Explore similar work

CardsList
  1. Beyond Pure Sampling: Hybrid Optimization Mechanisms for Non-Convex Model Predictive Control

    May 30, 2026Yuichiro Aoyama, Minchan Jung, Akash Ratheesh +1Model Predictive ControlModel-Free Reinforcement Learning

  2. Accelerating trajectory optimization with Sobolev-trained diffusion policies

    Apr 21, 2026Théotime Le Hellard, Franki Nguimatsia Tiofack, Quentin Le Lidec +1Robust TrajectoryDiffusion Policies

  3. Sampling-Based Control via Entropy-Regularized Optimal Transport

    May 4, 2026Vincent Pacelli, Akash Ratheesh, Evangelos A. TheodorouModel Predictive ControlTransport