Abstract
Stochastic differential equations (SDEs) can be studied via Itô calculus and rough path theory. For stochastic optimal control, these two frameworks give distinct Pontryagin Maximum Principle (PMP) optimality conditions with forward-backward SDEs (FBSDEs) or rough differential equations. We show that the adjoint equations of the Itô and rough PMPs are connected via the conditional expectation ptItoˆ=E[ptrough∣Ft], where Ft represents information available at time t. First, we derive a rough stochastic PMP for problems with adapted controls that does not use FBSDEs. Its proof extends the rough stochastic PMP over deterministic controls by considering stochastic needle variations. Second, we derive a unified PMP connecting the Itô and rough PMPs, using Itô-Stratonovich conversion formulas and duality identities between the forward tangent and backward adjoint SDEs. As a first application, we rederive the adjoint matching method for fine-tuning generative models. As a second application, we propose an indirect shooting method for a class of feedback problems. Overall, these results give a new conditional bridge connecting two popular frameworks for stochastic optimal control.
Explore similar work
Feb 10, 2025math.OC
We derive first-order Pontryagin optimality conditions for stochastic optimal control with deterministic controls for systems modeled by rough differential equations (RDE) driven by Gaussian rough paths. This Pontryagin Maximum Principle (PMP) applies to systems following stochastic differential equations (SDE) driven by Brownian motion, yet it does not rely on forward-backward SDEs and involves the same Hamiltonian as the deterministic PMP. The proof consists of first deriving various integrable error bounds for solutions to nonlinear and linear RDEs by leveraging recent results on Gaussian rough paths. The PMP then follows using standard techniques based on needle-like variations. As an application, we propose the first indirect shooting method for nonlinear stochastic optimal control and show that it converges 10x faster than a direct method on a stabilization task.
Thomas Lew
Toyota Research Institute
Mar 28, 2026math.OC
Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints. In this work, we revisit and generalize Adjoint Matching, a recently proposed SOC-based method for learning optimal controls, and place it on a rigorous footing by deriving it from the Stochastic Maximum Principle (SMP). We formulate a general Hamiltonian adjoint matching objective for SOC problems with control-dependent drift and diffusion and convex running costs, and show that its expected value has the same first variation as the original SOC objective. As a consequence, critical points satisfy the Hamilton--Jacobi--Bellman (HJB) stationarity conditions. In the important practical case of state- and control-independent diffusion, we recover the lean adjoint matching loss previously introduced, which avoids second-order terms and whose critical points coincide with the optimal control under mild uniqueness assumptions. Numerical experiments confirm that the extra terms it discards become necessary once the diffusion is state-dependent. Finally, we show that adjoint matching can be precisely interpreted as a continuous-time method of successive approximations induced by the SMP, yielding a practical and implementable alternative to classical SMP-based algorithms, which are obstructed by intractable martingale terms in the stochastic setting. These results are also of independent interest to the stochastic control community, providing new implementable objectives and a viable pathway for SMP-based iterations in stochastic problems.
Carles Domingo-Enrich, Jiequn Han
Microsoft Research New England · Flatiron Institute
Aug 2, 2026math.OC
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We then prove that this formulation enjoys strong duality, thereby enabling us to reformulate the problem as an equivalent unconstrained one in the Lagrange dual framework. We propose a dual-ascent algorithm to solve the resulting problem and show that it converges to a deterministic Markov policy defined over the augmented state space that is both optimal and feasible. To accommodate continuous state-input spaces, we propose a dedicated learning algorithm to approximate the value function in an offline training setting, thereby significantly reducing the computational complexity of the online control phase. We then test our approach on a numerical example and demonstrate its effectiveness compared to online predictive control methods in terms of performance and computational complexity.
Francesco Cordiano, Kanghui He, Bart De Schutter
Delft Center for Systems and Control, Delft University of Technology, 2628 CD, Delft, The Netherlands · Department of Engineering Science, University of Oxford, OX1 3PJ Oxford, U.K.