cs.LGJul 8, 2026

Avoiding unsafe sets when training with Langevin Dynamics

Authors: Adam M. Oberman

Organizations: Department of Mathematics and Statistics, McGill University, Montreal, QC, Canada; Mila, Quebec AI Institute; and LawZero

Abstract

Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics, and a natural safety question is to bound the probability νt(AH)=P(QtAH)ν_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in \mathcal{A}_H) that the trajectory lies in a designated failure region AH\mathcal{A}_H. We study this for a smooth, strongly convex loss in dd dimensions, with AH\mathcal{A}_H separated from the minimizer by an energy gap. At the end of training, the equilibrium mass π(AH)π(\mathcal{A}_H) is exponentially small in dd, with a complementary energy-barrier rate when the noise is small. Along the trajectory, a shape-free bound νt(AH)π(AH)(1+χ02/π(AH)emt)ν_t(\mathcal{A}_H) \le π(\mathcal{A}_H)(1 + \sqrt{χ_0^2/π(\mathcal{A}_H)}\,e^{-mt}) shows the in-set probability relaxes to (twice) the static value after a burn-in of order dd, using only the global spectral gap mm. A worked Ornstein-Uhlenbeck example shows this burn-in is necessary: an angular slice of the equilibrium shell can transiently swell by a factor exponential in dd, though its equilibrium mass is tiny. To rule this out we introduce a local relaxation rate, defined through the spectral measure of the region's centered indicator rather than a Dirichlet-form Rayleigh quotient. For geometrically isolated regions this rate exceeds the global one, shrinking the burn-in, and with a maximum-principle ceiling it caps the trajectory probability uniformly in time. Strong convexity sets how fast training relaxes, but the shape of the unsafe set decides whether the trajectory bulges through it on the way to equilibrium.

Explore similar work

CardsList