cs.LGJun 30, 2026

On the Convergence of Self-Improving Online LLM Alignment

Authors: Xudong WuPangpang LiuVaneet AggarwalJiayu Chen

Organizations: The University of Hong Kong, Hong Kong SAR · Yale University, New Haven, CT, USA · Purdue University, West Lafayette, IN, USA

Abstract

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the Polyak-Lojasiewicz (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and LLM alignment tasks.

Explore similar work

CardsList
  1. Theoretical Limits of Language Model Alignment

    May 8, 2026Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1Large Language Model AlignmentBregman Divergences