cs.CLOct 1, 2026

The Asymptotics of Language Model Alignment with Memory

Authors: Haricharan Balasundaram, V. Arvind Rameshwar

Organizations: School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA · IIT Madras · Department of Electrical Engineering, IIT Madras, Chennai, India

Abstract

Language model (LM) alignment broadly aims to perturb a given LM QQ into an aligned LM qq such that i) the outputs produced by qq and QQ are 'close' in probability, ii) qq has a higher expected reward than QQ. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-nn algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an mm--length i.i.d. token sequence output by the LM, in the limit as mm increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the mm--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when m=1m=1 -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

Figures & tables

Explore similar work

CardsList
  1. Theoretical Limits of Language Model Alignment

    May 8, 2026Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1Large Language Model AlignmentFundamental Limits

  2. Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models

    Sep 28, 2026Florian Eichin, Philipp Mondorf, Andrei Mircea +3MemorizationRecurrent Model