cs.LGOct 8, 2026

When KL Regularization Misfires in Group Policy Optimization

Authors: Fei Ding

Organizations: Alibaba Group

Abstract

Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Group Adaptive Clipping Policy Optimization

    Aug 31, 2026Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan +1Group Relative Policy OptimizationPolicy Optimization

  2. Constrained Group Relative Policy Optimization

    Feb 5, 2026Roger Girgis, Rodrigue de Schaetzen, Luke Rowe +3Group Relative Policy OptimizationConstrained RL