Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
Figures & tables
Figure 1: KL placement and seven types of mismatch. The right panel illustrates the coefficient decomposition ( di=⟨κ⟩i=Di , Kiκ=Ki ): a uniform shift represents the common KL offset, and the gold arrows indicate the relative calibration contribution. S denotes the base surrogate, sg stops gradients, r=πθ/πold is the probability ratio for a sampled token, and β≥0 is the KL weight; the independent term is illustrated using ℓKL=e−Δ+Δ−1 ( k3 ). Ai=Ri−Rˉ , A^i=Ai/(σq+εadv) , and sq=maxj∣Aj∣ ; R,σq,T,h,G,q denote the reward, reward standard deviation, length, prefix, group size, and prompt, respectively; overbars indicate group means, and εadv>0 . Conditions illustrated: F2 assumes the same context/token, all gates open, and equal weights; F3 assumes unfiltered groups; F4 assumes no length averaging and ℓi,tKL≈c>0 ; F5 assumes nonzero denominators; F7 assumes G≥2 independent single-token responses sampled from the current policy for the same prompt, with distinct, full-support πθ,πref . The coefficient objective first skips groups with G<2 or sq=0 .
Figure 2: ZCPO coefficient construction. Ai=Ri−Rˉ , Bq=1{G≥2,sq>0} . All tokens in a response share sg(Ci) . Zero-sum coefficients do not imply a zero total gradient; conditional KL eliminates only token sampling noise conditional on a given prefix.
Model/method
AIME24 (%)
AIME25 (%)
DAPO without KL ( β=0 )
47.3
37.1
DAPO + standard KL, β=0.04
31.2
18.9
DAPO + standard KL, β=0.001
44.5
34.7
ZCPO
56.8
48.5
No KL centering ( Ki←Di/Ti )
45.6
36.7
No KL length norm. ( Ki←Di−G1∑j=1GDj )
53.9
43.5
Table 1: Method comparisons, ablations, and F1 controls. Both AIME24 and AIME25 report avg@32 (%), with means over 3 random seeds. Each ZCPO ablation independently substitutes the definition shown in parentheses; all other training settings match full ZCPO.
Figure 3: KL failure-mode diagnostics. (a) F1: KL gradient share at clipped tokens. (b) F2: KL/total gradient norm ratios for groups in the top/bottom 10% by reward-gradient cancellation. (c) F4: at step 100, vanilla has mean lengths of 774 and 611 tokens at β=0.04 and 4 , versus 1695 for ZCPO at β=4 . (d) F5: KL gradient shares in the shortest/longest 10% of groups. (e,f) F6: KL signal concentration and first-token drift. (g,h) F7: within-group dispersion of mean response KL and negative-estimate fractions, both before centering. Dashed/dotted curves in (g) give small-KL sampling-noise reference scales; the dashed curve in (h) gives the normal approximation for the k1 control. Vanilla denotes DAPO with independent KL loss.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration item
Setting
Training framework
verl
Training data
DAPO-Math-17K: approximately 17K prompts with integer answers
Optimizer
AdamW
Learning rate
1×10−6
Learning-rate warmup
Linear warmup over the first 20 rollout steps
Learning-rate schedule after warmup
Constant
Appendix
Table 2: Base training and evaluation settings and KL coefficients for the main experiments.
β
AIME24 (%)
0
48.1
4
56.8
20
53.7
Appendix
Table 3: ZCPO sensitivity to β : AIME24 avg@32 (%) on Qwen2.5-32B, reporting means over 3 random seeds. The result at β=4 is reproduced from Table 1 .
Reward normalization
KL calibration off ( β=0 )
KL calibration on ( β=4 )
Standard deviation
47.3
55.4
Maximum absolute advantage
48.1
56.8
Appendix
Table 4: Joint comparison of reward normalization and KL-coefficient calibration. Entries are collected from Tables 1 and 3 , reporting Qwen2.5-32B AIME24 avg@32 (%) means over 3 random seeds.
Method
β
KL/total gradient norm ratio
vanilla
0.001
0.4%
vanilla
0.04
5.8%
ZCPO
4
7.0%
ZCPO
20
30.2%
Appendix
Table 5: Median ratio of the KL gradient norm to the total gradient norm in the diagnostic runs (steps 11–100).
Figure 4: Conditional sampling-noise comparisons. (a) Using the first mini-batch at each of the first 20 steps of ZCPO, the orange curve shows the residual RMSE of the k1 coefficient relative to ZCPO, and the green curve shows ZCPO’s own zero residual. ZCPO’s zero conditional sampling error refers only to the local estimation term obtained by replacing sampled Δi,t with κi,t for a given prefix; variation in sampled prefixes remains. (b) GVPO training curves. “Noise” denotes the sampling residual of the log-ratio relative to its per-prefix conditional mean. Its magnitude is the measured RMS of Dik1−Di=∑t(Δi,t−κi,t) after within-group centering, without multiplication by the KL weight. The green dashed line is the zero reference for conditional sampling error.
Model
DocMath
LBV2
Frames
MRCR
CorpusQA
LBV1-QA
Qwen3-4B-Thinking-2507
60.7
39.4
64.9
37.5
49.2
64.2
+ (GRPO + standard KL)
61.6
41.7
64.8
49.3
53.8
64.3
+ (GRPO without KL)
62.3
45.2
66.4
65.7
61.9
65.6
+ (GRPO-ZCPO)
63.8
48.2
68.2
66.0
69.6
66.4
+ (GRPO-ZCPO, Bq=1 )
63.3
48.0
67.6
66.3
68.5
66.1
Appendix
Table 6: Long-context evaluation results and group-gating ablation for Qwen3-4B. All methods trained in this work use the same data subset. Bq=1 denotes the ablation that removes gating for groups with identical rewards, with all other ZCPO settings unchanged.
Hyperparameter
Value
Data
Max prompt length
160K
Max response length
16K
Responses per prompt
16
Optimization
Learning rate
2e-6
Appendix
Table 7: Training hyperparameters for the long-context experiments.
Parameter
Setting
Model, Data, and Method
Starting model
Qwen3-4B-Thinking-2507
Training data
A fixed subset of 8,000 GoLongRL examples; the same subset for all methods
Training framework and entry point
verl; recipe.dapo.main_dapo
Trainer configuration
dapo_megatron_trainer
Update coefficient
GRPO uses group-relative advantages; ZCPO uses the joint coefficient Ci defined in the main text
Appendix
Table 8: Additional configuration for the Qwen3-4B long-context experiments. KL and group-gating settings that differ across methods are listed separately.
Model/method
AIME24 avg@32 (%)
GLM-4-9B-0414 + DAPO without KL
24.5
GLM-4-9B-0414 + ZCPO
36.2
Appendix
Table 9: Cross-model math results on GLM-4-9B-0414 (AIME24 avg@32, %).