Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by 2.88 points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
Figures & tables
Figure 1: A real unfiltered rollout that GeoMean accepts and CARM rejects. This 8,619-token response has scores sGeoMean=1.0025≤1.005 and sCARM=1.0248>1.02 . Top: positive and negative log-ratio mass and prefix deviations. Bottom: a 103-token window; red tokens have logrt>0 , green tokens have logrt<0 , and color intensity encodes magnitude.
Figure 2: CARM detects bidirectional token shifts that geometric-mean masking can hide. GeoMean averages signed log-ratios before taking an absolute value; CARM averages their magnitudes. The response-level gate complements the token-level PPO/GRPO surrogate.
Figure 3: Paired deviations (left) and rejection rates at shared thresholds (right) on 1,024 unfiltered responses. Every response satisfies dCARM≥dGeoMean . Green shading denotes rejection by both rules; orange shading denotes rejection only by CARM.
Model
Method
AIME 2024
AIME 2025
AIME 2026
BeyondAIME
Average
Qwen3.5-4B
Null
49.38
48.33
54.79
22.25
43.69
IcePop
47.08
40.83
40.42
19.31
36.91
TRM
66.67
55.21
65.00
32.31
54.80
GeoMean
72.50
64.38
70.62
38.88
61.59
CARM (ours)
77.08
64.38
69.79
42.56
63.45
Qwen3.5-9B
Null
76.67
66.88
67.71
45.25
64.13
Table 1: Mathematical-reasoning mean@16 across methods (%, higher is better). Average is the mean of AIME 2024, AIME 2025, AIME 2026, and BeyondAIME. Bold and underline denote the best and second-best values within each model block.
Method
TACO
LiveCodeBench-v6
HumanEval+
MBPP+
Average
Null
34.50
38.46
85.57
75.22
58.44
IcePop
45.87
43.36
92.68
79.81
65.43
TRM
36.60
44.97
87.80
76.72
61.52
GeoMean
46.03
44.50
92.28
78.31
65.28
CARM (ours)
46.70
51.20
94.11
81.22
68.31
Table 2: Evaluation pass@1 across code benchmarks (%, higher is better). TACO denotes the held-out test split, separate from the training split. Average is the mean of the four datasets. Bold and underline denote the best and second-best values.
Figure 4: Sequence filtering versus AIME accuracy. Each point is a masking configuration. AIME mean@16 averages the 2024/2025/2026 benchmarks.
Figure 7
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Value
Base models
Qwen3.5-4B and Qwen3.5-9B (math); Qwen3.5-9B (code)
Framework / algorithm
verl ; GRPO-style optimization
Training data
DAPO Math (math); TACO training split with executable rewards (code)
Evaluation data
AIME 2024/2025/2026 and BeyondAIME; TACO test split, LiveCodeBench-v6, HumanEval+, and MBPP+
Evaluation sampling
16 responses/problem (math); one response/problem per evaluation (code)
Prompt batch / rollout group
128 prompts / 8 responses per prompt
Appendix
Table 3: Complete experimental configuration for mathematical reasoning and code generation.
Qwen3.5-4B
Qwen3.5-9B
Method
τ
Avg.
Method
τ
Avg.
Null
–
50.83
Null
–
70.42
GeoMean
1.0025
69.03
GeoMean
1.0025
71.67
GeoMean
1.005
69.17
GeoMean
1.005
74.03
GeoMean
1.01
66.74
GeoMean
1.01
69.65
GeoMean
1.02
52.57
GeoMean
1.02
71.67
Appendix
Table 4: Threshold sweep (%, higher is better). Each row reports the three-year AIME mean@16 for that threshold. Bold and underline denote the best and second-best values within each model block.
Figure 9: Sequence-masked fraction during training for the selected GeoMean, TRM, and CARM (ours) configurations. The left and right panels show Qwen3.5-9B mathematical-reasoning and code-generation training, respectively. Faint curves are raw measurements and opaque curves apply an exponential moving average with weight 0.8 .
Figure 10: Training diagnostics for Qwen3.5-9B mathematical reasoning. Panels show response length, policy entropy, and gradient norm across the five methods. Faint curves show raw measurements; opaque curves show exponential moving averages with weight 0.8 . Gradient norm is displayed on a logarithmic scale.
Figure 11: Code-generation training diagnostics: response length, policy entropy, and gradient norm for Qwen3.5-9B trained on TACO. Faint and opaque curves show raw and smoothed measurements, respectively, using the same smoothing and axis conventions as the mathematical-reasoning diagnostics.
Metric
Method
k=1
k=2
k=4
k=8
k=16
Best-of- k
Null
70.42
78.07
83.34
86.42
88.26
TRM
73.96
81.47
86.43
89.39
90.98
GeoMean
74.03
80.19
84.05
86.73
88.54
CARM (ours)
77.50
83.94
87.74
90.05
91.28
Worst-of- k
Null
70.42
62.70
55.30
49.04
44.38
TRM
73.96
66.49
59.27
52.99
48.00
Appendix
Table 5: Exact Qwen3.5-9B values underlying Figure 6 . All entries are equal-weighted three-year AIME averages in percent.
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each masked position and which positions to remask. We formalize this as a two-stage action MDP, showing that the policy gradient naturally decomposes into a token term and a masking term. Combining optimization of both terms leads to state-of-the-art outcomes on mathematical reasoning and coding benchmarks, with scores of 87.1% on GSM8K and 53.4% on MBPP.
Haran Raajesh, Kulin Shah, Adam Klivans +1
Department of Computer Science The University of Texas at Austin
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio for two criteria: a proximity criterion, which asks whether the policy has moved too far from the behavior policy, and a direction criterion, which asks whether the update pushes it farther away. Recent work DPPO improves the proximity criterion by replacing PPO's ratio-based test with a probability divergence between the behavior and training policies. However, its direction criterion is still inherited from PPO. A token can be masked only when the sampled-token importance ratio moves away from one. We observe that this ratio-based direction criterion is a single-sample proxy that can disagree in sign with the change of the divergence that defines the proximity criterion. We therefore propose the predictive divergence mask, which asks whether the next policy-gradient step will increase or decrease the same divergence used by the trust region. For the discrete softmax policies used in LLM RL, we derive this prediction in closed form. Because production rollout engines expose only a truncated (top-K) view of the vocabulary, we develop two lightweight top-K estimators for this prediction. Detailed analysis shows the divergence-based direction is better aligned with the realized change of the divergence than the sampled ratio, and the resulting masks improve RL training across model scales and precision settings.
Diffusion Large Language Models (dLLMs) have emerged as an efficient alternative to autoregressive models, yet aligning them via Reinforcement Learning (RL) requires likelihood surrogates estimated from masked reconstruction subproblems under a small Monte Carlo budget per rollout. Existing methods construct these subproblems by uniform random masking, leaving open the question of which subproblems to prioritize. We identify a systematic upstream/downstream structure in dLLM rollouts. Some tokens, when revealed, trigger large confidence changes in nearby undecoded positions; we call them upstream. Others induce only small local changes and are therefore downstream. We find masking downstream tokens yields substantially better-posed subproblems than masking upstream tokens, a phenomenon we term subproblem difficulty asymmetry. Based on the observation, we propose Informed Masking (IM), which derives a per-token priority score from the denoising trajectory at zero extra inference cost and biases mask sampling toward downstream tokens. IM is plug-and-play: when plugged into three state-of-the-art dLLM RL methods on LLaDA-8B-Instruct, it delivers up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks with improved training stability.
Xiaoyi Yu, Enver Sangineto, Pei Fu +6
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · MiLM Plus, Xiaomi Inc., Beijing, China · University of Modena and Reggio Emilia, Italy +1