Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by 2.88 points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
Figures & tables
Figure 1: A real unfiltered rollout that GeoMean accepts and CARM rejects. This 8,619-token response has scores sGeoMean=1.0025≤1.005 and sCARM=1.0248>1.02 . Top: positive and negative log-ratio mass and prefix deviations. Bottom: a 103-token window; red tokens have logrt>0 , green tokens have logrt<0 , and color intensity encodes magnitude.
Figure 2: CARM detects bidirectional token shifts that geometric-mean masking can hide. GeoMean averages signed log-ratios before taking an absolute value; CARM averages their magnitudes. The response-level gate complements the token-level PPO/GRPO surrogate.
Figure 3: Paired deviations (left) and rejection rates at shared thresholds (right) on 1,024 unfiltered responses. Every response satisfies dCARM≥dGeoMean . Green shading denotes rejection by both rules; orange shading denotes rejection only by CARM.
Model
Method
AIME 2024
AIME 2025
AIME 2026
BeyondAIME
Average
Qwen3.5-4B
Null
49.38
48.33
54.79
22.25
43.69
IcePop
47.08
40.83
40.42
19.31
36.91
TRM
66.67
55.21
65.00
32.31
54.80
GeoMean
72.50
64.38
70.62
38.88
61.59
CARM (ours)
77.08
64.38
69.79
42.56
63.45
Qwen3.5-9B
Null
76.67
66.88
67.71
45.25
64.13
Table 1: Mathematical-reasoning mean@16 across methods (%, higher is better). Average is the mean of AIME 2024, AIME 2025, AIME 2026, and BeyondAIME. Bold and underline denote the best and second-best values within each model block.
Method
TACO
LiveCodeBench-v6
HumanEval+
MBPP+
Average
Null
34.50
38.46
85.57
75.22
58.44
IcePop
45.87
43.36
92.68
79.81
65.43
TRM
36.60
44.97
87.80
76.72
61.52
GeoMean
46.03
44.50
92.28
78.31
65.28
CARM (ours)
46.70
51.20
94.11
81.22
68.31
Table 2: Evaluation pass@1 across code benchmarks (%, higher is better). TACO denotes the held-out test split, separate from the training split. Average is the mean of the four datasets. Bold and underline denote the best and second-best values.
Figure 4: Sequence filtering versus AIME accuracy. Each point is a masking configuration. AIME mean@16 averages the 2024/2025/2026 benchmarks.
Figure 7
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Value
Base models
Qwen3.5-4B and Qwen3.5-9B (math); Qwen3.5-9B (code)
Framework / algorithm
verl ; GRPO-style optimization
Training data
DAPO Math (math); TACO training split with executable rewards (code)
Evaluation data
AIME 2024/2025/2026 and BeyondAIME; TACO test split, LiveCodeBench-v6, HumanEval+, and MBPP+
Evaluation sampling
16 responses/problem (math); one response/problem per evaluation (code)
Prompt batch / rollout group
128 prompts / 8 responses per prompt
Appendix
Table 3: Complete experimental configuration for mathematical reasoning and code generation.
Qwen3.5-4B
Qwen3.5-9B
Method
τ
Avg.
Method
τ
Avg.
Null
–
50.83
Null
–
70.42
GeoMean
1.0025
69.03
GeoMean
1.0025
71.67
GeoMean
1.005
69.17
GeoMean
1.005
74.03
GeoMean
1.01
66.74
GeoMean
1.01
69.65
GeoMean
1.02
52.57
GeoMean
1.02
71.67
Appendix
Table 4: Threshold sweep (%, higher is better). Each row reports the three-year AIME mean@16 for that threshold. Bold and underline denote the best and second-best values within each model block.
Figure 9: Sequence-masked fraction during training for the selected GeoMean, TRM, and CARM (ours) configurations. The left and right panels show Qwen3.5-9B mathematical-reasoning and code-generation training, respectively. Faint curves are raw measurements and opaque curves apply an exponential moving average with weight 0.8 .
Figure 10: Training diagnostics for Qwen3.5-9B mathematical reasoning. Panels show response length, policy entropy, and gradient norm across the five methods. Faint curves show raw measurements; opaque curves show exponential moving averages with weight 0.8 . Gradient norm is displayed on a logarithmic scale.
Figure 11: Code-generation training diagnostics: response length, policy entropy, and gradient norm for Qwen3.5-9B trained on TACO. Faint and opaque curves show raw and smoothed measurements, respectively, using the same smoothing and axis conventions as the mathematical-reasoning diagnostics.
Metric
Method
k=1
k=2
k=4
k=8
k=16
Best-of- k
Null
70.42
78.07
83.34
86.42
88.26
TRM
73.96
81.47
86.43
89.39
90.98
GeoMean
74.03
80.19
84.05
86.73
88.54
CARM (ours)
77.50
83.94
87.74
90.05
91.28
Worst-of- k
Null
70.42
62.70
55.30
49.04
44.38
TRM
73.96
66.49
59.27
52.99
48.00
Appendix
Table 5: Exact Qwen3.5-9B values underlying Figure 6 . All entries are equal-weighted three-year AIME averages in percent.
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · MiLM Plus, Xiaomi Inc., Beijing, China · University of Modena and Reggio Emilia, Italy +1