cs.LGOct 1, 2026

CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

Authors: Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang

Organizations: Moore Threads AI

Abstract

Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.133.13 percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by 2.882.88 points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mask-Aware Policy Gradients for Diffusion Language Models

    Jul 16, 2026Haran Raajesh, Kulin Shah, Adam Klivans +1Masked Diffusion Language ModelsLarge Language Model Reinforcement Learning

  2. Predictive Divergence Masks for LLM RL

    Jul 12, 2026Xiangxin Zhou, Jiarui Yao, Penghui Qi +4Large Language Model Reinforcement LearningDivergence

  3. Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

    Sep 22, 2026Xiaoyi Yu, Enver Sangineto, Pei Fu +6Diffusion Language ModelsLarge Language Model Reinforcement Learning