cs.LGAug 19, 2026

Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

Authors: Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan

Organizations: State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, BIGAI · Beijing Institute of Technology · Institute for Artificial Intelligence, Peking University

Abstract

Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

    May 17, 2026Qiuhe Hong, Yuyang Liu, Shuo Yang +3Reinforcement Learning With Verifiable RewardReplay-Based Continual Learning

  2. Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

    Jun 2, 2026Chuanyu Qin, Chenxu Yang, Qingyi Si +3Reinforcement Learning With Verifiable RewardVerifiable Rewards