Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR
Organizations: State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, BIGAI · Beijing Institute of Technology · Institute for Artificial Intelligence, Peking University
Abstract
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.
Figures & tables
| Reasoning tasks | Task sequence | RLVR training | CL methods | |
| TRACE ( Wang et al., 2023 ) | ✓ | ✓ | ✓ | |
| CLiMB ( Srinivasan et al., 2022 ) | ✓ | ✓ | ✓ | |
| Continual World ( Wolczyk et al., 2021 ) | ✓ | ✓ | ||
| LIBERO ( Liu et al., 2023 ) | ✓ | ✓ | ||
| Reasoning Gym ( Stojanovski et al., 2025 ) | ✓ | ✓ | ||
| VisuLogic ( Xu et al., 2026 ) | ✓ | ✓ |
| Metric | Seq. RLVR | Reset (ReDo) | Reset (FIRE) | Regularization (EWC) | Optimizer (Muon) | Isolation (OSFT) | Regularization (KL) | Replay (CPR) |
|---|---|---|---|---|---|---|---|---|
| FWT | 0.0 | -1.5 | -5.8 | -1.7 | +1.3 | -3.4 | +0.6 | +2.4 |
| TLG | 0.0 | -7.8 | -4.0 | +0.4 | -0.8 | -7.9 | -5.0 | +4.9 |
| BWT | 0.0 | +3.0 | -14.5 | +2.7 | +2.5 | +2.7 | +5.4 | -0.9 |
| Total | 0.0 | -6.3 | -24.4 | +1.4 | +3.0 | -8.5 | +1.1 | +6.4 |
| Mean CTM | 0.88 | 0.78 | 0.29 | 0.87 | 0.97 | 0.65 | 0.95 | 1.03 |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Modality | Stream | Ordered stages |
|---|---|---|
| LLM | Algorithmic Reasoning | Base Conversion Binary Alternation Binary Matrix Caesar Cipher Cryptarithm Isomorphic Strings Jugs Matrix Rotation String Manipulation A::B Rewriting |
| LLM | Algebra Reasoning | Complex Arithmetic Intermediate Integration Polynomial Equations Polynomial Multiplication Simple Equations Simple Integration |
| VLM | Quantitative Reasoning | Linear Quantity Prime Quantity Point Quantity Counter Quantity |
| VLM | Spatial Reasoning | Assembly Cross-sectional Hexahedron Three Views Spatial Orders Polyhedron Rotation |
| VLM | Positional Reasoning | Translation Rotation Comparative Flip |
| Modality | Setting | Stages | Total examples | Answer format |
|---|---|---|---|---|
| LLM | Algorithmic | 10 | 20,000 | Symbolic |
| LLM | Algebra | 6 | 20,000 | Symbolic |
| VLM | Quantitative | 4 | 1,386 | Four-way choice |
| VLM | Spatial | 6 | 1,043 | Four-way choice |
| VLM | Positional | 4 | 743 | Four-way choice |
| Total | 30 | — | — | |
| Group | Hyperparameter | Value |
| Model | LLM backbone | Qwen3-4B |
| VLM backbone | Qwen2.5-VL-7B | |
| RLVR | Advantage estimator | GRPO |
| Entropy coefficient | ||
| Rollouts per prompt | 5 | |
| PPO clip range |
| Method | Setting | Value |
| MTRL | Task access | Complete task pool available from the first update |
| LLM task sampling | ||
| CPR | Replay pool | Prompts from previous tasks |
| Replay fraction | ||
| Maximum replayed prompts per batch | for batch size | |
| Store capacity | One record per observed prompt with task identity and verifier metadata. Footprint reported in Appendix B.3 |
| Method | Setting | Value |
| EWC | Penalty coefficient | |
| Fisher decay | ||
| Numerical stabilizer | ||
| Fisher samples | All available samples | |
| FIRE | Newton–Schulz steps | 5 |
| Reset modules | Attention query and key projections |
| Setting | Per prompt (KiB) |
|---|---|
| LLM Algorithmic | 1.02 |
| LLM Algebra | 1.01 |
| VLM Quantitative | 1.42 |
| VLM Spatial | 1.39 |
| VLM Positional | 1.34 |
| Mean | 1.24 |
| Method | Mean runtime (h) |
|---|---|
| Seq. RLVR | 14.6 |
| CPR | 14.8 |
| Role | Method | Model | Steps | Base (%) | Final (%) | Final Base (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|---|---|---|
| baseline | Seq. RLVR | Qwen3-4B | 500 | 17.5 | 49.9 | +32.3 | +17.2 | +15.4 | +1.6 |
| MTRL reference | MTRL | Qwen3-4B | 500 | 17.5 | 65.3 | +47.7 | – | – | – |
| CL method | FIRE | Qwen3-4B | 500 | 17.5 | 5.6 | -11.9 | +10.0 | +12.9 | -37.5 |
| CL method | ReDo | Qwen3-4B | 500 | 17.5 | 26.8 | +9.2 | +9.5 | +4.6 | -4.4 |
| CL method | EWC | Qwen3-4B | 500 | 17.5 | 65.0 | +47.5 | +19.9 | +30.1 | -0.6 |
| CL method | Muon | Qwen3-4B | 500 | 17.5 | 57.4 | +39.9 | +20.0 | +23.7 | -2.0 |
| Role | Method | Model | Steps | Base (%) | Final (%) | Final Base (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|---|---|---|
| baseline | Seq. RLVR | Qwen3-4B | 500 | 43.2 | 88.2 | +45.1 | +20.3 | +29.1 | -1.1 |
| MTRL reference | MTRL | Qwen3-4B | 500 | 43.2 | 90.0 | +46.8 | – | – | – |
| CL method | CPR | Qwen3-4B | 500 | 43.2 | 90.4 | +47.2 | +19.6 | +30.9 | +0.1 |
| CL method | OSFT | Qwen3-4B | 500 | 43.2 | 88.4 | +45.2 | +23.6 | +21.9 | +4.3 |
| CL method | FIRE | Qwen3-4B | 500 | 43.2 | 80.3 | +37.1 | +15.0 | +33.0 | -10.0 |
| CL method | ReDo | Qwen3-4B | 500 | 43.2 | 81.7 | +38.5 | +19.1 | +23.2 | -0.6 |
| Role | Method | Model | Steps | Base (%) | Final (%) | Final Base (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|---|---|---|
| baseline | Seq. RLVR | Qwen2.5-VL-7B | 500 | 24.8 | 25.5 | +0.7 | +1.4 | +5.1 | -7.2 |
| MTRL reference | MTRL | Qwen2.5-VL-7B | 500 | 24.8 | 30.9 | +6.1 | – | – | – |
| CL method | CPR | Qwen2.5-VL-7B | 500 | 24.8 | 37.2 | +12.5 | +13.2 | +16.0 | -18.0 |
| CL method | ReDo | Qwen2.5-VL-7B | 500 | 24.8 | 25.4 | +0.6 | -1.4 | -3.0 | +6.2 |
| CL method | FIRE | Qwen2.5-VL-7B | 500 | 24.8 | 0.4 | -24.4 | +0.1 | -3.2 | -28.3 |
| CL method | EWC | Qwen2.5-VL-7B | 500 | 24.8 | 24.0 | -0.8 | +2.9 | -10.7 | +10.3 |
| Role | Method | Model | Steps | Base (%) | Final (%) | Final Base (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|---|---|---|
| baseline | Seq. RLVR | Qwen2.5-VL-7B | 500 | 24.1 | 26.3 | +2.2 | +7.4 | -2.7 | -1.5 |
| MTRL reference | MTRL | Qwen2.5-VL-7B | 500 | 24.1 | 31.6 | +7.5 | – | – | – |
| CL method | CPR | Qwen2.5-VL-7B | 500 | 24.1 | 32.5 | +8.4 | +9.9 | +4.0 | -4.6 |
| CL method | ReDo | Qwen2.5-VL-7B | 500 | 24.1 | 28.7 | +4.7 | +1.9 | -3.7 | +8.1 |
| CL method | FIRE | Qwen2.5-VL-7B | 500 | 24.1 | 2.5 | -21.6 | -1.7 | -6.5 | -16.4 |
| CL method | EWC | Qwen2.5-VL-7B | 500 | 24.1 | 24.7 | +0.6 | -7.7 | +8.7 | -2.0 |
| Role | Method | Model | Steps | Base (%) | Final (%) | Final Base (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|---|---|---|
| baseline | Seq. RLVR | Qwen2.5-VL-7B | 500 | 28.7 | 34.7 | +6.0 | -8.9 | +15.8 | -4.2 |
| MTRL reference | MTRL | Qwen2.5-VL-7B | 500 | 28.7 | 35.3 | +6.6 | – | – | – |
| CL method | CPR | Qwen2.5-VL-7B | 500 | 28.7 | 33.2 | +4.5 | -11.3 | +10.4 | +3.4 |
| CL method | ReDo | Qwen2.5-VL-7B | 500 | 28.7 | 30.5 | +1.7 | +0.7 | +2.3 | -1.6 |
| CL method | FIRE | Qwen2.5-VL-7B | 500 | 28.7 | 14.0 | -14.8 | -21.8 | +6.5 | -6.6 |
| CL method | Muon | Qwen2.5-VL-7B | 500 | 28.7 | 27.9 | -0.8 | -8.2 | +3.6 | +2.4 |
| Variant | FinalAvg (%) | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|
| No replay (Seq. RLVR) | 49.9 | +17.2 | +15.4 | +1.6 |
| Sample replay | 47.5 | +19.6 | +10.3 | +2.3 |
| CPR | 63.3 | +20.7 | +26.0 | +1.2 |
| MTRL reference | 65.3 | – | – | – |
| Sample replay no replay | -2.4 | +2.4 | -5.1 | +0.6 |
| CPR no replay | +13.5 | +3.5 | +10.7 | -0.4 |
| Setting | Replay pool | FinalAvg (%) | CTM | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|---|
| LLM algorithmic | No replay | 49.9 | 0.76 | |||
| LLM algorithmic | Immediately previous task | 66.3 | 1.02 | |||
| LLM algorithmic | All previous tasks (CPR) | 63.3 | 0.97 | |||
| LLM algorithmic | Previous and current tasks | 64.3 | 0.98 | |||
| LLM algebra | No replay | 88.2 | 0.98 | |||
| LLM algebra | All previous tasks (CPR) | 90.4 | 1.00 |
| Replay ratio | FinalAvg (%) | CTM | FWT (%) | TLG (%) | BWT (%) |
|---|---|---|---|---|---|
| (Seq. RLVR) | 49.9 | 0.76 | +17.2 | +15.4 | +1.6 |
| 64.5 | 0.99 | +22.7 | +28.9 | ||
| 63.2 | 0.97 | +20.1 | +29.2 | ||
| 63.3 | 0.97 | +20.7 | +26.0 | +1.2 | |
| MTRL reference | 65.3 | 1.00 | – | – | – |
| Cosine to | Inner product | ||||
| Stage | Current task | Current | Mixture | Current | Mixture |
| 2 | Binary Alternation | 0.410 | 0.628 | 1.122 | 5.421 |
| 3 | Binary Matrix | 0.467 | 0.679 | 7.811 | 6.616 |
| 4 | Caesar Cipher | 0.351 | 0.804 | 1.384 | 3.801 |
| 5 | Cryptarithm | 0.372 | 0.811 | 0.964 | 2.987 |
| 6 | Isomorphic Strings | 0.751 | 0.849 | 17.877 | 11.039 |
| Task | Policy | pass@1 | pass@4 | pass@8 | pass@16 |
|---|---|---|---|---|---|
| Base Conversion | Base | 57.7 | 81.0 | 87.5 | 90.6 |
| CPR + current | 98.1 | 99.8 | 100.0 | 100.0 | |
| Binary Alternation | Base | 0.6 | 1.8 | 2.7 | 3.1 |
| CPR + current | 43.1 | 55.4 | 60.1 | 64.1 | |
| Binary Matrix | Base | 17.9 | 27.7 | 31.7 | 34.4 |
| CPR + current | 51.0 | 56.6 | 58.9 | 60.9 |
| Setting | Opus 4.8 | GPT-5.5 | Qwen-32B | Qwen-72B | Base | Seq. RLVR | CPR | MTRL |
|---|---|---|---|---|---|---|---|---|
| LLM algorithmic | 53.7 | 77.7 | 29.3 | – | 17.5 | 49.9 | 63.3 | 65.3 |
| LLM algebra | 84.0 | 94.8 | 66.6 | – | 43.2 | 88.2 | 90.4 | 90.0 |
| VLM quantitative | 10.5 | 5.5 | 27.7 | 26.6 | 24.8 | 25.5 | 37.2 | 30.9 |
| VLM spatial | 19.9 | 0.8 | 27.7 | 28.9 | 24.1 | 26.3 | 32.5 | 31.6 |
| VLM positional | 19.1 | 8.6 | 27.7 | 25.8 | 28.7 | 34.7 | 33.2 | 35.3 |