Organizations: State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, BIGAI · Beijing Institute of Technology · Institute for Artificial Intelligence, Peking University
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.
Figures & tables
Figure 1: Shared reasoning explains modest forgetting while CPR harnesses it to improve performance. Later-task updates benefit earlier tasks on average, helping explain modest forgetting. CPR replaces some current-task prompts with previous-task prompts to harness shared reasoning and improve learning on the arriving and future tasks.
Reasoning tasks
Task sequence
RLVR training
CL methods
TRACE ( Wang et al., 2023 )
✓
✓
×
✓
CLiMB ( Srinivasan et al., 2022 )
✓
✓
×
✓
Continual World ( Wolczyk et al., 2021 )
×
✓
×
✓
LIBERO ( Liu et al., 2023 )
×
✓
×
✓
Reasoning Gym ( Stojanovski et al., 2025 )
✓
×
✓
×
VisuLogic ( Xu et al., 2026 )
✓
×
✓
×
Table 1: Positioning CRG at the intersection of CL and RLVR. CRG combines reasoning tasks, a task sequence, RLVR training, and multiple CL methods. ✓ = present, × = absent.
Figure 2: Policy training proceeds stage by stage through a sequence of verifiable reasoning tasks. Top : the LLM algorithmic sequence contains T=10 tasks, with jugs as the current task. Left : at each stage, the policy rolls out responses y1,…,yG to a prompt. A verifier scores the responses, and Group Relative Policy Optimization (GRPO) uses these scores to produce the stage policy Mt . Wheel : the five task sequences cover two text-reasoning settings and three visual-reasoning settings. Bottom : a jugs example shows the policy’s reasoning and verifier-checked answer.
Figure 3: Seq. RLVR remains below MTRL despite modest forgetting. Each marker represents one task sequence. Across task sequences, performance on earlier tasks falls by only 2.47 percentage points on average, yet final performance reaches 88% of MTRL.
Figure 4: Task gradients are positively aligned on average under Seq. RLVR. The heatmap shows pairwise cosine similarities between task gradients on the ten-task LLM algorithmic sequence. The mean cosine is 0.345 , and 82.2% of task pairs have positive cosine.
Figure 5: A jugs case illustrates shared reasoning in both performance and reasoning behavior. Jugs performance improves before its own stage and again during later-task training. The aligned excerpts show a progression from unresolved reasoning to valid and more direct solutions.
Metric
Seq. RLVR
Reset (ReDo)
Reset (FIRE)
Regularization (EWC)
Optimizer (Muon)
Isolation (OSFT)
Regularization (KL)
Replay (CPR)
FWT
0.0
-1.5
-5.8
-1.7
+1.3
-3.4
+0.6
+2.4
TLG
0.0
-7.8
-4.0
+0.4
-0.8
-7.9
-5.0
+4.9
BWT
0.0
+3.0
-14.5
+2.7
+2.5
+2.7
+5.4
-0.9
Total
0.0
-6.3
-24.4
+1.4
+3.0
-8.5
+1.1
+6.4
Mean CTM
0.88
0.78
0.29
0.87
0.97
0.65
0.95
1.03
Table 2: CPR is the only method to reach MTRL-level performance on average. For each method, FWT, TLG, and BWT report their percentage-point contributions to the change in FinalAvg relative to Seq. RLVR. Total is their sum.
Figure 6: Previous-task gradients better align updates with the full task set. For Qwen3-4B, each line compares a current-task gradient with an equal-weight combination of that gradient and the mean previous-task gradient. Dark diamonds show the averages across the nine stages.
Figure 7: Current-policy regeneration drives CPR’s replay gain. On LLM algorithmic reasoning, CPR approaches MTRL. Sample replay reuses trajectories generated by an earlier policy and underperforms no replay.
Table 3: Five task sequences define the Continual Reasoning Gym streams. Arrows show the training sequence. The LLM algorithmic, LLM algebra, VLM quantitative, VLM spatial, and VLM positional streams contain 10, 6, 4, 6, and 4 stages, respectively.
Modality
Setting
Stages
Total examples
Answer format
LLM
Algorithmic
10
20,000
Symbolic
LLM
Algebra
6
20,000
Symbolic
VLM
Quantitative
4
1,386
Four-way choice
VLM
Spatial
6
1,043
Four-way choice
VLM
Positional
4
743
Four-way choice
Total
30
—
—
Appendix
Table 4: Continual Reasoning Gym contains 30 stages across five task sequences. Each setting is split 8:2 into training and test sets. LLM examples are generated procedurally, with at most 20,000 total examples sampled for each text setting. Symbolic and four-way answers are scored by the corresponding task verifier.
Figure 8: Representative CRG tasks. One example is shown for each setting.
Group
Hyperparameter
Value
Model
LLM backbone
Qwen3-4B
VLM backbone
Qwen2.5-VL-7B
RLVR
Advantage estimator
GRPO
Entropy coefficient
1×10−3
Rollouts per prompt
5
PPO clip range
0.2
Appendix
Table 5: All runs use the same core RLVR configuration. The text experiments use Qwen3-4B, and the visual experiments use Qwen2.5-VL-7B. Training updates all model parameters, including the VLM vision tower.
Method
Setting
Value
MTRL
Task access
Complete task pool available from the first update
LLM task sampling
PMTRL(k)=1/T
CPR
Replay pool
Prompts from previous tasks
Replay fraction
ρ=0.5
Maximum replayed prompts per batch
min{⌊ρb⌋,⌊b/2⌋} for batch size b
Store capacity
One record per observed prompt with task identity and verifier metadata. Footprint reported in Appendix B.3
Appendix
Table 6: MTRL accesses the complete task pool from the first update, whereas CPR replays previous-task prompts. Both use the shared configuration in Table 5 .
Method
Setting
Value
EWC
Penalty coefficient
1.0
Fisher decay
1.0
Numerical stabilizer ϵ
10−8
Fisher samples
All available samples
FIRE
Newton–Schulz steps
5
Reset modules
Attention query and key projections
Appendix
Table 7: Each CL baseline changes a specific part of the shared training protocol. The table lists every baseline-specific override. All remaining settings follow Table 5 .
Figure 9: CPR turns previous-task prompts into current-policy RLVR updates. The top row contrasts stale trajectory replay with CPR’s response regeneration. The lower panel shows one CPR update during a later stage. CPR selects previous-task prompts, mixes them with current-task prompts, generates every response with the current policy, and scores each response with its task verifier before the GRPO update.
Setting
Per prompt (KiB)
LLM Algorithmic
1.02
LLM Algebra
1.01
VLM Quantitative
1.42
VLM Spatial
1.39
VLM Positional
1.34
Mean
1.24
Appendix
Table 8: Average serialized footprint per CPR replay record. Each row reports the mean over prompts in one CRG setting. The final row weights the five settings equally.
Method
Mean runtime (h)
Seq. RLVR
14.6
CPR
14.8
Appendix
Table 9: Average wall-clock time per 500-update run on eight NVIDIA B200 GPUs.
Figure 10: CPR is the only evaluated CL method that reaches MTRL-level performance on average. Bars show CTM for eight sequential methods in each setting. Seq. RLVR is dark gray. The six other interventions are light gray. CPR is blue. This level is reached or exceeded for LLM algebra, VLM quantitative reasoning, and VLM spatial reasoning. Its mean CTM is 1.03 , compared with 0.88 for Seq. RLVR.
Figure 11: CPR’s FinalAvg gains arise from FWT, TLG, or both. Each pair compares Seq. RLVR with CPR. Stacked bars show the four terms BaseAvg, TT−1FWT , TLG, TT−1BWT . Black markers show FinalAvg. The BWT effect varies across task sequences. On VLM quantitative reasoning, larger FWT and TLG outweigh a more negative BWT contribution.
Role
Method
Model
Steps
Base (%)
Final (%)
Final − Base (%)
FWT (%)
TLG (%)
BWT (%)
baseline
Seq. RLVR
Qwen3-4B
500
17.5
49.9
+32.3
+17.2
+15.4
+1.6
MTRL reference
MTRL
Qwen3-4B
500
17.5
65.3
+47.7
–
–
–
CL method
FIRE
Qwen3-4B
500
17.5
5.6
-11.9
+10.0
+12.9
-37.5
CL method
ReDo
Qwen3-4B
500
17.5
26.8
+9.2
+9.5
+4.6
-4.4
CL method
EWC
Qwen3-4B
500
17.5
65.0
+47.5
+19.9
+30.1
-0.6
CL method
Muon
Qwen3-4B
500
17.5
57.4
+39.9
+20.0
+23.7
-2.0
Appendix
Table 10: Complete results for LLM algorithmic reasoning. Base reports pre-training accuracy. Final reports endpoint accuracy. Final − Base, FWT, TLG, BWT use percentage points. The largest reported value in each performance column is bold.
Role
Method
Model
Steps
Base (%)
Final (%)
Final − Base (%)
FWT (%)
TLG (%)
BWT (%)
baseline
Seq. RLVR
Qwen3-4B
500
43.2
88.2
+45.1
+20.3
+29.1
-1.1
MTRL reference
MTRL
Qwen3-4B
500
43.2
90.0
+46.8
–
–
–
CL method
CPR
Qwen3-4B
500
43.2
90.4
+47.2
+19.6
+30.9
+0.1
CL method
OSFT
Qwen3-4B
500
43.2
88.4
+45.2
+23.6
+21.9
+4.3
CL method
FIRE
Qwen3-4B
500
43.2
80.3
+37.1
+15.0
+33.0
-10.0
CL method
ReDo
Qwen3-4B
500
43.2
81.7
+38.5
+19.1
+23.2
-0.6
Appendix
Table 11: Complete results for LLM algebra. Base reports pre-training accuracy. Final reports endpoint accuracy. Final − Base, FWT, TLG, BWT use percentage points. The largest reported value in each performance column is bold.
Role
Method
Model
Steps
Base (%)
Final (%)
Final − Base (%)
FWT (%)
TLG (%)
BWT (%)
baseline
Seq. RLVR
Qwen2.5-VL-7B
500
24.8
25.5
+0.7
+1.4
+5.1
-7.2
MTRL reference
MTRL
Qwen2.5-VL-7B
500
24.8
30.9
+6.1
–
–
–
CL method
CPR
Qwen2.5-VL-7B
500
24.8
37.2
+12.5
+13.2
+16.0
-18.0
CL method
ReDo
Qwen2.5-VL-7B
500
24.8
25.4
+0.6
-1.4
-3.0
+6.2
CL method
FIRE
Qwen2.5-VL-7B
500
24.8
0.4
-24.4
+0.1
-3.2
-28.3
CL method
EWC
Qwen2.5-VL-7B
500
24.8
24.0
-0.8
+2.9
-10.7
+10.3
Appendix
Table 12: Complete results for VLM quantitative reasoning. Base reports pre-training accuracy. Final reports endpoint accuracy. Final − Base, FWT, TLG, BWT use percentage points. The largest reported value in each performance column is bold.
Role
Method
Model
Steps
Base (%)
Final (%)
Final − Base (%)
FWT (%)
TLG (%)
BWT (%)
baseline
Seq. RLVR
Qwen2.5-VL-7B
500
24.1
26.3
+2.2
+7.4
-2.7
-1.5
MTRL reference
MTRL
Qwen2.5-VL-7B
500
24.1
31.6
+7.5
–
–
–
CL method
CPR
Qwen2.5-VL-7B
500
24.1
32.5
+8.4
+9.9
+4.0
-4.6
CL method
ReDo
Qwen2.5-VL-7B
500
24.1
28.7
+4.7
+1.9
-3.7
+8.1
CL method
FIRE
Qwen2.5-VL-7B
500
24.1
2.5
-21.6
-1.7
-6.5
-16.4
CL method
EWC
Qwen2.5-VL-7B
500
24.1
24.7
+0.6
-7.7
+8.7
-2.0
Appendix
Table 13: Complete results for VLM spatial reasoning. Base reports pre-training accuracy. Final reports endpoint accuracy. Final − Base, FWT, TLG, BWT use percentage points. The largest reported value in each performance column is bold.
Role
Method
Model
Steps
Base (%)
Final (%)
Final − Base (%)
FWT (%)
TLG (%)
BWT (%)
baseline
Seq. RLVR
Qwen2.5-VL-7B
500
28.7
34.7
+6.0
-8.9
+15.8
-4.2
MTRL reference
MTRL
Qwen2.5-VL-7B
500
28.7
35.3
+6.6
–
–
–
CL method
CPR
Qwen2.5-VL-7B
500
28.7
33.2
+4.5
-11.3
+10.4
+3.4
CL method
ReDo
Qwen2.5-VL-7B
500
28.7
30.5
+1.7
+0.7
+2.3
-1.6
CL method
FIRE
Qwen2.5-VL-7B
500
28.7
14.0
-14.8
-21.8
+6.5
-6.6
CL method
Muon
Qwen2.5-VL-7B
500
28.7
27.9
-0.8
-8.2
+3.6
+2.4
Appendix
Table 14: Complete results for VLM positional reasoning. Base reports pre-training accuracy. Final reports endpoint accuracy. Final − Base, FWT, TLG, BWT use percentage points. The largest reported value in each performance column is bold.
Variant
FinalAvg (%)
FWT (%)
TLG (%)
BWT (%)
No replay (Seq. RLVR)
49.9
+17.2
+15.4
+1.6
Sample replay
47.5
+19.6
+10.3
+2.3
CPR
63.3
+20.7
+26.0
+1.2
MTRL reference
65.3
–
–
–
Δ Sample replay − no replay
-2.4
+2.4
-5.1
+0.6
Δ CPR − no replay
+13.5
+3.5
+10.7
-0.4
Appendix
Table 15: Current-policy regeneration drives CPR’s replay gain. On LLM algorithmic reasoning at 500 steps and ρ=0.5 , sample replay reuses earlier-policy trajectories with importance sampling. CPR regenerates responses with the current policy. Delta rows are relative to no replay and are computed from unrounded values. The best value among the three sequential variants in each metric is bold.
Setting
Replay pool
FinalAvg (%)
CTM
FWT (%)
TLG (%)
BWT (%)
LLM algorithmic
No replay
49.9
0.76
+17.2
+15.4
+1.6
LLM algorithmic
Immediately previous task
66.3
1.02
+20.8
+30.1
+0.0
LLM algorithmic
All previous tasks (CPR)
63.3
0.97
+20.7
+26.0
+1.2
LLM algorithmic
Previous and current tasks
64.3
0.98
+20.9
+29.1
−1.3
LLM algebra
No replay
88.2
0.98
+20.3
+29.1
−1.1
LLM algebra
All previous tasks (CPR)
90.4
1.00
+19.6
+30.9
+0.1
Appendix
Table 16: Replaying prompts from all previous tasks gives the highest mean FinalAvg. No replay corresponds to Seq. RLVR. Replay from only the immediately previous task is reported for LLM algorithmic reasoning. Bold values mark the maximum within each setting and among the mean rows.
Replay ratio ρ
FinalAvg (%)
CTM
FWT (%)
TLG (%)
BWT (%)
0.00 (Seq. RLVR)
49.9
0.76
+17.2
+15.4
+1.6
0.05
64.5
0.99
+22.7
+28.9
−2.6
0.20
63.2
0.97
+20.1
+29.2
−1.8
0.50
63.3
0.97
+20.7
+26.0
+1.2
MTRL reference
65.3
1.00
–
–
–
Appendix
Table 17: CPR is robust to replay ratio ρ on LLM algorithmic reasoning. With Qwen3-4B, every tested nonzero ratio reaches CTM 0.97 – 0.99 , compared with 0.76 without replay. FinalAvg varies by only 1.3 percentage points across the nonzero ratios. The best sequential value in each metric is bold.
Figure 12: CPR gains accompany both entropy increases and decreases. Each row pairs CPR with Seq. RLVR in the same setting. The left panel reports the ratio of final-policy generated-token entropy, measured on fixed prompts under greedy decoding. The right panel reports the corresponding CTM change. Dashed lines mark no change. Among the four settings where CPR improves CTM, entropy decreases in both LLM settings and increases in VLM quantitative and spatial reasoning.
Cosine to gˉ⋆
Inner product (10−5)
Stage
Current task
Current
Mixture
Current
Mixture
2
Binary Alternation
0.410
0.628
1.122
5.421
3
Binary Matrix
0.467
0.679
7.811
6.616
4
Caesar Cipher
0.351
0.804
1.384
3.801
5
Cryptarithm
0.372
0.811
0.964
2.987
6
Isomorphic Strings
0.751
0.849
17.877
11.039
Appendix
Table 18: Adding previous-task gradients improves alignment with the all-task mean at the Qwen3-4B base model. Each row uses the same fixed set of ten task gradients. For each stage, the mixture assigns half of its weight to the current task and distributes the other half equally across previous tasks. The all-task mean gˉ⋆ serves only as the evaluation direction. Inner products are reported in units of 10−5 .
Figure 13: LLM gradient alignment weakens across the algorithmic–algebra boundary. Pairwise cosines are shown for the ten algorithmic tasks and six algebra tasks. Black lines mark the domain boundary. Cross-domain task pairs have mean cosine −0.01 , while the cosine between the two domain-average vectors is 0.006 . The panel combines separately collected CPR diagnostics.
Figure 14: The measured VLM task gradients are positively aligned across all three domains. The matrix combines separately collected CPR diagnostics for quantitative, spatial, and positional reasoning. Every displayed off-diagonal task pair has positive cosine, including pairs that cross domain boundaries. Black lines mark those boundaries.
Task
Policy
pass@1
pass@4
pass@8
pass@16
Base Conversion
Base
57.7
81.0
87.5
90.6
CPR + current
98.1
99.8
100.0
100.0
Binary Alternation
Base
0.6
1.8
2.7
3.1
CPR + current
43.1
55.4
60.1
64.1
Binary Matrix
Base
17.9
27.7
31.7
34.4
CPR + current
51.0
56.6
58.9
60.9
Appendix
Table 19: CPR improves both pass@ 1 and higher- K success rates. CPR with current-task replay outperforms the base model throughout the sampling curve. Both policies use the same 64 held-out prompts per task and sample 16 responses at temperature 0.7 . The higher value within each task and K is bold.
Setting
Opus 4.8
GPT-5.5
Qwen-32B
Qwen-72B
Base
Seq. RLVR
CPR
MTRL
LLM algorithmic
53.7
77.7
29.3
–
17.5
49.9
63.3
65.3
LLM algebra
84.0
94.8
66.6
–
43.2
88.2
90.4
90.0
VLM quantitative
10.5
5.5
27.7
26.6
24.8
25.5
37.2
30.9
VLM spatial
19.9
0.8
27.7
28.9
24.1
26.3
32.5
31.6
VLM positional
19.1
8.6
27.7
25.8
28.7
34.7
33.2
35.3
Appendix
Table 20: CPR outperforms larger same-family zero-shot models across all five settings. Scores are mean verifier accuracy on CRG’s fixed evaluation sets with a 1024-token response budget. Claude Opus 4.8, GPT-5.5, and the larger Qwen backbones are evaluated zero-shot. Base is the untrained small model. The Seq. RLVR, CPR, and MTRL columns report trained policies. The best score in each row is bold.
Vision-Language Models in Continual Learning (VLM-CL) aim to continuously adapt to new multimodal tasks while retaining prior knowledge. The emerging paradigm that couples Multimodal Large Language Models (MLLMs) with Reinforcement Learning with Verifiable Rewards (RLVR) calls for a new pattern to guide continual adaptation. Advances in reasoning capability now make it feasible to impose constraints at the reasoning level. We formalize portability, a sample-level measure of how reusable the previous policy's behavior is on a new task, and empirically show that reasoning-level signals remain reliable on out-of-distribution samples while answer-level signals do not. We instantiate this as Reasoning Portability (RP) and propose Reasoning-based Dynamic Balance Continual Learning (RDB-CL), which modulates the per-sample Kullback-Leibler regularization in RLVR according to RP: a tight anchor preserves reusable reasoning on high-RP samples, while a relaxed anchor on low-RP samples permits exploration of new reasoning pathways. Experiments show that RDB-CL consistently outperforms baselines, improving Last accuracy by +12.0% over the vanilla RLVR baseline.
Qiuhe Hong, Yuyang Liu, Shuo Yang +3
1Shenzhen Graduate School of Peking University · Centre for Artificial Intelligence and Robotics, HKISI, CAS · 3Peng Cheng Laboratory
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLVR reliably represent how a model gets to its answer. In this paper, we develop two metrics for critically examining this assumption: Causal Importance of Reasoning (CIR), which measures the cumulative effect of reasoning tokens on the final answer, and Sufficiency of Reasoning (SR), which measures whether a verifier can arrive at an unambiguous answer based on the reasoning alone. Through experiments with the Qwen2.5 model series and ReasoningGym tasks, we find that: (1) while RLVR does improve task accuracy, it does not reliably improve CIR or SR, calling the role of reasoning in model performance into question; (2) a small amount of SFT before RLVR can be a remedy for low CIR and SR; and (3) CIR and SR can be improved even without SFT by applying auxiliary CIR/SR rewards on top of the outcome-based reward. This joint reward matches the accuracy of RLVR while also leading to causally important and sufficient reasoning. These results show that RLVR does not always lead models to rely on reasoning in the way that is commonly thought, but this issue can be remedied with simple modifications to the post-training procedure.
Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved problems quietly become unsolvable as training proceeds. We frame this phenomenon as \emph{correct-set turnover}, representing the coupled dynamics of solution acquisition and regression over the mastered set. Under this view, retention becomes an explicit optimization target alongside acquisition. We analytically and empirically establish the \emph{repair-window principle}: the cost of restoring a regressed prompt grows sharply with review delay, defining a low-cost window that standard RLVR pipelines fail to exploit. To address this, we propose \textbf{\method{}}, a retention-aware review mechanism that tracks mastered prompts and periodically reintroduces them to \textbf{remind} the model of previous solutions. By utilizing pre-rollout batch replacement, \method{} incurs zero additional rollout overhead. Evaluated across 20 benchmarks spanning image-text, video, and text-only tasks with Qwen3-VL and Qwen2.5-Math, \method{} consistently improves performance over GRPO, DAPO, and replay baselines, demonstrating robust generalizability across modalities and algorithms.
Chuanyu Qin, Chenxu Yang, Qingyi Si +3
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China