Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.
Figures & tables
Figure 1: Isolated Spike and Coherent Forgetting.
Figure 2: Overview of SPHERE.
Buffer size 300
Buffer size 3000
Method
FP ↑
Worst Forget ↓
Tail Forget ↓
Min. ↑
FP ↑
Worst Forget ↓
Tail Forget ↓
Min. ↑
No replay
.415 ± .146
.811 ± .156
.693 ± .173
.130
.415 ± .146
.811 ± .156
.693 ± .173
.130
ER
.661 ± .075
.245 ± .149
.200 ± .137
.423
.641 ± .071
.329 ± .143
.269 ± .112
.479
Loss-prop
.689 ± .045
.188 ± .061
.142 ± .043
.538
.618 ± .050
.325 ± .085
.269 ± .076
.502
MIR-prop
.675 ± .045
.188 ± .083
.152 ± .072
.570
.635 ± .042
.296 ± .111
.237 ± .079
.561
MIR-top
.621 ± .023
.244 ± .055
.211 ± .060
.588
.633 ± .066
.375 ± .205
.285 ± .155
.500
Table 1: Replay performance with limited memory and replay. Values are mean ± SD except Min.
Metric
S=I
SPHERE
Gain
Final acc. ↑
64.90
69.32
+4.42
Incremental acc. ↑
66.08
68.54
+2.47
New task acc. ↑
70.34
72.43
+2.08
Mean forgetting ↓
8.88
6.18
+2.70
Worst forgetting ↓
24.75
16.85
+7.90
Tail forgetting ↓
20.59
13.54
+7.05
Table 2: Matched neighborhood-block ablation (%).
Variant
FP ↑
Worst ↓
Tail ↓
Min. ↑
Full SPHERE
.690 ± .021
.180 ± .053
.143 ± .035
.661
Consensus only
.680 ± .021
.189 ± .054
.154 ± .030
.647
No transport
.640 ± .063
.275 ± .129
.227 ± .110
.478
Permuted geometry
.666 ± .056
.213 ± .082
.188 ± .078
.497
Matched MIR
.682 ± .048
.192 ± .113
.161 ± .087
.517
No uniform floor
.673 ± .034
.230 ± .082
.183 ± .063
.608
Table 3: Component controls with scarce memory and replay. Values are mean ± SD except Min.
CIFAR-100 ( K=2,000 )
TinyImageNet ( K=4,000 )
Method
Clean
20%
40%
60%
Clean
20%
40%
60%
ER
19.6 ± 3.4
14.3 ± 1.1
10.2 ± 1.1
5.1 ± 0.8
16.3 ± 1.2
12.0 ± 0.6
7.4 ± 0.6
3.3 ± 0.4
ER + SPHERE
20.9 ± 1.8
15.7 ± 1.2
12.2 ± 0.8
8.0 ± 0.7
15.3 ± 1.4
13.5 ± 0.9
10.1 ± 0.6
6.1 ± 0.7
MIR-prop
19.8 ± 1.4
15.0 ± 2.3
10.2 ± 1.6
6.2 ± 0.8
17.1 ± 0.5
11.6 ± 1.1
8.0 ± 0.8
4.7 ± 0.2
DER++
15.3 ± 4.1
10.9 ± 4.1
9.0 ± 2.0
5.6 ± 1.0
9.8 ± 1.0
7.4 ± 0.9
4.8 ± 0.7
3.4 ± 0.3
ER-ACE
20.4 ± 3.4
16.8 ± 2.1
12.9 ± 1.5
8.1 ± 1.1
17.9 ± 1.7
14.4 ± 0.8
9.6 ± 0.8
5.5 ± 0.4
Table 4: Replay across hosts under symmetric label noise.
Method
Class-IL ↑
Task-IL ↑
AAA ↑
avgF ↓
worst-class ↑
CVaR 0.2↓
ER
33.4 ± 4.5
76.1 ± 5.1
51.4 ± 4.8
42.9 ± 5.6
5.6 ± 4.5
72.5 ± 5.7
ER-ACE
36.1 ± 4.5
77.9 ± 6.5
53.2 ± 2.1
25.3 ± 5.5
11.1 ± 7.2
54.8 ± 8.9
GDumb
36.8 ± 0.9
77.0 ± 2.8
28.1 ± 0.7
20.2 ± 7.4
8.4 ± 5.4
51.2 ± 19.0
DER++
28.2 ± 4.3
80.7 ± 4.8
48.4 ± 4.7
57.5 ± 9.3
1.3 ± 1.7
81.3 ± 7.0
MIR
32.1 ± 4.4
75.6 ± 4.1
48.5 ± 3.5
36.9 ± 7.0
5.5 ± 3.6
71.1 ± 6.9
X-DER
28.1 ± 8.2
75.1 ± 8.0
44.9 ± 5.3
25.0 ± 9.9
0.0 ± 0.0
60.0 ± 16.1
Table 5: CIFAR-10 at 20% noise, K=500 (mean ± SD).
Metric
SPHERE
S=I
Final acc ↑
46.61 ± 2.19
45.42 ± 2.99
Stage-average acc ↑
59.01 ± 3.00
58.93 ± 3.30
Mean forgetting ↓
35.10 ± 5.90
36.55 ± 7.71
Worst forgetting ↓
62.66 ± 9.63
69.85 ± 12.42
Top-2 forgetting ↓
59.53 ± 8.81
61.84 ± 10.69
Table 6: Neighborhood-block ablation on clean Split CIFAR-10.
Method
Final ↑
Old ↑
Worst ↑
F↓
BWT ↑
Uniform replay
42.75 ± 2.92
43.56
40.00
4.00
-2.78
MIR
44.50 ± 1.63
45.78
38.67
4.56
-3.89
No Transport
42.67 ± 1.70
41.78
39.67
6.56
-5.56
Permuted geometry
48.58 ± 1.88
45.78
41.67
3.56
-2.89
SPHERE
49.08 ± 0.80
46.89
43.67
3.33
-1.56
Table 7: Code-generation results.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Setting
Representation
Current T5-Large with its LoRA adapters; last decoder layer, mean-pooled over target-label tokens (1024 dimensions), then ℓ2 -normalized. Refreshed with the candidate pool.
Distance
Cij=∥hi−hj∥2/ν , where ν=mediank<l∥hk−hl∥2 within the pool.
Allocation
ϵs=0.2 , ϵT=0.5 ; τ=cβstd([Sδ]+)/ϵT with cβ=1 ; α=0.5 .
Optimization
AdamW learning rate 10−3 for training; one virtual SGD step on LoRA parameters with ηv=10−2 .
Memory and scoring
K∈{300,3000} ; up to 512 distinct candidates, refreshed every four replay events.
Replay
Four examples per update, without replacement; unit-mass mean replay loss, λ=1 .
Appendix
Table 8: Instruction-tuning configuration. Distances and allocation scales are computed on the scored candidate pool.
Table 9: Task orders for instruction-tuning evaluation.
Dataset
Host
ks
kT
Tfrac
Interval
η
CIFAR-10, MNIST
ER
5
25
0.5
4
0.1
CIFAR-10
ER-ACE
10
50
1.0
4
0.1
CIFAR-10
X-DER
5
50
0.5
8
0.03
CIFAR-100
ER, ER-ACE
5
25
0.5
4
0.1
CIFAR-100
X-DER
5
25
0.5
4
0.03
TinyImageNet
ER
5
25
0.5
4
0.1
Appendix
Table 10: Vision allocation settings. The last two columns give the rescoring interval in stream steps and the virtual-step size. α=0.5 ; ER and ER-ACE use λ=1 .
Table 12: Final Class-IL accuracy on Split CIFAR-10 (%, memory 500). Values are mean ± SD. Best means are bold within each protocol.
Method
Class-IL ↑
Task-IL ↑
AAA ↑
avgF ↓
Worst-class ↑
CVaR 0.2↓
40% symmetric noise
ER
25.1 ± 2.8
71.6 ± 3.4
45.3 ± 3.6
49.7 ± 5.1
2.6 ± 1.5
81.1 ± 3.5
ER-ACE
27.9 ± 3.4
72.7 ± 4.1
47.0 ± 3.3
28.4 ± 4.5
8.3 ± 6.0
59.4 ± 8.6
ER-ACE + SPHERE
31.3 ± 2.4
74.3 ± 5.4
50.2 ± 3.0
18.4 ± 6.4
16.4 ± 3.6
52.2 ± 8.0
DER++
23.3 ± 4.8
76.8 ± 4.7
39.7 ± 4.9
60.6 ± 9.5
2.0 ± 4.6
86.9 ± 7.5
MIR-prop
24.7 ± 2.1
69.3 ± 3.8
38.7 ± 2.4
37.9 ± 5.2
2.5 ± 2.5
72.8 ± 5.3
Appendix
Table 13: CIFAR-10 retention (%, K=500 ; mean ± SD). Bold compares ER-ACE with its SPHERE integration.
Control
Operation
S=I
Use the identity for aggregation and omit its pullback; retain clipping, transport with C , and the uniform floor.
Consensus only
Replace q∗(v) by softmaxT(v) , with T chosen to match the ESS of the full method’s weights. Set wSC=S⊤(softmaxT(v)⊙χ) , and retain the uniform mixture and normalization in Eq. ( 83 ).
No transport
Normalize the supported field directly, N(v) , without transport or the S⊤ pullback.
Permuted geometry
At every pool refresh, randomly permute feature rows before constructing C and S , breaking their alignment with the memories while retaining the full allocation procedure.
Matched MIR
Use softmaxT([δ]+) with T chosen to match the ESS of the full method’s weights.
No uniform floor
Set α=1 ; remove only the uniform mixture from the full allocation rule.
Appendix
Table 14: Instruction-tuning controls in Tables 5.1 and 5.1 . The two transport-free controls use different allocation rules.
Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation samples by their remaining learning potential and replay samples by their observed forgetting. Their competition for a shared training budget yields an online curriculum that determines what to train on at each step. Across models, scales, and adaptation domains, RoD reaches or improves upon the adaptation-forgetting frontier of tuned fixed-replay baselines and model merging without prescribing a replay allocation in advance. Replay concentrates on sources that are more vulnerable to forgetting and dynamically increases and redistributes as forgetting emerges during training. Together, our results show that replay can be allocated online from the model's evolving state, targeting what is needed, when it is needed.
Lukas Thede, Shengzhuang Chen, Stefan Winzeck +3
University of Tübingen, Tübingen AI Center · Helmholtz Munich · Munich Center for Machine Learning (MCML) +3
Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (replaying stored past data) or rehearsal-free (regularising or isolating parameters), overwhelmingly target one objective: preventing catastrophic forgetting. Forward transfer, the past helping the future, has meanwhile been pursued almost exclusively through parameter reuse, with no explicit account of when transfer should be expected at all. We begin one step earlier: before designing a transfer mechanism, we ask when transfer should exist at all. We answer with a framework of three measurable conditions: the target task must leave room for improvement beyond its own limited supervision, transferable information must survive continued optimisation, and replay must come from compatible previous tasks. We instantiate this view as Transfer-Selective Replay (TSR), which selects replay data predicted to benefit the incoming task rather than replaying past examples indiscriminately. Selection is guided by a zero-training task signature, while distillation preserves stability on previous tasks. Under the standard continual learning protocol in the low-budget regime, TSR consistently improves forward transfer while maintaining stability, outperforming existing replay baselines across heterogeneous and homogeneous task streams. More broadly, the results argue for treating transfer as a first-class objective of continual learning, to be understood before it is engineered.
Yang Meng, Zhenya Liu, Zhuokai Zhao +1
Department of Computer Science, University of Chicago
Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: finite memory replaces each past distribution with an empirical proxy, and repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory. We develop a layer-wise information-theoretic framework that separates these effects at every depth. Our main result decomposes the expected generalization gap into a replay-induced representation drift and an optimization-dependence term, the latter further resolved into stability, plasticity, interaction, and residual-coupling components. Two refinements make the framework operational. A Wasserstein relaxation of the drift term, valid under support mismatch, yields a depth-dependent drift--sensitivity trade-off whose minimizer identifies which interior layer to stabilize. An SGLD instantiation of the optimization term reduces it to a trajectory-level log-determinant budget, exposing a curvature-aware gradient-alignment statistic that serves as an online diagnostic of task-wise forgetting. Controlled and benchmark experiments confirm the predicted memory scaling, the interior funnel, and the alignment signal's link to forgetting.
Tieliang Gong, Zhongbo Zhang, Wen Wen +1
Tieliang Gong, Zhongbo Zhang and Wen Wen are with the School of Computer Science and Technology, Xi’an Jiaotong University, China · Yong-Jin Liu is with Department of Computer Science and Technology, Tsinghua University, China