Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.
Figures & tables
Figure 1: Isolated Spike and Coherent Forgetting.
Figure 2: Overview of SPHERE.
Buffer size 300
Buffer size 3000
Method
FP ↑
Worst Forget ↓
Tail Forget ↓
Min. ↑
FP ↑
Worst Forget ↓
Tail Forget ↓
Min. ↑
No replay
.415 ± .146
.811 ± .156
.693 ± .173
.130
.415 ± .146
.811 ± .156
.693 ± .173
.130
ER
.661 ± .075
.245 ± .149
.200 ± .137
.423
.641 ± .071
.329 ± .143
.269 ± .112
.479
Loss-prop
.689 ± .045
.188 ± .061
.142 ± .043
.538
.618 ± .050
.325 ± .085
.269 ± .076
.502
MIR-prop
.675 ± .045
.188 ± .083
.152 ± .072
.570
.635 ± .042
.296 ± .111
.237 ± .079
.561
MIR-top
.621 ± .023
.244 ± .055
.211 ± .060
.588
.633 ± .066
.375 ± .205
.285 ± .155
.500
Table 1: Replay performance with limited memory and replay. Values are mean ± SD except Min.
Metric
S=I
SPHERE
Gain
Final acc. ↑
64.90
69.32
+4.42
Incremental acc. ↑
66.08
68.54
+2.47
New task acc. ↑
70.34
72.43
+2.08
Mean forgetting ↓
8.88
6.18
+2.70
Worst forgetting ↓
24.75
16.85
+7.90
Tail forgetting ↓
20.59
13.54
+7.05
Table 2: Matched neighborhood-block ablation (%).
Variant
FP ↑
Worst ↓
Tail ↓
Min. ↑
Full SPHERE
.690 ± .021
.180 ± .053
.143 ± .035
.661
Consensus only
.680 ± .021
.189 ± .054
.154 ± .030
.647
No transport
.640 ± .063
.275 ± .129
.227 ± .110
.478
Permuted geometry
.666 ± .056
.213 ± .082
.188 ± .078
.497
Matched MIR
.682 ± .048
.192 ± .113
.161 ± .087
.517
No uniform floor
.673 ± .034
.230 ± .082
.183 ± .063
.608
Table 3: Component controls with scarce memory and replay. Values are mean ± SD except Min.
CIFAR-100 ( K=2,000 )
TinyImageNet ( K=4,000 )
Method
Clean
20%
40%
60%
Clean
20%
40%
60%
ER
19.6 ± 3.4
14.3 ± 1.1
10.2 ± 1.1
5.1 ± 0.8
16.3 ± 1.2
12.0 ± 0.6
7.4 ± 0.6
3.3 ± 0.4
ER + SPHERE
20.9 ± 1.8
15.7 ± 1.2
12.2 ± 0.8
8.0 ± 0.7
15.3 ± 1.4
13.5 ± 0.9
10.1 ± 0.6
6.1 ± 0.7
MIR-prop
19.8 ± 1.4
15.0 ± 2.3
10.2 ± 1.6
6.2 ± 0.8
17.1 ± 0.5
11.6 ± 1.1
8.0 ± 0.8
4.7 ± 0.2
DER++
15.3 ± 4.1
10.9 ± 4.1
9.0 ± 2.0
5.6 ± 1.0
9.8 ± 1.0
7.4 ± 0.9
4.8 ± 0.7
3.4 ± 0.3
ER-ACE
20.4 ± 3.4
16.8 ± 2.1
12.9 ± 1.5
8.1 ± 1.1
17.9 ± 1.7
14.4 ± 0.8
9.6 ± 0.8
5.5 ± 0.4
Table 4: Replay across hosts under symmetric label noise.
Method
Class-IL ↑
Task-IL ↑
AAA ↑
avgF ↓
worst-class ↑
CVaR 0.2↓
ER
33.4 ± 4.5
76.1 ± 5.1
51.4 ± 4.8
42.9 ± 5.6
5.6 ± 4.5
72.5 ± 5.7
ER-ACE
36.1 ± 4.5
77.9 ± 6.5
53.2 ± 2.1
25.3 ± 5.5
11.1 ± 7.2
54.8 ± 8.9
GDumb
36.8 ± 0.9
77.0 ± 2.8
28.1 ± 0.7
20.2 ± 7.4
8.4 ± 5.4
51.2 ± 19.0
DER++
28.2 ± 4.3
80.7 ± 4.8
48.4 ± 4.7
57.5 ± 9.3
1.3 ± 1.7
81.3 ± 7.0
MIR
32.1 ± 4.4
75.6 ± 4.1
48.5 ± 3.5
36.9 ± 7.0
5.5 ± 3.6
71.1 ± 6.9
X-DER
28.1 ± 8.2
75.1 ± 8.0
44.9 ± 5.3
25.0 ± 9.9
0.0 ± 0.0
60.0 ± 16.1
Table 5: CIFAR-10 at 20% noise, K=500 (mean ± SD).
Metric
SPHERE
S=I
Final acc ↑
46.61 ± 2.19
45.42 ± 2.99
Stage-average acc ↑
59.01 ± 3.00
58.93 ± 3.30
Mean forgetting ↓
35.10 ± 5.90
36.55 ± 7.71
Worst forgetting ↓
62.66 ± 9.63
69.85 ± 12.42
Top-2 forgetting ↓
59.53 ± 8.81
61.84 ± 10.69
Table 6: Neighborhood-block ablation on clean Split CIFAR-10.
Method
Final ↑
Old ↑
Worst ↑
F↓
BWT ↑
Uniform replay
42.75 ± 2.92
43.56
40.00
4.00
-2.78
MIR
44.50 ± 1.63
45.78
38.67
4.56
-3.89
No Transport
42.67 ± 1.70
41.78
39.67
6.56
-5.56
Permuted geometry
48.58 ± 1.88
45.78
41.67
3.56
-2.89
SPHERE
49.08 ± 0.80
46.89
43.67
3.33
-1.56
Table 7: Code-generation results.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Setting
Representation
Current T5-Large with its LoRA adapters; last decoder layer, mean-pooled over target-label tokens (1024 dimensions), then ℓ2 -normalized. Refreshed with the candidate pool.
Distance
Cij=∥hi−hj∥2/ν , where ν=mediank<l∥hk−hl∥2 within the pool.
Allocation
ϵs=0.2 , ϵT=0.5 ; τ=cβstd([Sδ]+)/ϵT with cβ=1 ; α=0.5 .
Optimization
AdamW learning rate 10−3 for training; one virtual SGD step on LoRA parameters with ηv=10−2 .
Memory and scoring
K∈{300,3000} ; up to 512 distinct candidates, refreshed every four replay events.
Replay
Four examples per update, without replacement; unit-mass mean replay loss, λ=1 .
Appendix
Table 8: Instruction-tuning configuration. Distances and allocation scales are computed on the scored candidate pool.
Table 9: Task orders for instruction-tuning evaluation.
Dataset
Host
ks
kT
Tfrac
Interval
η
CIFAR-10, MNIST
ER
5
25
0.5
4
0.1
CIFAR-10
ER-ACE
10
50
1.0
4
0.1
CIFAR-10
X-DER
5
50
0.5
8
0.03
CIFAR-100
ER, ER-ACE
5
25
0.5
4
0.1
CIFAR-100
X-DER
5
25
0.5
4
0.03
TinyImageNet
ER
5
25
0.5
4
0.1
Appendix
Table 10: Vision allocation settings. The last two columns give the rescoring interval in stream steps and the virtual-step size. α=0.5 ; ER and ER-ACE use λ=1 .
Table 12: Final Class-IL accuracy on Split CIFAR-10 (%, memory 500). Values are mean ± SD. Best means are bold within each protocol.
Method
Class-IL ↑
Task-IL ↑
AAA ↑
avgF ↓
Worst-class ↑
CVaR 0.2↓
40% symmetric noise
ER
25.1 ± 2.8
71.6 ± 3.4
45.3 ± 3.6
49.7 ± 5.1
2.6 ± 1.5
81.1 ± 3.5
ER-ACE
27.9 ± 3.4
72.7 ± 4.1
47.0 ± 3.3
28.4 ± 4.5
8.3 ± 6.0
59.4 ± 8.6
ER-ACE + SPHERE
31.3 ± 2.4
74.3 ± 5.4
50.2 ± 3.0
18.4 ± 6.4
16.4 ± 3.6
52.2 ± 8.0
DER++
23.3 ± 4.8
76.8 ± 4.7
39.7 ± 4.9
60.6 ± 9.5
2.0 ± 4.6
86.9 ± 7.5
MIR-prop
24.7 ± 2.1
69.3 ± 3.8
38.7 ± 2.4
37.9 ± 5.2
2.5 ± 2.5
72.8 ± 5.3
Appendix
Table 13: CIFAR-10 retention (%, K=500 ; mean ± SD). Bold compares ER-ACE with its SPHERE integration.
Control
Operation
S=I
Use the identity for aggregation and omit its pullback; retain clipping, transport with C , and the uniform floor.
Consensus only
Replace q∗(v) by softmaxT(v) , with T chosen to match the ESS of the full method’s weights. Set wSC=S⊤(softmaxT(v)⊙χ) , and retain the uniform mixture and normalization in Eq. ( 83 ).
No transport
Normalize the supported field directly, N(v) , without transport or the S⊤ pullback.
Permuted geometry
At every pool refresh, randomly permute feature rows before constructing C and S , breaking their alignment with the memories while retaining the full allocation procedure.
Matched MIR
Use softmaxT([δ]+) with T chosen to match the ESS of the full method’s weights.
No uniform floor
Set α=1 ; remove only the uniform mixture from the full allocation rule.
Appendix
Table 14: Instruction-tuning controls in Tables 5.1 and 5.1 . The two transport-free controls use different allocation rules.
Tieliang Gong, Zhongbo Zhang and Wen Wen are with the School of Computer Science and Technology, Xi’an Jiaotong University, China · Yong-Jin Liu is with Department of Computer Science and Technology, Tsinghua University, China