Federated continual learning must integrate new tasks over time without losing earlier-task knowledge. Most existing methods attach an anti-forgetting mechanism to the client-trained, server-aggregated loop of federated learning, which holds back new learning to preserve earlier knowledge and burdens resource-constrained clients. We propose ReSCENE, which structurally mitigates catastrophic forgetting by having each client upload a small condensed surrogate of its local data while the server keeps the surrogates of past tasks and trains the global model on them together with the current task surrogates. For efficient server memory, we introduce temporal herding, which selects the more recent surrogates from the pool accumulated over a task into a compressed buffer. Our study provides a theoretical analysis showing that this buffer can represent the original task data more closely than full accumulation of all surrogates. Across CIFAR-10, CIFAR-100, and TinyImageNet, ReSCENE achieves the strongest accuracy over seven baselines, by up to 31.1 points of average accuracy, while requiring as little as 0.11× of the client computation and up to 179× less upload than the model-update baselines. ReSCENE further demonstrates its effectiveness when scaled to larger client populations and larger models while remaining efficient, which makes it a practical method for federated continual learning.
Figures & tables
Figure 1: Illustration of ReSCENE. During task t , clients upload condensed surrogates every round, and the server trains the global model on the accumulated pool together with the buffers of earlier tasks. At the end of the task, temporal herding compresses the pool into a fixed-size buffer centered on the recent rounds, which joins the past buffer. Darker tiles denote more recent rounds.
Figure 2
Dataset
β
Metric
FedAvg
FedEWC
FedLwF
TARGET
GLFC
FedCBDR
Re-Fed
ReSCENE
CIFAR-10
0.1
AA
10.00 ± 0.00
10.00 ± 0.00
10.41 ± 0.61
9.95 ± 0.34
14.91 ± 5.36
25.94 ± 3.63
30.67 ± 3.72
61.78 ± 1.14
AIA
23.34 ± 0.88
23.37 ± 0.92
22.93 ± 0.09
22.80 ± 0.04
24.65 ± 2.41
37.23 ± 3.36
36.74 ± 5.41
68.35 ± 0.62
0.5
AA
15.84 ± 3.04
17.54 ± 0.94
16.94 ± 5.70
21.66 ± 7.12
35.38 ± 1.60
46.78 ± 0.93
50.73 ± 3.02
67.96 ± 1.33
AIA
32.90 ± 7.78
36.81 ± 5.03
33.57 ± 11.77
36.76 ± 13.40
48.90 ± 6.83
57.64 ± 4.59
62.69 ± 6.83
73.53 ± 0.73
1.0
AA
14.70 ± 4.25
15.12 ± 4.53
22.82 ± 7.36
18.86 ± 1.39
43.13 ± 0.96
50.06 ± 0.91
56.57 ± 1.26
69.64 ± 0.84
AIA
30.83 ± 1.89
31.78 ± 4.52
33.51 ± 5.96
32.37 ± 3.58
52.56 ± 5.15
57.07 ± 2.63
66.83 ± 2.33
74.57 ± 0.67
Table 1: AA and AIA (%) of the baselines and ReSCENE. Small gray values are standard deviations over three seeds.
Figure 3: Old-task versus recent-task accuracy at β=0.1 . Points high on both axes and near the diagonal forget less, performing well on both old and recent tasks.
Dataset
FedDM
FedAF
ReSCENE
CIFAR-10
AA
15.69
15.99
61.78
AIA
37.06
37.06
68.35
CIFAR-100
AA
6.27
6.37
44.89
AIA
16.31
16.35
49.67
TinyImageNet
AA
2.52
2.58
23.84
AIA
10.03
10.37
31.80
Table 2: AA and AIA (%) of the SA-FL baselines and ReSCENE at β=0.1 .
Client comp.
Uplink
Accuracy
Method
(PFLOP/run)
(MB/rd)
AA
AIA
FedAvg
46.33
44.73
14.58 ↓ 1.26
28.49 ↓ 4.41
FedEWC
48.65
44.73
15.97 ↓ 1.57
31.62 ↓ 5.19
FedLwF
58.69
44.73
14.38 ↓ 2.56
29.30 ↓ 4.27
TARGET
132.25
44.73
16.01 ↓ 5.65
31.29 ↓ 5.47
GLFC
108.45
44.73
34.52 ↓ 0.86
49.21 ↑ 0.31
Table 3: Accuracy and client cost with (a) more clients and (b) larger models. In (a), the arrows and the numbers beside them give the change from N=20 in Table 1 .
Accuracy
Method
AA
AIA
Earliest-1K
64.80
69.82
Latest-1K
66.56
72.67
Random-1K
65.28
70.81
Full-Pool Herd.
66.70
72.86
Temporal Herd.
67.96
73.53
Table 4: (a) Herding-buffer ablation with a fixed budget of B=1000 per class. (b) Comparison with full accumulation.
Accuracy
Esrv
Server comp.
AA
AIA
100
1.00 ×
61.78
68.35
50
0.50 ×
58.26
66.91
25
0.25 ×
53.91
64.81
10
0.10 ×
53.45
65.03
Re-Fed
30.67
36.74
Table 5: AA and AIA (%) in the server-epoch sweep.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Final accuracy of each task after the complete sequence for the methods of Table 1 , one column per benchmark and one row per β (values in Table 7 ).
Domain
Accuracy
Method
Amazon
DSLR
Webcam
AA
AIA
FedAvg
18.83
23.23
65.19
35.75
36.21
FedEWC
19.18
21.21
61.39
33.93
35.94
FedLwF
25.75
37.37
56.96
40.03
37.07
TARGET
29.48
20.20
39.87
29.85
34.37
GLFC
36.94
17.17
11.39
21.84
22.81
Appendix
Table 6: Domain-incremental results on Office-31 (%). Final accuracy on each domain, AA, and AIA.
β=0.1
β=0.5
β=1.0
Method
T0
T1
T2
T3
T4
T0
T1
T2
T3
T4
T0
T1
T2
T3
T4
FedAvg
0.0
0.0
0.0
0.0
50.0
0.0
0.0
0.0
0.0
79.2
0.0
0.0
0.0
0.0
73.5
FedEWC
0.0
0.0
0.0
0.0
50.0
0.0
0.0
0.0
0.0
87.6
0.0
0.0
0.0
0.0
75.6
FedLwF
31.7
0.2
0.0
0.4
19.6
25.1
11.3
30.0
2.8
15.5
32.6
18.4
13.3
45.5
4.3
TARGET
35.8
0.0
0.2
0.4
13.4
38.7
6.7
20.9
22.2
19.6
29.0
18.9
6.3
30.6
9.5
GLFC
0.4
0.2
3.3
20.5
50.1
4.9
4.3
16.8
58.2
92.7
13.9
7.0
28.9
74.3
91.6
Appendix
Table 7: Final accuracy (%) on each task after the complete sequence. T0 is the oldest task, and each row averages to the AA of the run.
Setting of Table 1
Matched setting (10,000 in total)
Method
Budget
Per client
Server +
AA
AIA
Budget
Server +
AA
AIA
clients
clients
TARGET
25,600
25,600
512,000
21.66
36.76
500
10,000
19.62
37.62
GLFC
20
200
4,000
35.38
48.90
50
10,000
40.67
51.69
FedCBDR
150
1,500
30,000
46.78
57.64
50
10,000
32.85
49.86
Re-Fed
1,000
≈ 513
10,253
50.73
62.69
1,000
10,253
50.73
62.69
Appendix
Table 8: Replay images and accuracy on CIFAR-10 ( β=0.5 ) at the budgets of Table 1 and after matching every baseline to ReSCENE’s 10,000 images. Budget denotes the memory parameter of each method, defined in the text.
(a) Target fraction
(b) Budget
λ
AA
AIA
B
AA
AIA
0.25
66.92
73.15
500
61.11
69.12
0.50
67.16
73.35
1000
67.96
73.53
0.75
67.96
73.53
1500
69.43
74.07
1.00
66.70
72.86
2000
68.36
73.86
Appendix
Table 9: Temporal herding sweeps on CIFAR-10 ( β=0.5 ). (a) varies λ at B=1000 , and (b) varies B at λ=0.75 . Bold marks the default.
Accuracy
Variant
α
τ
AA
AIA
Natural
1.0
0.0
67.26
72.55
Sampling only
0.5
0.0
67.64
72.65
Logit only
1.0
1.0
67.70
73.26
ReSCENE
0.5
1.0
67.96
73.53
Appendix
Table 10: Ablation of class-imbalance resolution on CIFAR-10 ( β=0.5 ).
Dataset
Method
Client comp.
State mem.
Uplink
Downlink
(PFLOP/run)
(MB)
(MB/rd)
(MB/rd)
CIFAR-10
FedAvg
32.68 ( × 1.0)
44.73
44.73
44.73
FedEWC
34.31 ( × 1.0)
134.13
44.73
44.73
FedLwF
41.32 ( × 1.3)
89.47
44.73
44.73
TARGET
72.20 ( × 2.2)
359.31
44.73
359.31
GLFC
48.43 ( × 1.5)
91.93
44.73
44.73
Appendix
Table 11: Client-side resources with ResNet-18. Multipliers are relative to FedAvg on the same dataset.
Asymptotic compute (per task)
CIFAR-10 (per task)
Method
All clients
Server
All clients
Server
Total
Clients
Server
(PFLOP)
(PFLOP)
(PFLOP)
share
share
FedAvg
O(RKEDlocF)
O(RKP)
6.54
≈ 0
6.54
100%
0%
FedEWC
O(RKEDlocF)
O(RKP)
6.86
≈ 0
6.86
100%
0%
FedLwF
O(RKEDlocF)
O(RKP)
8.26
≈ 0
8.26
100%
0%
TARGET
O(RKEDlocF)
O(RsynSkdBsynF)
14.44
38.2
52.64
27.4%
72.6%
Appendix
Table 12: Total computation across the server and all clients, asymptotically per task and on CIFAR-10 at task 2.
Method
T0
T1
T2
T3
T4
AA
FedAvg
0.00
0.00
0.00
0.00
50.00
10.00
FedCBDR
0.00
0.00
0.00
0.00
49.15
9.83
Re-Fed
8.10
0.00
0.00
0.00
46.35
10.89
ReSCENE
64.15
36.45
43.75
62.15
69.15
55.13
Appendix
Table 13: Privacy robustness on CIFAR-10 ( β=0.5 ). Top, AA and AIA (%) at each noise multiplier σ . Bottom, final per-task accuracy (%) at σ=5 .