Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
Figures & tables
Figure 1: Measured return landscapes of neuroevolution and reinforcement learning at a task switch. Each panel shows the return landscape around a method’s solution at a switch between two alternating tasks. The circle marks the solution of the previous task and the diamond that of the new task. Blue and orange are the neighborhoods of the previous and the new task, the weights around the solution whose policies still solve it, and dark grey is their overlap. The arrow shows the move from one solution to the other. ES has the widest neighborhoods and overlap, and solves both tasks at the switch. The GA’s overlap is smaller and at the switch the solution escapes it. PPO has a very small neighborhood and has lost plasticity: it does not move at the switch. Continual RL variants reach the neighborhood of the new task but leave that of the previous one. The panels show empirical landscapes measured in our study (see Section 4 for a definition of a neighborhood, Section 5 for a deeper empirical discussion and Appendix H.6 for more examples of such landscapes).
Figure 2: Stability-plasticity trade-off. Learning accuracy (LA, higher is better) against forgetting (F, lower is better) a) with a sequence of twenty tasks and b) with two alternating tasks. Grey lines are equal LA − F. A ring marks the best method on the stability-plasticity trade-off (LA-F). GA is the most plastic method and ES the one with the best trade-off.
Figure 3: Return and zero-shot transfer. Cumulative return of the centroid (Cum.) and its zero-shot transfer to the next task (ZT), with a sequence of tasks (top) and with two alternating tasks. Mean and 95% bootstrap interval over trials.
Figure 4: Neighborhood width and the shared neighborhood. Left: neighborhood width relative to PPO, one point per method and classic control or MiniGrid setting; bars are medians and stars a Wilcoxon test over settings. Right: shared neighborhood (random moves of radius 0.1 ) against LA − F, with the Spearman correlation over all points and within each family. ES exhibits the largest neighborhood and neighborhood overlap correlates with the stability-plasticity trade-off.
Figure 5: Effect of novelty search. Return of the best individual (elite) and of the centroid, behavioural diversity and genomic diversity; dashed lines mark task switches.
Figure 6: Dormant units and weight drift. Fraction of dormant hidden units of the centroid (top) and root mean square of its weights (bottom). Dashed lines mark task switches. The dormant fraction of NE depends on the task and stays constant; that of RL rises.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Suite
Environment
Obser-
Actions
Policy
Episode
Steps /
vations
params
length
task
gymnax
CartPole-v1
R4
D(2)
386
500
1.54×108
gymnax
Acrobot-v1
R6
D(3)
435
500
1.54×108
gymnax
MountainCar-v0
R2
D(3)
371
500
1.54×108
gymnax
DeepSea ( 12×12 )
{0,1}144
D(2)
2,626
12
3.69×106
MiniGrid
Empty rooms
{0,1}7×7×25
D(6)
7,758
1024
3.15×108
Appendix
Table 1: The environments. Observations are real-valued ( Rn ) or binary ( {0,1}n ): DeepSea is a one-hot position on the grid, MiniGrid a one-hot encoding of the tile and colour in each cell of the egocentric view, and Kinetix an RGB image. D(k) : k discrete actions. Policy params count the actor only. Steps / task is the environment-step budget of one task.
Policy network (all methods)
Value network (RL only)
Environment
hidden
act.
output
params
hidden
params
CartPole-v1
16,16
ReLU
linear
386
128,128,128
33,793
Acrobot-v1
16,16
ReLU
linear
435
128,128,128
34,049
MountainCar-v0
16,16
ReLU
linear
371
128,128,128
33,537
MiniGrid
conv 4 ; 64
ReLU
linear
7,758
128,128,128
190,081
HalfCheetah
128,128
tanh
tanh
19,590
128,128
18,945
Appendix
Table 2: Policy and value network architectures. Hidden widths are listed in order and the activation column is the hidden activation; every layer carries a bias. The policy is shared by all methods; the value network exists only for the gradient-based methods. On HalfCheetah the PPO actor is the evolved policy plus a single state-independent logσ vector (initialised at −0.5 ), adding 6 parameters; on the discrete-action environments the PPO actor is the evolved policy, with a categorical head.
shared
ES
GA
Environment
P
R
G
shaping
optimiser
σ
α
σ
ρ
CartPole-v1
512
3
200
z-score
SGD
0.1
0.05
0.5
0.5
Acrobot-v1
512
3
200
z-score
SGD
0.1
0.05
0.5
0.5
MountainCar-v0
512
3
200
z-score
SGD
0.1
0.05
0.5
0.5
DeepSea
512
3
200
z-score
SGD
0.1
0.05
0.5
0.5
MiniGrid
512
3
200
z-score
SGD
0.1
0.05
0.01
0.1
Appendix
Table 3: Hyperparameters for NE. P : candidates per generation (the re-scored archive included). R : episodes per candidate. G : generations per task. ES: fitness shaping (z-score or centred ranks), optimiser, width σ and learning rate α . GA: mutation width σ and archive fraction ρ=μ/P . GA + Novelty uses the GA’s σ , ρ=0.5 and k=3 , and is run only on the classic-control environments and DeepSea. On MiniGrid the GA uses P=256 and R=6 , the same budget. On MountainCar and Kinetix the GA is the variant described above, and σ is its initial width.
Environment
N
L
K
M
γ
GAE λ
α
clip ε
U
CartPole-v1
2048
50
10
32
0.95
0.95
3×10−4
0.2
1500
Acrobot-v1
2048
50
10
32
0.99
0.95
1×10−4
0.2
1500
MountainCar-v0
2048
50
10
32
0.99
0.95
3×10−4
0.2
1500
DeepSea
2048
50
10
32
0.99
0.95
3×10−4
0.2
36
MiniGrid
2048
50
4
16
0.99
0.95
5×10−4
0.2
3072
HalfCheetah
512
20
10
32
0.97
0.95
3×10−4
0.3
2400
Appendix
Table 4: Hyperparameters for RL, shared by PPO and its four continual variants. N : parallel environments. L : rollout length. K : epochs per update. M : minibatches per epoch. U : updates per task. Value coefficient 0.5 and entropy coefficient 0.01 in every environment. The settings specific to each variant are in Table 5 .
Variant
Hyperparameter
Value
Environments
TRAC-PPO
tuner discounts
0.9,…,0.999999 (six)
all
initial scale S0
≈0
all
ReDo-PPO
dormancy threshold τ
0.025
all
interval Tredo (updates)
50
classic control
1000
other environments
networks reset
policy and value
all
Appendix
Table 5: Hyperparameters specific to the continual variants of PPO, on top of Table 4 . Classic control: CartPole, Acrobot and MountainCar. C-CHAIN follows the reference implementation of Tang et al. (2025) : its categorical settings on the discrete-action environments and its continuous-control settings on HalfCheetah. No variant is told when the task changes: the C-CHAIN coefficient, the ReDo schedule and the PBT-PPO exploit schedule all run unchanged across task switches. PBT-PPO’s N is the one with the higher cumulative return in each figure.
Figure 7: Stationary tasks. Top: return of the centroid without task switches (MiniGrid: first 250 generations; bottom: the twenty Kinetix levels). Table: area under the centroid curve (reward × generations / 1000) and, in grey, the median generation at which the smoothed curve covers 95% of its rise to its final value; Kinetix is averaged over levels. † PBT-PPO with N=8 members.
Figure 8: Return of the centroid over training. The settings of Figure 2 , in its layout: (a) a sequence of tasks (ten observation offsets, each visited twice; twenty Kinetix levels) and (b) two alternating tasks. Dashed lines are task switches.
Figure 9: Symptoms of plasticity loss in the continual settings. Figure 6 in ten continual settings (top block: observation noise and MiniGrid; bottom block: action reversal and Kinetix). HalfCheetah and Kinetix have almost no dormant units, so their dormancy lines sit at zero. All rows as defined in Appendix A.3 .
Setting
Untrained
GA
ES
PPO, task 1
PPO, task 20
CartPole, noise
0.05 ± 0.04
0.07
0.09
0.06
0.24
CartPole, action reversal
0.09 ± 0.05
0.24
0.24
0.14
0.50
Acrobot, noise
0.03 ± 0.03
0.05
0.06
0.02
0.08
Acrobot, action reversal
0.08 ± 0.04
0.10
0.10
0.06
0.15
MountainCar, noise
0.42 ± 0.08
0.46
0.36
0.23
0.31
MountainCar, action reversal
0.49 ± 0.08
0.46
0.42
0.22
0.38
Appendix
Table 6: Dormant units of an untrained network against the trained methods, classic control. Fraction of hidden units below ReDo’s threshold on the pooled probe states of Figure 9 : fifty freshly initialised 16×16 ReLU policies (mean ± s.d.), the GA and ES averaged over the twenty tasks, and PPO at the end of the first and of the last task.
Figure 10: Effect of novelty search in all seven settings. Figure 5 for CartPole, Acrobot and MountainCar under observation noise (top; MountainCar at σ=0.5 ) and under action reversal, and DeepSea (bottom). Rows as in Figure 5 , defined in Appendix A.3 .
Figure 11: Learning accuracy against forgetting as the minibatch size changes. One panel per setting. The four values of M , the number of minibatches per epoch, are joined in order of learning accuracy and labelled. A larger M means a smaller minibatch and more gradient steps per task. The reported value is outlined. Both axes are in the environment’s own return units. Lower is more stable and further right is more plastic, so a path along the diagonal trades one for the other.
Figure 12: Learning accuracy and forgetting of the GA and ES against the search width. One panel pair per setting: learning accuracy above forgetting, against the multiple of the reported width. Both axes are rescaled so that 0 is the untrained network and 1 the best reported method in that setting.
Figure 13: Stability-plasticity trade-off against the length of a task, classic control under observation noise. One row per length, one column per environment; the middle row is the reported setting, as in Figure 2 . Learning accuracy against forgetting, the region, lines and ring as in Figure 2 , except that both axes are rescaled per environment on one scale shared by the three lengths, so the rows compare directly.
Figure 14: Population diversity in the continual settings. One row per setting, one column per property: standard deviation of fitness, behavioural diversity (disagreement of the members’ greedy actions on probe states; mean action distance on HalfCheetah) and genomic diversity per weight. HalfCheetah, MiniGrid and Kinetix are left out: GA + Novelty has no diversity record there.
Figure 15: Two alternating tasks in toy landscapes. Top: smooth landscape ( h=0.1 ). Bottom: rugged landscape ( a=1.6 ), with local optima as white dots; at this depth the score between neighbouring dots is zero. Left: task A in blue and task B in red; the shared region is dashed. On the rugged landscape, the centroid path of one seed per method. On the smooth landscape (zoomed), the centroid at the end of each of the last four tasks, marked by the task’s letter: the GA switches sides, while ES stays in the shared region. Middle: share of seeds whose centroid ends in the shared region, at every level (grey: the drawn level), with both methods at σ=0.05 (smooth) and σ=0.2 (rugged). Right: the same share against the search width σ , at the drawn level (grey: the width used in the other panels).
intrinsic dimensionality k
Method
a
2
4
8
16
32
64
128
ES
0.4
24
24
24
24
24
24
24
0.8
24
24
24
24
24
24
18
1.6
24
24
24
24
17
0
0
GA
0.4
24
24
24
24
0
0
0
0.8
24
24
23
1
0
0
0
Appendix
Table 7: Intrinsic dimensionality in the rugged landscape ( 128 parameters). Seeds out of 24 whose centroid ends in the shared region, for intrinsic dimensionality k (the number of rippled coordinates) and ripple depth a . The other 128−k coordinates do not affect the score. Both methods use σ=0.4 , their best width among 0.1 , 0.2 and 0.4 .
Curve
Level
GA
ES
TRAC-PPO
ReDo-PPO
C-CHAIN
PBT-PPO
ρ
Share lost, <0.5
0.25
2.3
4.6 ∗∗∗
1.6 ∗∗∗
1.1
1.3 ∗∗∗
1.8 ∗∗∗
0.99
0.5 †
2.3
3.9 ∗∗∗
1.5 ∗∗∗
1.2
1.4 ∗∗∗
2.0 ∗∗∗
–
0.75
2.2
3.6 ∗∗∗
1.7 ∗∗
1.2
1.2 ∗∗∗
2.0 ∗∗∗
0.99
0.9
1.6
2.2 ∗∗∗
1.7 ∗∗
1.0
1.2 ∗
1.9 ∗∗
0.92
Share lost, <0.8
0.5
2.5
4.3 ∗∗∗
1.6 ∗∗
1.1
1.3 ∗∗∗
1.9 ∗∗∗
0.99
Median return
0.9
2.8
5.7 ∗∗∗
1.6 ∗∗∗
1.2
1.4 ∗∗∗
2.0 ∗∗∗
0.98
Appendix
Table 8: Neighborhood width under other levels and references. Width relative to PPO, median over the thirteen classic-control and MiniGrid panels of Figure 4 , when the width is read at another point of the same curves: the radius at which the given share of the 64 perturbed copies has lost the task ( share lost : rescaled return below 0.5 , or below 0.8 ), or the radius at which the median return of the copies relative to the unperturbed policy’s own falls below the given level ( median return ), the robustness score of Lehman et al. (2018) . Stars: Wilcoxon signed-rank test over panels ( ∗ p<0.05 , ∗∗ p<0.01 , ∗∗∗ p<0.001 ). ρ : Spearman correlation of the 78 per-panel ratios with those of the reported definition ( † ).
Figure 16: The curves the width is read from. Share of the 64 random perturbations of the centroid that still solve the task against the radius of the noise, relative to each tensor’s norm; mean over the solved checkpoints of the last five tasks, on every classic-control and MiniGrid panel of Figure 4 . The width of Figure 4 is the radius at which a method’s curve crosses the dotted line.
Figure 17: Neighborhood width per panel. Width of the neighborhood around the centroid relative to PPO over the last five tasks, at checkpoints that solve their task, for every method on each classic-control and MiniGrid panel of Figure 2 : geometric mean over trials divided by PPO’s. Above 1 is wider than PPO. Stars: Holm-corrected Mann-Whitney U test against PPO over trials ( ∗ p<0.05 , ∗∗ p<0.01 ).
Figure 18: Neighborhood width over training, by the action-change proxy. Top four rows: the change of the greedy action under weight noise of 0.1 times each tensor’s norm, for the centroid saved at the end of every task, on the thirteen ReLU panels of Figure 4 , solved checkpoints only. The vertical axis is inverted, so up is a wider neighborhood. Rows are the axis of change, columns the environment. Bottom row: the width relative to PPO, the median over the four panels of an environment, with a running mean over three tasks; MiniGrid has one panel and no line. The legend gives the median width relative to PPO over all thirteen panels by this proxy.
ES
GA
Panel
0.02
0.03
0.1
0.02
0.03
0.1
HalfCheetah, offset 0.5
–
1.48∗∗
1.10
–
–
–
HalfCheetah, ground friction
–
1.78∗∗
1.35∗∗
–
1.43∗∗
1.19∗∗
HalfCheetah, sign flipped
–
1.64∗∗
1.19∗∗
–
1.28∗∗
1.04
HalfCheetah, 10 noise tasks
–
1.34∗∗
1.05
–
1.10
0.96
Kinetix, 20 levels
1.24∗
1.30∗
1.26∗
0.81∗
0.85∗
1.12
Appendix
Table 9: Neighborhood width relative to PPO on the tanh panels. Every weight is perturbed by Gaussian noise of standard deviation 0.02 , 0.03 or 0.1 , the same for every method; last five tasks, every checkpoint. Above 1 is wider than PPO. Stars: Mann-Whitney U test against PPO over trials, Holm-corrected over the five panels ( ∗ p<0.05 , ∗∗ p<0.01 ). 0.02 is the search width of ES on Kinetix and is measured there only. The GA has no run on HalfCheetah under noise with two tasks.
Method
σ=0.005
0.01
0.02
0.04
RMS weight
ES
0.00
0.00
0.06
0.23
3.1
GA
0.15
0.17
0.28
0.31
4.8
PPO
0.00
0.02
0.07
0.20
0.12
Appendix
Table 10: Width on the return on Kinetix. Share of random perturbations of absolute standard deviation σ that lose a solved level (score more than 2% of the score range below the unperturbed policy; median over the solved checkpoints at the end of the last two levels, 120 antithetic directions per checkpoint). Lower is wider. 0.02 is the search width of ES. Root-mean-square weight in the last column.
Figure 19: The shared neighborhood in every setting. Share of random moves after which the centroid solves both the task just trained and the next one, in the gymnax and MiniGrid settings of Figure 2 . Left: moves of radius 0.1 of every tensor’s norm, the definition, which Figure 4 (right) plots against LA − F. Right: moves as large in every tensor as the method’s own step over the next task. Crossed: frozen, the method only ever solves one of the two tasks.
Figure 20: Neighborhood width along the method’s own path, relative to random directions of the same size. One point per method and panel of Figure 4 : the multiple of the method’s step over the next task at which half of the re-scored points on the path no longer solve the task just trained, over the same multiple for random moves as large as the step in every tensor; the bar is the median, a filled point a panel where the paired Wilcoxon test over trials is significant (Holm-corrected), stars a Wilcoxon test over panels. Right of 1 , the path stays inside the neighborhood further than random moves do.
Figure 21: Return landscapes around the centroid. One slice per method (columns) and per two-task setting of Figure 2 b with a ReLU network (rows). The circle is the centroid before the switch and the diamond the centroid at the end of the next task; the vertical axis is a random orthogonal direction in units of 0.1 times the norm of the centroid’s weights. Colour marks where a policy solves the previous task (blue), the one just trained (orange) or both (dark), as in Figure 1 . Each checkpoint at a switch is classed by its scores on the two tasks: it keeps both (it learns the new task, reaching half the best method’s gain over an untrained network, and keeps 90% of its gain on the previous one; filled marker), switches (learns the new task but keeps less; white), is frozen (only ever solves one of the two tasks; crossed) or did not learn (grey). A method takes the class of its mean scores over runs, and the checkpoint shown is, among its checkpoints of that class, the one nearest those mean scores; the corner marks the class. A run that did not move over the task is drawn along its last movement.
Acrobot
MountainCar
CartPole
LA
kept
LA
kept
LA
kept
PPO (reported)
0.44
0.29
0.68
0.46
0.79
0.29
PPO, trainer of PBT-PPO
0.61
0.48
0.60
0.47
0.65
0.28
PPO, an eighth of the updates
1.00
0.76
0.88
0.61
0.93
0.18
PPO, learning rate ×0.1
0.91
0.70
0.78
0.61
1.00
0.19
PBT-PPO
0.96
0.65
1.00
0.69
0.96
0.23
Appendix
Table 11: Learning accuracy (LA) and kept return of PPO variants on the ten-task noise settings, rescaled per setting as described at the start of this appendix.
School of Computer Science, McGill University, Montr´eal, Canada · Mila - Quebec Artificial Intelligence Institute, Montr´eal, QC, Canada · CIFAR Learning in Machines and Brains +3