RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
Organizations: University of the Chinese Academy of Sciences · Baidu · Tsinghua University · Institute of Automation, Chinese Academy of Sciences · Peking University · Shanghai University · University of Bristol
Abstract
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
Figures & tables
| Task | Init | Baseline Return | Evolved Return | Fitness | Improvement |
|---|---|---|---|---|---|
| SAC, own-init (17 tasks): evolved vs. our own SAC baseline; evolved § — median | |||||
| dm-cheetah-run | Ours | 302.3 37.3 | 646.9 67.4 | 0.647 | +114.0% |
| dm-walker-walk | Ours | 679.6 41.0 | 971.0 3.3 † | 0.971 | +42.9% |
| dm-walker-run | Ours | 487.7 56.3 | 585.7 7.1 | 0.586 | +20.1% |
| dm-quadruped-walk | Ours | 931.6 12.5 | 939.8 4.1 † | 0.940 | +0.9% |
| gym-HalfCheetah-v5 | Ours | 5093.6 209.3 | 9053.7 1554.1 | 0.604 | +77.7% |
| Ablating PCE | Single-fidelity PPE | |||||
|---|---|---|---|---|---|---|
| Full | w/o decomp. | w/o Phase 1 | L0 | L2 | L1 | |
| (1) | 0.784 0.09 | 0.762 0.13 | 0.758 0.14 | 0.728 0.17 | 0.700 0.08 | 0.691 0.05 |
| (2) | 0.731 0.23 | 0.687 0.25 | 0.704 0.24 | 0.687 0.25 | 0.700 0.21 | 0.694 0.22 |
| (3) | 0.789 0.08 | 0.731 0.17 | 0.748 0.13 | 0.727 0.15 | 0.728 0.10 | 0.736 0.08 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Algorithm | encode | value target | critic loss | policy loss | exploration |
|---|---|---|---|---|---|
| DQN | MLP | TD target | MSE | greedy ∗ | -greedy |
| Rainbow | MLP | -step | distributional | greedy ∗ | NoisyNet |
| DDPG / TD3 | MLP | twin TD target | MSE | DPG | OU / Gaussian |
| A2C / PPO | MLP | GAE | MSE | clipped surrogate | entropy |
| SAC | MLP | soft Bellman | MSE | reparam. + entropy | entropy |
| Level | Steps | Seeds | Eval Eps | Measured wall-clock | Purpose |
|---|---|---|---|---|---|
| L0 | 50K | 1 | 5 | 8–11 min (9.3) | Rapid screening |
| L1 | 200K | 2 | 10 | 60–79 min (78) | Moderate validation |
| L2 | 1M | 3 | 20 | 4.1–9.6 h (5.0) | Final confirmation |
| Task | Seed gap | own gain | SOTA gain | Final gap | Outcome |
| dm-cheetah-run | SOTA-init | ||||
| dm-walker-walk | tie | ||||
| dm-walker-run | SOTA-init | ||||
| dm-quadruped-walk | tie | ||||
| gym-HalfCheetah-v5 | SOTA-init | ||||
| gym-Hopper-v5 | SOTA-init |
| Task | Source | Obs Dim | Act Dim | Norm Bound |
|---|---|---|---|---|
| dm-cheetah-run | DMControl | 17 | 6 | 1000 |
| dm-walker-walk | DMControl | 24 | 6 | 1000 |
| dm-walker-run | DMControl | 24 | 6 | 1000 |
| dm-quadruped-walk | DMControl | 78 | 12 | 1000 |
| gym-HalfCheetah-v5 | Gymnasium | 17 | 6 | 15000 |
| gym-Hopper-v5 | Gymnasium | 11 | 3 | 4000 |
| Source program ( fitness) | |||||
| Task | walker-walk* | quadruped* | walker2d* | Humanoid* | Baseline Fitness |
| dm-cheetah-run | +0.265 | +0.269 | 0.129 | +0.313 | 0.302 |
| dm-walker-walk | +0.292 * | +0.290 | +0.243 | +0.221 | 0.680 |
| dm-walker-run | +0.091 | +0.025 | 0.319 | +0.142 | 0.488 |
| dm-quadruped-walk | +0.019 | +0.006* | 0.669 | 0.082 | 0.932 |
| gym-HalfCheetah-v5 | +0.058 | +0.154 | 0.148 | +0.141 | 0.340 |
| Metric | Mean | Min | Max |
|---|---|---|---|
| L0 evaluations per run | 230 | 230 | 230 |
| L0 L1 promotion rate | 7.6% | 4.8% | 12.2% |
| Code deduplication rate | 18% | 2.5% | 30% |
| Fraction of generated candidate programs discarded as AST-duplicates before screening; | |||
| the denominator is the budget, so 230 survive to L0. | |||
| L1 evaluations per run | 17.5 | 11 | 28 |
| Baseline | Evolved | ||||||
| Task | Seed 0 | Seed 42 | Seed 123 | Seed 0 | Seed 42 | Seed 123 | |
| SAC — DMControl | |||||||
| dm-cheetah-run | 344.0 | 309.4 | 253.5 | 612.4 | 741.1 | 587.3 | |
| dm-walker-walk | 626.0 | 687.4 | 725.5 | 974.3 | 972.4 | 966.5 | |
| dm-walker-run | 553.1 | 415.6 | 494.4 | 583.7 | 595.2 | 578.1 | |
| dm-quadruped-walk | 936.2 | 914.5 | 944.0 | 938.2 | 935.8 | 945.4 | |
| Ablating PCE | Ablating PPE | |||||
| Task | Full | w/o decomp. | w/o Phase 1 | L0-only | L2-only | L1-only |
| SAC | ||||||
| dm-cheetah-run | 0.685 | 0.639 | 0.622 | 0.556 | 0.613 | 0.636 |
| gym-HalfCheetah-v5 | 0.763 | 0.702 | 0.694 | 0.682 | 0.649 | 0.671 |
| md-100envs | 0.907 | 0.945 | 0.954 | 0.957 | 0.754 | 0.728 |
| mw-drawer-open-v3 | 0.780 | 0.763 | 0.763 | 0.718 | 0.782 | 0.730 |
| Condition | dm-cheetah-run | HalfCheetah-v5 | md-100envs | mw-drawer-open | Avg |
|---|---|---|---|---|---|
| Full method | 0.685 0.025 | 0.763 0.048 | 0.907 0.095 | 0.780 0.027 | 0.784 |
| Ablating PCE | |||||
| w/o decomposition | 0.639 0.001 | 0.702 0.083 | 0.945 0.017 | 0.763 0.051 | 0.762 |
| w/o Phase 1 | 0.622 0.018 | 0.694 0.109 | 0.954 0.011 | 0.763 0.029 | 0.758 |
| Ablating PPE | |||||
| L0-only selection | 0.556 0.059 | 0.682 0.071 | 0.957 0.044 | 0.718 0.035 | 0.728 |
| Task | Full | w/o Phase 1 | w/o decomp. |
|---|---|---|---|
| LunarLander-v3 | 0.752 0.027 | 0.638 0.064 | 0.465 0.084 |
| MountainCar-v0 (saturated) | 0.908 0.006 | 0.948 0.009 | 0.909 0.049 |
| Task | Encoding | Value Target | Critic Loss | Policy Loss | Exploration |
| DeepMind Control | |||||
| dm-cheetah-run | , no rescale | Fixed | Huber ( ) | Standard SAC mean L2 | reward, EMA curiosity, action L1 |
| dm-walker-walk | , clip | Sine anneal mean min, then variance penalty | Asym. Huber, EMA-adaptive , 0.55–0.75 | Cosine min/max blend; asym. LR / | Dual-timescale EMA novelty, cosine decay |
| dm-walker-run | + decaying noise | Per-sample min/mean from Q-disagreement | Asym. Huber, quantile-EMA | Softmin Q ( 2 10), PI controller | Convexified reward + movement synergy |
| dm-quadruped-walk | Clip , then symlog | Cosine-annealed max-Q weight | Scaled log-cosh ( ) | Standard SAC action L2 | reward + decaying movement bonus |
| Gym MuJoCo | |||||
| Task | Encoding | Value Target | Critic Loss | Policy Loss | Exploration |
| Meta-World | |||||
| mw-push-v3 | Unused dims distance/velocity features, clip | Softmin blend, , | Asym. Huber, EMA , | Batch-normalized Q + UCB disagreement bonus | Potential shaping ( ) + EMA curiosity |
| mw-drawer-open-v3 | Inject distance/contact/alignment scalars, then symlog | Variance-aware min/mean, pessimism | Quantile-EMA Huber (0.85), TD reweighting | Min-Q decaying disagreement; action penalties | Adaptive potential shaping, decay floor 0.2 |
| mw-door-open-v3 | Velocity rebuild + proximity features, scaled symlog | Min-Q cosine-decayed optimism | Log-cosh Huber, EMA , | ; 0.05 disagreement penalty | Potential shaping + object motion + grasp term |
| mw-peg-insert-side-v3 | History dims velocity/direction features | Min-Q , | Standard Huber, MAD-derived EMA | Standard SAC (+ inactive saturation hinge) | Potential shaping (inverse/exp/grasp), |
| mw-pick-place-v3 | Overwrite dims 18–29 with TCP/object/goal vectors | Min-Q disagreement-gated optimism | Huber, EMA 80th-pct TD error | Standard SAC + mean L2 annealed to 0 | Potential shaping on reach/place/ distances |
| Task | Encoding | Value Target | Critic Loss | Policy Loss | Exploration |
|---|---|---|---|---|---|
| gym-Hopper-v5 | Piecewise symlog above , clip | Cosine-annealed GAE + MC-return blend (weight ) | Smooth-L1, , | Dual clipping ( ) + multiplicatively adapted KL coef. (target 0.015) + decaying entropy § | EMA-normalized state-difference novelty action L2 + survival bonus |
| gym-HalfCheetah-v5 | Symlog | Linear GAE | Huber ( ) decaying | Clip anneal , adv. clamp , KL penalty , entropy | + decaying state-magnitude bonus |
| gym-Ant-v5 | + decaying Gaussian noise | Linear GAE + advantage mean-centering | Adaptive Huber, , | Dual clipping ( ) + KL penalty 0.005 + entropy | Dimension-invariant action penalty + decaying survival bonus |
| gym-Walker2d-v5 | Symmetric clip | Linear GAE | Adaptive- Huber, batch median TD , clamp | Dual clipping ( ) + hinge KL penalty , entropy | Decaying action-energy penalty |
| gym-Humanoid-v5 | Piecewise log above , linear below | Dual- GAE : parallel and streams, averaged 50/50 | Huber, fixed (MSE-like in range, linear on outliers) | Clip anneal + dual clipping + KL-mask pseudo early-stopping + KL penalty | Decaying L1 and L2 action penalties, clip |
| Task | Encoding | Value Target | Critic Loss | Policy Loss | Exploration |
|---|---|---|---|---|---|
| CartPole-v1 | Welford online standardization symlog clip | Unchanged [dead: aggregation] | MSE—unchanged up to a constant [dead: 2nd-critic term] | Never called | Welford-normalized transition distance, positive part, |
| MountainCar-v0 | Symlog EMA mean/var ( ) clip | Unchanged [dead: uncertainty-weighted min/max mixing] | Adaptive expectile regression, weights | Never called | Online RND with random Fourier features, |
| LunarLander-v3 | Symmetric clip | Unchanged [dead: adaptive min/mean blend and the advertised adaptive- warm-up] | Smooth-L1 ( )—a genuine live change from seed MSE | Never called | State-difference bonus (L2, cap 5.0), first 20% of training only, linearly decayed |
| Acrobot-v1 | Clip + negligible decaying noise | Unchanged + live reward clamp (inert for Acrobot) [dead: optimistic min/max mixing] | Asymmetric Huber, under/over weights , | Never called | Online RND (RFF) + NLMS predictor, EMA novelty, decay |