Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
Organizations: Purdue University
Abstract
As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dependence. To address this, we propose REUSE (Risk-controlled Evaluation Under Sequential Evolution), a certified evaluation and promotion framework that allows a fixed evaluation set to support repeated adaptive decisions while providing statistical guarantees. For a user-specified error level , with probability at least , every promoted modification is a genuine population improvement on the underlying task distribution. REUSE achieves this by strictly limiting the evaluation feedback returned to the search process and accounting for possible promotion histories within the error budget. We develop detailed statistical theory for RSI evaluation in this setting, including simultaneous error control, valid lower bounds on cumulative improvement, and a characterization of the fundamental limits of adaptive evaluation reuse. In live self-improvement experiments, REUSE commits substantially fewer false promotions than evaluation frameworks from current RSI systems and error-controlled baselines, reducing the proportion of false promotions from up to 20.7% to 0%, while achieving final true population performance comparable to the best baselines.
Figures & tables
| Method | Select. | Repeat. | Adapt. | Pop. Impr. | Prom. | False | Opt. Gap | Pop. Impr. | Prom. | False | Opt. Gap | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Empirical best-of- | ✗ | ✗ | ✗ | 7.04 0.08 | 13.0 | 75 | 0.43 0.11 | 7.19 0.06 | 14.2 | 71 | 0.15 0.06 | |
| SICA-style | ✗ | ✗ | ✗ | 6.36 0.18 | 12.3 | 72 | 0.34 0.10 | 6.56 0.16 | 12.5 | 62 | 0.14 0.05 | |
| DGM-style | ✗ | ✗ | ✗ | 3.89 0.17 | 6.4 | 34 | 0.06 0.11 | 4.03 0.18 | 7.2 | 18 | 0.01 0.05 | |
| Elite | ✗ | ✗ | ✗ | 6.72 0.14 | 14.3 | 89 | 0.47 0.12 | 6.85 0.12 | 13.7 | 59 | 0.16 0.05 | |
| Niche-elite | ✗ | ✗ | ✗ | 6.46 0.15 | 12.3 | 71 | 0.40 0.13 | 6.37 0.15 | 12.4 | 44 | 0.13 0.05 | |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Setting |
|---|---|
| Data | |
| Training set | 50,000 examples |
| Development set | 20,000 examples |
| Evaluation pool | 60,000 examples |
| Held-out set | 50,000 examples |
| Evaluation-set size | or |
| Information in the prompt | Compared methods | Reuse |
|---|---|---|
| Parent configuration | ✓ | ✓ |
| Accuracy and slice error rates on | ✓ | ✓ |
| Accuracies of past proposals on | ✓ | ✓ |
| Evaluation accuracies on | ✓ | ✗ |
| Promotion decisions | ✓ | ✓ |
| Development-draft updates and submissions | ✗ | ✓ |
| Method | Feedback from | Promotion rule | Select. | Repeat. | Adapt. |
|---|---|---|---|---|---|
| Empirical best-of- | scores | largest positive gain on | ✗ | ✗ | ✗ |
| SICA-style | scores | largest positive gain on | ✗ | ✗ | ✗ |
| DGM-style | scores | largest positive gain on | ✗ | ✗ | ✗ |
| Elite | scores | largest positive gain on | ✗ | ✗ | ✗ |
| Niche-elite | scores | largest positive gain on | ✗ | ✗ | ✗ |
| PACE-style | scores | McNemar test at level | ✗ | ✗ | ✗ |
| Diff. vs Reuse | Worst | ||||
|---|---|---|---|---|---|
| Method | |||||
| Empirical best-of- | 0.25 | 0.15 | |||
| SICA-style | 0.29 | 0.19 | |||
| DGM-style | 0.31 | 0.23 | |||
| Elite | 0.28 | 0.18 | |||
| Niche-elite | 0.34 | 0.19 | |||
| tests | Prom. | False / | Pop. Impr. | Diff. | Changed | at default’s | |
| Setting | per run | per run | all | (pp) | (pp) | runs | promotions (pp) |
| Default | 4.30 0.37 | 2.90 0.28 | 0 /29 | 6.38 0.23 | N/A | N/A | 0.74 0.11 |
| Split | |||||||
| 4.00 0.30 | 3.00 0.15 | 0 /30 | 7.23 0.36 | 0.40 | 9/10 | 0.89 0.11 | |
| 3.90 0.18 | 2.60 0.16 | 0 /26 | 6.52 0.27 | 0.37 | 4/10 | 0.84 0.11 | |
| 4.80 0.33 | 3.00 0.26 | 0 /30 | 6.52 0.35 | 0.29 | 1/10 | 0.59 0.10 | |
| Method | Prom. | False | Runs | Prom. | False | Runs | |
|---|---|---|---|---|---|---|---|
| Empirical best-of- | 13.0 | 75 | 30/30 | 14.2 | 71 | 28/30 | |
| SICA-style | 12.3 | 72 | 30/30 | 12.5 | 62 | 26/30 | |
| DGM-style | 6.4 | 34 | 22/30 | 7.2 | 18 | 15/30 | |
| Elite | 14.3 | 89 | 27/30 | 13.7 | 59 | 25/30 | |
| Niche-elite | 12.3 | 71 | 28/30 | 12.4 | 44 | 24/30 | |
| Empirical | SICA- | DGM- | Niche- | PACE- | Bonferroni- | Reuse | ||||||||||
| Seed | best-of- | style | style | Elite | elite | style | style | (ours) | ||||||||
| 3 | 7.14 | ( 1 /9) | 5.76 | ( 3 /11) | 4.65 | ( 1 /8) | 7.14 | ( 3 /13) | 7.17 | ( 3 /20) | 7.22 | (0/9) | 0.00 | (0/0) | 7.24 | (0/3) |
| 4 | 7.25 | ( 2 /12) | 7.24 | ( 3 /14) | 5.26 | (0/5) | 7.25 | (0/9) | 7.24 | ( 2 /9) | 5.75 | (0/5) | 3.21 | (0/2) | 5.73 | (0/3) |
| 5 | 7.06 | ( 3 /14) | 7.41 | ( 3 /16) | 2.77 | ( 2 /7) | 4.65 | ( 4 /15) | 6.81 | ( 1 /14) | 3.79 | (0/5) | 0.00 | (0/0) | 7.26 | (0/4) |
| 6 | 6.70 | ( 3 /16) | 5.55 | ( 2 /11) | 2.33 | ( 2 /4) | 6.70 | ( 5 /19) | 6.70 | ( 6 /19) | 6.72 | ( 1 /8) | 4.11 | (0/2) | 6.70 | (0/4) |
| 7 | 7.14 | ( 3 /15) | 5.92 | ( 2 /14) | 3.21 | ( 1 /9) | 5.92 | ( 5 /18) | 6.50 | ( 3 /13) | 4.65 | (0/4) | 0.86 | (0/1) | 10.00 | (0/5) |
| Empirical | SICA- | DGM- | Niche- | PACE- | Bonferroni- | Reuse | ||||||||||
| Seed | best-of- | style | style | Elite | elite | style | style | (ours) | ||||||||
| 3 | 7.17 | ( 1 /17) | 4.65 | (0/5) | 2.92 | ( 1 /7) | 7.41 | ( 4 /20) | 7.24 | ( 2 /10) | 7.24 | (0/9) | 7.10 | (0/7) | 6.70 | (0/5) |
| 4 | 7.37 | ( 1 /13) | 7.44 | ( 2 /14) | 5.27 | ( 1 /8) | 7.27 | ( 3 /16) | 5.61 | (0/7) | 7.37 | (0/7) | 7.24 | (0/7) | 7.25 | (0/8) |
| 5 | 6.64 | ( 2 /14) | 7.37 | ( 2 /13) | 2.19 | (0/5) | 6.45 | ( 2 /8) | 5.80 | (0/9) | 7.24 | (0/9) | 7.10 | (0/7) | 7.37 | (0/6) |
| 6 | 7.41 | ( 2 /15) | 7.34 | ( 2 /13) | 5.46 | (0/7) | 7.24 | ( 1 /9) | 6.57 | ( 2 /20) | 7.27 | (0/10) | 4.65 | (0/4) | 7.34 | (0/7) |
| 7 | 7.41 | ( 3 /17) | 5.92 | ( 1 /13) | 4.65 | ( 1 /8) | 6.62 | ( 4 /21) | 5.18 | (0/13) | 7.24 | (0/9) | 7.05 | (0/7) | 7.24 | (0/6) |
| Final population improvement | Promotions per run | Running certificate | Direct certificate | Valid / | |
|---|---|---|---|---|---|
| 7.02 | 3.2 | 0.92 [0.13, 1.76] | 3.10 [1.16, 5.49] | 30 / 30 | |
| 7.19 | 5.9 | 2.67 [2.06, 3.88] | 5.33 [4.34, 6.92] | 30 / 30 |