The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Organizations: Meta · Microsoft
Abstract
Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused. In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner's curse of a generation's best candidate. With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting. In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256. Held-out gains grew with the selection set on TREC but not on GSM8K, and the tested acceptance rules did not beat greedy acceptance over whole runs. Gains measured on the selection set also exceeded held-out gains when a current model refined a competent instruction, and in the validation scores of GEPA and MIPROv2. Scoring the starting and the current instruction on 64 items never used for selection removes the average bias of a loop's reported gain, but single estimates remain off by about 6 points. Self-improvement studies should report held-out gains with their uncertainty.
Figures & tables
| Setting | difference (pts) | wins / pairs | |
|---|---|---|---|
| H1: gain( ) gain( ), greedy | |||
| 1.5B TREC | +10.5 2.6 | 7 / 8 | 0.010 |
| 7B TREC | +6.4 0.7 | 8 / 8 | 0.001 |
| 1.5B GSM8K | +1.6 2.2 | 4 / 8 | 0.483 |
| H2: gap( ) gap( ), greedy | |||
| 1.5B TREC | +15.9 2.7 | 8 / 8 | 0.002 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| generation 1 | later generations | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Proposer | Solver, task | cand. | better | better | exc. kurt. | ||||
| 1.5B | 1.5B TREC | 64 | 100% | +6.8 | 8% | -13.5 | 7.1 | 0.44 | -0.4 |
| 1.5B | 1.5B Banking77 | 77 | 0% | -11.5 | 4% | -10.0 | 8.2 | 0.60 | -0.2 |
| 1.5B | 1.5B GSM8K | 55 | 0% | -7.7 | 39% | -2.3 | 7.5 | 0.52 | 3.4 |
| 7B | 7B TREC | 120 | 100% | +4.9 | 15% | -7.2 | 7.1 | 0.58 | 2.7 |
| 7B | 7B Banking77 | 120 | 25% | -0.3 | 38% | -0.2 | 0.6 | 0.57 | 2.5 |
| Setting | Rule ( ) | runs | held-out gain | proxy held-out | harmful / commits | pts | cand. |
|---|---|---|---|---|---|---|---|
| 1.5B TREC | Greedy (16) | 8 | +13.7 2.9 | +17.3 2.5 | 4 / 22 | 3 | 19% |
| Greedy (64) | 8 | +18.5 1.5 | +8.0 1.2 | 8 / 35 | 5 | 20% | |
| Greedy (256) | 8 | +24.2 0.7 | +1.4 0.8 | 8 / 46 | 4 | 12% | |
| McNemar (16) | 8 | +6.3 4.1 | +5.9 2.4 | 0 / 2 | 0 | 2% | |
| McNemar (64) | 8 | +14.6 2.7 | +6.2 2.5 | 2 / 12 | 1 | 9% | |
| Select-confirm (16) | 8 | +15.8 3.2 | +11.3 4.8 | 0 / 10 | 0 | 13% |
| Study | gens 1–3 | gens 4–10 | gens 11–20 | |
|---|---|---|---|---|
| exploratory | 16 | 9.4 | 13.3 | 16.5 2.6 |
| 32 | 6.2 | 8.6 | 10.2 2.8 | |
| 64 | 3.7 | 5.3 | 7.3 1.7 | |
| 128 | 2.3 | 3.0 | 4.2 1.4 | |
| 256 | 1.6 | 2.1 | 2.5 0.7 | |
| confirmatory | 16 | 7.6 | 11.8 | 13.6 1.7 |
| commits | number | runs | observed | model, accepted | model, unconditional | obs. model (95% CI) | |
|---|---|---|---|---|---|---|---|
| 16 | generation 1 | 21 | 21 | 9.1 | 9.0 | 6.0 | +0.0 [-3.4, +4.0] |
| delayed first | 10 | 10 | 6.3 | 9.6 | 6.2 | -3.3 [-8.0, +1.8] | |
| later | 38 | 22 | 5.2 | 9.7 | 5.7 | -4.5 [-6.1, -2.6] | |
| 64 | generation 1 | 27 | 27 | 3.2 | 3.8 | 2.2 | -0.5 [-2.5, +1.5] |
| delayed first | 4 | 4 | 3.5 | 3.8 | 2.1 | -0.2 [-1.9, +1.3] | |
| later | 80 | 28 | 1.9 | 3.7 | 1.9 | -1.8 [-2.5, -1.1] |
| Setting | generation 1 | delayed first | later | |
|---|---|---|---|---|
| 1.5B TREC | 16 | 1.9 / 5.9 (6) | 9.3 / 10.3 (2) | 5.7 / 11.8 (14) |
| 64 | -0.4 / 1.5 (7) | 1.1 / 3.6 (1) | 1.9 / 4.4 (27) | |
| 256 | -0.5 / 0.5 (8) | – | -0.0 / 1.4 (38) | |
| 1.5B Banking77 | 16 | 12.9 / 7.0 (4) | 11.1 / 9.6 (3) | 4.1 / 7.5 (5) |
| 64 | 6.6 / 3.2 (5) | 3.4 / 3.6 (2) | 1.2 / 2.5 (8) | |
| 256 | 0.4 / 1.0 (3) | 1.0 / 1.3 (5) | 1.4 / 1.0 (13) |
| Estimate of the gain | extra evals | |||
|---|---|---|---|---|
| self-report (selection set) | 0 | +10.8 / 15.5 | +6.1 / 9.0 | +1.8 / 4.7 |
| shrinkage by the loop’s own | 0 | +0.7 / 10.7 | -0.7 / 7.6 | -1.4 / 5.3 |
| audit, 16 items | 32 | -0.0 / 10.7 | -0.3 / 11.9 | +0.1 / 12.3 |
| audit, 32 items | 64 | -0.1 / 7.7 | -0.2 / 8.7 | -0.0 / 9.0 |
| audit, 64 items | 128 | -0.0 / 5.7 | -0.1 / 6.3 | -0.0 / 6.7 |
| audit, 128 items | 256 | -0.1 / 4.4 | -0.1 / 4.9 | +0.0 / 5.1 |
| GEPA, metric calls | greedy loop, 20 generations | ||||||
|---|---|---|---|---|---|---|---|
| Task | reported | held-out | inflation | reported | held-out | gap | |
| TREC | 16 | +25.0 4.4 | +1.6 0.6 | +16.4 3.1 | +14.8 3.7 | -1.5 2.7 | +9.1 2.5 |
| 256 | +3.0 1.1 | +0.8 0.8 | -0.9 0.8 | +6.0 1.0 | +4.6 0.9 | +0.7 1.0 | |
| Banking77 | 16 | +25.0 5.2 | +4.1 0.7 | +22.1 2.1 | +11.7 2.2 | +0.2 0.4 | +9.0 3.6 |
| 256 | +3.7 0.6 | +1.9 1.1 | +3.1 1.3 | +3.6 0.7 | +2.0 0.5 | +1.2 0.6 | |
| Start | Task | cand. | reported gain | held-out gain | overstatement | Prop. 1 | ||
|---|---|---|---|---|---|---|---|---|
| one-line seed | TREC | 16 | 2250 | 91 | +45.0 7.2 | +18.5 1.2 | +26.5 6.9 | 14.1 |
| 64 | 32 | +26.6 4.0 | +20.4 1.6 | +6.2 3.9 | 5.4 | |||
| 256 | 9 | +17.0 2.6 | +19.0 1.9 | -1.9 2.2 | 0.8 | |||
| Banking77 | 16 | 2250 | 91 | +22.5 3.2 | +3.7 1.4 | +18.8 2.5 | 12.6 | |
| 64 | 32 | +11.2 2.3 | +1.3 1.1 | +10.0 2.1 | 5.5 | |||
| 256 | 9 | +4.5 1.0 | +0.8 0.6 | +3.7 0.7 | 1.8 |
| Task | runs | inflation | reported gain | held-out gain | regret | |
|---|---|---|---|---|---|---|
| TREC | 16 | 5 | +14.1 2.7 | +38.8 4.6 | +19.2 2.0 | +4.9 1.9 |
| 64 | 5 | +2.4 2.6 | +24.4 4.3 | +19.3 1.7 | +2.1 0.6 | |
| 256 | 5 | -4.5 1.3 | +15.9 2.9 | +18.9 1.5 | +0.2 0.1 | |
| Banking77 | 16 | 5 | +24.6 5.0 | +18.8 2.8 | +1.2 1.7 | +1.8 0.6 |
| 64 | 5 | +7.3 2.9 | +7.5 2.2 | +2.4 1.0 | +0.3 0.3 | |
| 256 | 5 | +2.4 0.6 | +2.9 0.7 | +1.2 0.6 | +0.1 0.1 |
| Task | runs | inflation | reported gain | held-out gain | regret | |
|---|---|---|---|---|---|---|
| TREC | 64 | 5 | +4.5 3.0 | +28.8 2.8 | +21.6 1.0 | +1.4 0.6 |
| 256 | 5 | -0.0 1.1 | +19.5 2.4 | +18.1 1.8 | +2.0 0.8 | |
| Banking77 | 64 | 5 | +12.7 1.6 | +13.4 2.8 | +2.9 1.3 | +2.3 0.7 |
| 256 | 5 | +5.2 0.5 | +5.1 1.0 | +0.5 0.7 | +1.8 0.3 |
| Task | cand. | inflation | reported gain | held-out gain | overstatement | Prop. 1 | |
|---|---|---|---|---|---|---|---|
| TREC | 16 | 15.8 | +0.2 6.5 | +11.2 4.1 | +5.8 2.1 | +5.5 2.6 | 9.0 |
| 256 | 14.4 | -0.8 0.6 | +12.1 3.1 | +11.4 2.3 | +0.7 1.3 | 0.9 | |
| Banking77 | 16 | 15.8 | +13.8 2.9 | +5.0 2.3 | -1.6 0.9 | +6.6 3.3 | 7.7 |
| 256 | 15.4 | +3.0 0.4 | +2.5 0.3 | +0.1 0.5 | +2.4 0.4 | 1.9 |
| Rule | ||||
|---|---|---|---|---|
| Greedy keep-if-better | -0.47 [-0.82, -0.03] | -0.08 [-0.46, +0.37] | +0.24 [-0.15, +0.71] | +0.50 [+0.11, +0.96] |
| McNemar, | +0.06 [+0.00, +0.14] | +0.20 [+0.04, +0.43] | +0.34 [+0.09, +0.67] | +0.49 [+0.18, +0.88] |
| McNemar, | +0.12 [-0.05, +0.34] | +0.31 [+0.04, +0.63] | +0.43 [+0.12, +0.82] | +0.58 [+0.25, +0.99] |
| PACE e-process | +0.02 [+0.00, +0.04] | +0.09 [+0.01, +0.23] | +0.25 [+0.04, +0.52] | +0.43 [+0.13, +0.79] |
| Select-then-confirm | +0.06 [-0.09, +0.26] | +0.21 [-0.03, +0.52] | +0.35 [+0.07, +0.71] | +0.55 [+0.21, +0.95] |
| McNemar, pilot-tuned | +0.19 [-0.02, +0.43] | +0.44 [+0.12, +0.82] | +0.56 [+0.17, +1.02] | +0.64 [+0.25, +1.10] |
| Rule | ||||
|---|---|---|---|---|
| Pilot from the other model (same task; 4 settings) | ||||
| Greedy | -0.62 [-1.08, -0.08] | -0.10 [-0.57, +0.47] | +0.24 [-0.26, +0.83] | +0.48 [+0.01, +1.06] |
| McNemar | +0.06 [+0.01, +0.15] | +0.16 [+0.01, +0.34] | +0.33 [+0.03, +0.75] | +0.51 [+0.09, +1.06] |
| McNemar, tuned | -0.05 [-0.19, +0.14] | +0.28 [-0.08, +0.74] | +0.45 [+0.01, +1.00] | +0.56 [+0.12, +1.13] |
| Threshold, tuned | +0.00 [-0.20, +0.28] | +0.16 [-0.15, +0.56] | +0.46 [+0.02, +1.03] | +0.51 [+0.07, +1.08] |
| Pilot Bayes gate | +0.25 [-0.08, +0.66] | +0.39 [-0.00, +0.90] | +0.50 [+0.05, +1.06] | +0.62 [+0.18, +1.19] |
| Setting | Rule | runs | held-out gain | proxy held-out | harmful / commits |
|---|---|---|---|---|---|
| 1.5B TREC, | Greedy | 3 | +15.8 3.6 | 15.8 | 2 / 10 |
| McNemar | 3 | +11.3 5.8 | 7.7 | 0 / 2 | |
| E-process | 3 | +0.0 0.0 | -1.8 | 0 / 0 | |
| EB gate | 3 | +5.6 1.5 | 17.6 | 0 / 3 | |
| 1.5B TREC, | Greedy | 3 | +17.8 0.7 | 7.4 | 2 / 13 |
| McNemar | 3 | +12.6 2.4 | 1.6 | 0 / 3 |
| Setting | Rule | runs | held-out gain | proxy held-out | harmful / commits |
|---|---|---|---|---|---|
| 7B TREC, | Greedy | 3 | +5.9 1.7 | 16.0 | 1 / 5 |
| McNemar | 3 | +1.0 1.0 | 10.7 | 0 / 1 | |
| EB gate | 3 | +8.9 4.2 | 25.7 | 3 / 7 | |
| 7B TREC, | Greedy | 3 | +15.4 3.2 | 10.9 | 4 / 14 |
| McNemar | 3 | +8.9 3.0 | 6.9 | 0 / 3 | |
| E-process | 3 | +5.9 5.9 | 2.6 | 0 / 1 |
| Model | |||||
|---|---|---|---|---|---|
| 1.5B TREC | +16.3 1.9 (3) | +13.6 2.0 (3) | +11.6 0.9 (3) | +14.8 1.3 (3) | +13.7 2.2 (3) |
| 7B TREC | +8.8 4.2 (2) | +11.5 7.3 (2) | +13.9 4.2 (3) | +12.6 5.3 (2) | +6.6 3.8 (2) |
| Setting | Rule | runs | held-out gain | proxy held-out | harmful / commits |
|---|---|---|---|---|---|
| 1.5B TREC, | Greedy | 3 | +9.1 2.5 | 14.4 | 1 / 9 |
| McNemar | 3 | +0.0 0.0 | -3.6 | 0 / 0 | |
| E-process | 3 | +0.0 0.0 | -3.6 | 0 / 0 | |
| EB gate | 3 | +10.8 4.0 | 4.3 | 0 / 5 | |
| EB, no safeguard | 2 | +12.8 7.4 | 8.7 | 0 / 5 | |
| 1.5B TREC, | Greedy | 3 | +21.3 1.4 | 6.9 | 4 / 16 |