Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent frameworks often have every model work toward a shared goal. We ask a different question. When each agent pursues its own reward, can self-improving LLMs learn from one another well enough to improve the whole population? We call this capability recursive social improvement. We study populations that revise skill files and choose whether, when, and whom to copy from. Independent search, learning from peers, and acting all share one token budget. In controlled environments, established social-learning algorithms benefit from peers, but three LLMs do not. They earn less reward per token than solo learners, and explore too narrowly or run out of tokens before acting. We then let the models write and revise their own skills. Observing peers changes how they improve, helping one model find useful skills sooner and another spend less on private search. Neither, however, outperforms independent learners at the same cost. Skills are copied, revised, and passed on, so one discovery can seed further search. Yet these exchanges concentrate the population around fewer independent discoveries. Together, these results show that LLMs can make learning more efficient by copying from peers, but not yet more effective.
Figures & tables
Figure 1: Recursive social improvement. An agent can copy a peer’s skill file, revise one of its own, or skip learning. Copying and revision add alternatives to its library of reusable task-solving instructions. It then chooses any saved skill, including an earlier version, and uses it to solve a task. Feedback informs the next round. The entire process shares one finite token budget. Peers follow the same loop independently, each pursuing its own task reward.
Figure 2: Mean reward per agent at each round. Panels (a)–(c) show Qwen3-14B, Ministral-3-14B, and GPT-OSS-20B; panel (d) shows solo UCB and hierarchical social UCB, which use no LLM inference. All panels use the same eight environment seeds. Blue denotes solo and orange social; the black dashed line shows the oracle. Shading shows seed SE, and missing pulls earn zero.
Figure 3
Figure 5: Held-out whole-answer accuracy during skill evolution. Panels (a) and (b) compare LLM-controlled solo and social learning with full solo OpenEvolve for each model. The dotted reference is the LLM-solo initial skill. Lines show means and shaded bands show SE across six population seeds.
LLM social − LLM solo accuracy (pp)
Model
Trajectory
Final
ΔJ(C)
Observe rate (social)
GPT-OSS-120B
+0.19±0.59
−0.58±0.88
+0.14±0.35
35.92±1.40%
GLM-5.3-Flash
+1.31±0.61
+0.95±1.15
+0.98±0.85
8.37±0.90%
Table 1: Social minus solo held-out whole-answer accuracy, in percentage points. Trajectory averages rounds 0, 10, 25, and 49; Final is round 49; ΔJ(C) compares equal population-wide learning tokens (prompts and completions, excluding hidden evaluation). Observe rate is the fraction of learning rounds social agents spend observing. Trajectory and final SE account for both population seeds and held-out examples (Appendix D.2 ); other SE are across six seeds.
Figure 6: The same number of observations has different effects depending on timing. Panels (a)–(b) compare four assigned early observations with four distributed observations and solo learning. Models choose whom to observe and which skill to execute. Shading is SE across six population seeds. The optional early-access control and paired statistics appear in Appendix G .
Figure 7: Final held-out accuracy (a–b) and end-to-end token use, including held-out evaluation (c–d), per five-agent population. OpenEvolve denotes full solo OpenEvolve. Bars: six-seed means ± SE.
Appendix figures & tables61 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Audit of the stationary finite-bandit payoff distribution over eight matched seeds. Panel (a) shows arm means sorted within seed, averaged with seed SE. Panel (b) shows the best-minus-second and best-minus-median gaps for every seed; the dotted line is the standard deviation of one reward draw.
Figure 9: Realized per-round social-minus-solo reward at matched token expenditure. Panels (a)–(c) show Qwen, Ministral, and GPT-OSS. Each dot is one paired environment seed; black markers and bars show mean ± SE across eight seeds. The primary comparison interpolates against completion tokens; alternatives use the last completed round or interpolate against prompt-plus-completion tokens. Missing pulls earn zero.
Figure 10: Cumulative reward per agent over the same eight environment seeds in all panels. Panels (a)–(c) show the three LLM families; (d) shows solo and hierarchical social UCB. Every panel includes the same oracle, shown with a black dashed line and light-gray SE shading. Colored shading is seed SE.
Figure 11: Precision sensitivity for hierarchical versus solo UCB. The original eight environments are a subset of the fixed 100-seed extension. Paired mean differences and SE quantify uncertainty across environments, not additional information available to agents.
Figure 12: Cumulative expected regret for the policies in Figure 2 . For an invalid or missing pull, the selected arm has expected payoff zero. Lines average agents within population seed and bands are seed SE over the same eight environments for all policies.
Figure 13: Local copy quality under the primary expiring-budget condition. Panel (a) shows the percentage of socially learned pulls whose arm mean exceeds the best arm the copier had personally discovered. Panel (b) shows the share of socially learned pulls for which that personal comparator exists. Bars are seed means with seed SE over eight populations.
Figure 14: Finite-bandit reward and arm-use trajectories. Panels (a)–(c) show reward per round for Qwen, GPT-OSS-20B, and Ministral. Panels (d)–(f) show the number of distinct arms used by the corresponding populations. Lines are means and bands are seed SE.
Figure 15: Primary finite-bandit outcomes under the stationary, neutral-prompt, expiring-budget condition. Columns show Qwen, GPT-OSS, and Ministral. Panels (a)–(c) show population reward; (d)–(f) reward conditional on a valid pull; (g)–(i) pull completion. Missing pulls earn zero in population reward. Bars show eight-seed means and SE.
Figure 16: Mechanism estimates for the finite bandit. Panel (a) is PE’s effect on completion within social populations. Panel (b) is its completion interaction from Equation 4 . Panel (c) is IS’s effect on effective selected-arm diversity within social populations. Panel (d) is its reward interaction. Bars show paired means with seed SE over eight seeds.
Figure 17: Effects of independent search on Ministral. Panels (a)–(b) separate the treatment effect within solo, the effect within social, and their difference for reward and completion. Each uses eight matched causal-control seeds, without pooling the original default replications. Bars show paired mean effects and SE.
Figure 18: Hierarchical social UCB versus solo UCB under stationary and changing payoffs. Panel (a) shows cumulative reward; (b) shows the paired social-minus-solo contrast. Bars show means and seed SE. Both algorithms reset their estimates at the known payoff-change boundaries, so this comparison does not test their ability to detect changes.
Figure 19: Mean reward per round for the interventions in Figure 4 , averaged over all 100 rounds rather than measured only at the endpoint. Panels (a)–(c) show the three models. Multiplying these means and SE by 100 gives the main cumulative-reward figure; the two summaries therefore yield the same ranking.
Figure 20: Cumulative expected regret through round 100 for the same populations and interventions as Figure 4 . Panels (a)–(c) show the three models. Lower regret is better; bars show means and SE across run replicates.
Figure 21: Cumulative reward trajectories underlying Figure 4 . Columns show the three models; rows show default, protected execution, and independent search. Lines show means and shading shows SE across run replicates. These curves distinguish improvement throughout learning from an effect confined to its endpoint.
Figure 22: Continuous-bandit stationary structured environment. Panel (a) shows absolute reward for solo and payoff-visible social populations. Panel (b) shows the independent discovery supply. Error bars are seed SE over matched populations.
Figure 23: Continuous-bandit payoff-visible social minus solo reward when payoffs are stationary or change every 25 or 10 rounds. Panels (a)–(c) show Qwen, GPT-OSS-20B, and Ministral. Points are paired means with seed SE.
Figure 24: Continuous-bandit trajectories in the stationary structured environment. Panels (a)–(c) show reward per round for Qwen, GPT-OSS-20B, and Ministral. Panels (d)–(f) show cumulative independently discovered arms for the same models. Lines are means and bands are seed SE.
Figure 25: Cumulative reward in the stationary continuous landscape. Panels (a)–(c) show Qwen, GPT-OSS, and Ministral under the primary expiring-budget condition, with population-seed means and SE.
Figure 26: Cumulative expected regret relative to the best value on the saved dense reference grid. Panels (a)–(c) show the same models and actions as Figure 25 . The continuous-space oracle is approximate; these are not exact regret bounds over infinitely many arms.
Figure 27: Paired social-minus-solo reward for DEOE across completed finite-bandit and continuous-bandit environments. Blue points expose actions only; pink points also expose payoff. Error bars are seed SE over eight matched populations.
Figure 28: Illustration of the fixed-skill scheduling task. The benchmark uses 16 jobs and up to eight slots. Invalid schedules receive zero.
Figure 29: Fixed-skill population reward before and after equalizing access to all eight skills. Panels (a)–(c) show Qwen, GPT-OSS-20B, and Ministral. Gray bars are the original restricted-acquisition design; blue bars enable independent acquisition in both solo and social conditions. Error bars are seed SE.
Figure 30: Fixed-skill reward when the LLM or DEOE selects a skill and the same LLM executes the schedule. Panels (a)–(c) show Qwen, GPT-OSS-20B, and Ministral. Bars are means with seed SE over eight seeds.
Figure 31: Ancestry of deployed skills, with population means and seed SE. Columns show GPT-OSS and GLM. Panels (a)–(b) count private revision steps in a deployed file’s ancestry; (c)–(d) count deployments whose ancestry includes copying a non-initial skill followed by another revision; (e)–(f) count distinct founding revision lineages currently deployed by the five agents. This is recorded first-acquisition ancestry, not all possible paths to identical file content.
Figure 32: Training-score changes from private versus socially derived parents in LLM-social populations. Panels (a)–(b) show GPT-OSS and GLM. Points and SE first average scored proposals within population seed. Parent fitness is the recorded score available before revision, not a simultaneous fresh evaluation. These are proposal gains, not gains of the best selected skill.
Figure 33: Seed and test-item uncertainty for LLM-social minus LLM-solo whole-answer accuracy. Panels (a)–(b) show GPT-OSS and GLM. The same estimates are shown with conventional seed SE, item-only bootstrap SE, and joint seed-and-item bootstrap SE. Trajectory is the trapezoidal average over rounds 0, 10, 25, and 49; initial-adjusted subtracts each condition’s starting accuracy.
Figure 34: Earlier single-agent weak-skill capability checks. Panels (a)–(b) show whole-answer accuracy for GPT-OSS and GLM; (c)–(d) show individual-constraint accuracy. Frozen and evolved use eight seeds each and the same 200 hidden tasks. Replay evaluates the starting file saved inside the evolved run: eight GPT-OSS seeds and seven complete GLM seeds. The remaining GLM replay is incomplete and excluded rather than averaged over a different task denominator. Error bars are seed SE. These runs are not pooled with the corrected population experiment.
Figure 35: Final minus initial held-out accuracy, paired within population seed. Panels (a)–(b) show whole-answer accuracy gains for GPT-OSS and GLM; (c)–(d) show gains on individual constraints. Bars show six-seed means with SE.
Figure 36: Final-minus-initial constraint-accuracy gains on training and held-out examples. Panels (a)–(b) show GPT-OSS and GLM. Bars show matched within-population gains averaged across six seeds, with SE. Training gains are larger, but differences in token caps and reuse of selection scores prevent interpreting the gap as overfitting alone.
Figure 37: Recorded training fitness of deployed skills over learning, for GPT-OSS (a) and GLM (b). This is the score available to the learner, not an independently re-estimated hidden score. All curves average agents within population and show six-seed mean and SE.
Figure 38: Paired controller differences in percentage points. Panels (a)–(b) show GPT-OSS final and trajectory-average accuracy; (c)–(d) show the corresponding GLM contrasts. Small points are matched population seeds; larger points and whiskers are the paired mean and SE. UCB and uniform are compared with full solo OpenEvolve, and with each other.
Figure 39: Held-out individual-constraint accuracy at the same checkpoints as the main whole-answer metric. Panels (a)–(b) show GPT-OSS and GLM. The dotted reference is the initial LLM-solo skill, and shaded regions are population-seed SE.
Figure 40: Overall observation and copied-file deployment rates. Panels (a)–(b) show observation as a fraction of learning rounds; (c)–(d) show deployments of directly acquired social files. Denominators include all post-initialization agent-rounds. Bars show population-seed mean and SE.
Figure 41: Trajectory-average contrasts after subtracting each run’s initial whole-answer accuracy. Panels (a)–(b) show GPT-OSS and GLM. Points and error bars show the paired mean and SE across six populations. This controls the starting offset, not every possible difference in subsequent search.
Figure 42: Observation and use of copied skills over learning. Panels (a)–(b) show the fraction of rounds spent observing; (c)–(d) show deployment of a file first acquired socially. Curves average non-overlapping five-round blocks within seed (four rounds in the last block), then show seed mean and SE. Revising a copied file creates a locally acquired descendant, which is not counted as direct copied-file deployment.
Figure 43: When observations occur. Panels (a)–(b) show the mean observation round; (c)–(d) show the percentage of all observations occurring in rounds 1–10. Each statistic is computed within a population before averaging across six seeds. Error bars are seed SE.
Figure 44: How often agents observe within a time interval. Panels (a)–(b) divide observations by the available agent-rounds in rounds 1–10; (c)–(d) use rounds 11–49. Unlike Figure 43 c–d, the denominator includes rounds with no observation. Bars show mean and SE across population seeds.
Figure 45: Private and social additions to skill libraries. Panels (a)–(b) show successful revisions per agent; (c)–(d) show newly acquired social skill versions per agent. Repeated observations of an already acquired file do not count as new acquisitions. Agents are averaged within population before computing six-seed means and SE.
Figure 46: Saved selected skills versus training-ranked alternatives. Panels (a)–(b) show selected-skill test accuracy; (c)–(d) show alternative-skill accuracy. Both use the same held-out tasks. Bars are population means with SE; these marginal bars do not by themselves estimate a paired-effect SE.
Figure 47: Original versus corrected population experiments, displayed separately. Panels (a)–(b) show GPT-OSS and GLM final whole-answer accuracy. Each bar uses six population seeds with SE. These are different implementations, not additional interchangeable replicates.
Model
Learning tokens
All tokens
Learning USD
All USD
GPT-OSS-120B
−17.34±2.08∗∗∗
−22.80±9.04 n.s.
−2.25±0.15∗∗∗
−2.13±0.19∗∗∗
GLM-5.3-Flash
−1.74±3.32 n.s.
−4.64±11.67 n.s.
−0.51±0.23 n.s.
−0.30±0.27 n.s.
∗∗∗p<0.001 ; n.s.: p≥0.05 . Negative values mean social uses fewer resources.
Appendix
Table 2: Resource use for LLM-social minus LLM-solo populations. Tokens include both prompts and completions and are reported in millions. “Learning” includes search, revision, selection, and training-time execution; “all” additionally includes every held-out evaluation. USD uses the provider-reported cost for the same calls. Effects are paired means ± SE across six seeds.
Figure 48: Matched learning-token comparison from saved checkpoints. Panels (a)–(b) show GPT-OSS and GLM. Colored points are paired population-seed contrasts; black points and whiskers are means and SE. The middle comparison interpolates both policies at their common token endpoint. The neighboring comparisons retain social at that target and use solo’s preceding or following checkpoint instead. Tokens include prompts and completions across the population, excluding hidden evaluation. No curve is extrapolated.
Figure 49: Realized completion-token use as a percentage of the round allowance. Panels (a)–(b) show GPT-OSS and GLM. Bars average all initialization and learning agent-rounds within seed, then show seed means and SE. Crosses show each seed’s maximum agent-round fraction, all below the dotted aggregate limit. This does not measure individual-call cap hits or the tokens reserved before a call.
Model
Prompt
Completion
Total
GPT-OSS-120B
−6.87±1.55∗∗
−10.46±0.74∗∗∗
−17.34±2.08∗∗∗
GLM-5.3-Flash
+2.51±2.75 n.s.
−4.25±0.76∗∗
−1.74±3.32 n.s.
∗∗p<0.01 , ∗∗∗p<0.001 ; n.s.: p≥0.05 .
Appendix
Table 3: Learning-token decomposition for LLM-social minus LLM-solo populations, in millions of tokens. Prompt and completion columns sum to the total column. Values are paired means ± SE across six seeds.
Figure 50: Performance and spending during skill evolution. Panels (a)–(b) show GPT-OSS and GLM on separate vertical scales for comparison with each model’s own solo policy. Each point is a saved checkpoint at mean cumulative learning expense per five-agent population, with vertical accuracy SE across six seeds. Higher and further left is preferable. Costs include private discovery throughout the population and exclude hidden evaluation. OE denotes OpenEvolve.
Figure 51: Reported API expense by phase, averaged over six five-agent populations per condition; error bars show SE of the total. Panels (a)–(b) show GPT-OSS and GLM. Hidden evaluation is included here but excluded from the matched-learning-token comparison in Appendix E.1 . Unconfirmed usage reservations are not attributed to these reported-cost components.
Figure 52: Held-out accuracy against cumulative learning tokens, including prompts and completions across all five agents and excluding hidden evaluation. Panels (a)–(b) show GPT-OSS and GLM, with separate vertical scales. Points are saved checkpoints at mean token expenditure; shading shows accuracy SE across six population seeds. Matched-token contrasts are computed within each seed, not from these averaged curves.
Figure 53: Completion tokens used to execute the selected skill, by acquisition origin and learning period. Early is rounds 1–10; late is 11–49. Columns compare LLM social, source UCB, and uniform source selection; rows show GPT-OSS and GLM. Conditional means are computed within population before taking six-seed means and SE. All cells have six contributing populations; their numbers of deployments differ. These are execution costs, not tokens spent deciding to observe.
Figure 54: LLM-social tokens spent choosing a skill to deploy, conditional on the selected file’s acquisition origin. Panels (a)–(b) show GPT-OSS and GLM; periods and uncertainty match Figure 53 . This decision considers a library, so its tokens cannot be attributed solely to reading the selected file. UCB and uniform controllers choose deployment algorithmically and make no corresponding LLM decision call.
Figure 55: Source quality at observed rounds in LLM-social populations. Panels (a)–(b) show GPT-OSS and GLM. First peer minus others compares the first eligible ID with the remaining peers; chosen minus available compares the selected peer with the mean across all four eligible peers. Small points are population-seed averages; black points show their mean and SE. These are retrospective diagnostics, not additional information revealed before source selection.
Figure 56: Separation of peer training scores at realized observations. Panels (a)–(b) show GPT-OSS and GLM. For each observation, we compute the standard deviation across four available scores, then average observations within seed. Bars show seed means and SE. Different controllers generate different observation times and candidate pools.
Figure 57: Whom LLM agents observe and how much computation their decisions use. Panels (a)–(b) show each peer’s share of observations in rounds 1–10 and 11–49. IDs follow the order of the available-peer list; they do not encode skill quality. Panels (c)–(d) show completion tokens conditional on observation: the learning decision before receiving the file and the deployment choice afterward. Bars show population means with seed SE; decision-cost panels have separate vertical scales. All calls, including unusually long ones, contribute to the means.
Figure 58: Source repetition and learning-decision cost over time. Panels (a)–(b) condition on an observation with an earlier source to compare against; the dotted line is the uniform-repeat probability. Panels (c)–(d) compare learning-selection completion tokens for social observation, other social actions, and solo actions. Curves show means with SE across eligible population seeds in five-round blocks. Late GLM observations are sparse, so conditional estimates can use fewer seeds and should be read with the coverage saved in the accompanying data.
Figure 59: Breadth and concentration of observation. Panels (a)–(b) show the cumulative fraction of possible directed links used; (c)–(d) show the largest incoming observation share within each five-round block. Curves show population-seed means with SE shading. Conditional concentration is undefined, rather than zero, in blocks without observations. Sparse observation can produce high concentration even under random choice, so the realized uniform baseline is more informative than a concentration threshold alone.
Figure 60: Directed observation graphs for LLM-controlled social learning. Panels (a)–(b) show rounds 1–10; (c)–(d) show rounds 11–49. Each annotated cell is the mean percentage of all observations assigned to that observer–peer pair, with identical color scales across panels. Cells are descriptive edge weights, not independent samples. Late GLM graphs summarize substantially fewer observations; the observation-rate comparison is shown in Figure 44 .
Figure 61: Decision costs when observing familiar and new sources. Panels (a)–(b) show conditional means with seed SE, separately for learning selection before observation and deployment selection afterward; solid lines denote familiar sources and dashed lines new ones. Panels (c)–(d) show familiar-minus-new differences matched within observer and five-round block, averaged within seed. Black points are population estimates and bars show their mean with SE. No unobserved cell is imputed as zero.
Figure 62: Paired effects of copying schedules, with means and seed SE. Columns show GPT-OSS and GLM; rows show final accuracy, trajectory-average accuracy, and learning expense excluding hidden tests. Early and distributed both assign four observations; early access only permits optional observation in rounds 1–10. Solo and unrestricted controls are reused, not new independent replications.
Figure 63: Held-out whole-answer accuracy under randomized copying schedules. Panels (a)–(b) compare four forced observations in rounds 1–10 with four observations distributed across learning; solo is shown as a reference. Panels (c)–(d) compare optional access only in rounds 1–10 with unrestricted social access and solo learning. Lines show six-seed means and shaded SE. All comparisons use the common checkpoints at rounds 0, 10, 25, and 49; test feedback never enters learning.
Figure 64: How the timing interventions change learning behavior. Panels (a)–(b) show observations per agent; panels (c)–(d) show the share of observations that copy a non-initial skill version; panels (e)–(f) show successful private revisions. Each statistic is first computed within a five-agent population, then averaged over six seeds with SE. Forced conditions contain exactly four observations per agent by design. “Early access” is optional rounds 1–10; “Unrestricted” permits observation throughout learning.
Comparison
Outcome
Effect
Significance
Hierarchical vs. solo UCB
Cumulative reward
+25.68±2.87
p<0.001
Qwen social vs. solo
Cumulative reward
+2.8±20.3
n.s.
GPT-OSS social vs. solo
Cumulative reward
−61.6±39.2
n.s.
Ministral social vs. solo
Cumulative reward
+5.0±27.1
n.s.
Qwen social vs. solo
Efficiency relative to solo
−29.6±13.4%
n.s.
GPT-OSS social vs. solo
Efficiency relative to solo
−51.9±7.8%
p<0.01
Appendix
Table 4: Key paired results in the finite-option environment. Effects are treatment minus comparison, reported as mean ± SE across population seeds. Reward is cumulative through round 100. Efficiency is the within-seed percentage change in E(π) (Equation 3 ) relative to solo, estimated using realized rewards and charged completion tokens. Intervention effects compare treated and default social populations. Significance uses two-sided paired t -tests with Bonferroni correction across all ten comparisons; n.s. denotes adjusted p≥0.05 .
Figure 65: Finite-bandit mean reward against completion tokens in the primary stationary, neutral-prompt, expiring-budget comparison. Panels (a)–(c) show Qwen, GPT-OSS, and Ministral. Each point is one population seed. These are realized outcomes, not effects of assigning a token budget.
LLM-powered AI assistants acting on behalf of users can produce poor collective outcomes at scale. We introduce a framework for evaluating their emergent behaviour in social dilemmas, applied to three iterated games (Public Goods, Collective Risk, Common Pool Resource). We prompt each model to produce a natural-language strategy, then have the same model translate it into code. This aims to isolate strategic reasoning from input-parsing, enables pre-deployment inspection, and scales to populations of hundreds of agents. We propose three analyses: behavioural fingerprinting via exhaustive evaluation over opponent histories; self-play robustness across mixtures of a model's strategies with either a Selfish or Collective disposition; and cultural evolution under payoff-biased imitation. Applied to three state-of-the-art LLMs, we find substantial cross-model differences in self-play welfare, and that cultural evolution converges to low-welfare, Selfish-dominant equilibria in larger groups.
Richard Willis, Jianing Zhao, Yali Du +1
King’s College London London, United Kingdom · King’s College London · Turing Institute +1
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{$4.03} in search cost, it produces a harness that solves \textbf{82.0}% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
Jungwoo Yang, Injin Kong, Yohan Jo
Graduate School of Data Science, Seoul National University
LLM agents increasingly operate in open-ended environments spanning hundreds of sequential episodes, yet they remain largely stateless: each task is solved from scratch without converting past experience into better future behavior. The central obstacle is not \emph{what} to remember but \emph{how to use} what has been remembered, including which retrieval policy to apply, how to interpret prior outcomes, and when the current strategy itself must change. We introduce \emph{Agent Evolving Learning} (\ael{}), a two-timescale framework that addresses this obstacle. At the fast timescale, a Thompson Sampling bandit learns which memory retrieval policy to apply at each episode; at the slow timescale, LLM-driven reflection diagnoses failure patterns and injects causal insights into the agent's decision prompt, giving it an interpretive frame for the evidence it retrieves. On a sequential portfolio benchmark (10 sector-diverse tickers, 208 episodes, 5 random seeds), \ael{} achieves a Sharpe ratio of 2.13±0.47, outperforming five published self-improving methods and all non-LLM baselines while maintaining the lowest variance among all LLM-based approaches. A nine-variant ablation reveals a ``less is more'' pattern: memory and reflection together produce a 58% cumulative improvement over the stateless baseline, yet every additional mechanism we test (planner evolution, per-tool selection, cold-start initialization, skill extraction, and three credit assignment methods) \emph{degrades} performance. This demonstrates that the bottleneck in agent self-improvement is \emph{self-diagnosing how to use} experience rather than adding architectural complexity. Code and data: https://github.com/WujiangXu/AEL.