Agent Plasticity: Measuring Self-Improvement Through Experience
Organizations: UC Berkeley · Meta Superintelligence Labs · University of Washington · Princeton University
Abstract
AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize past experience into reusable artifacts that are inherited by future instances. At each checkpoint, we measure performance on training and held-out environment interactions while accounting for learning cost. We introduce agent plasticity, the efficiency with which an agent converts experience into gains in future held-out performance. Across multiple environments, frontier models exhibit sharply different improvement trajectories despite comparable opportunities to learn. Some achieve substantial and persistent gains, while others remain near or below their initial performance, and gains within the training regime often transfer only partially to out-of-distribution conditions. Endpoint capability and acquisition efficiency also diverge: the agent that ultimately performs best need not be the one that improves most efficiently. Tracing failures through the improvement loop further reveals different candidate bottlenecks. Agents with low plasticity often fail to reuse relevant artifacts, whereas more plastic agents may still fail despite reusing relevant artifacts, pointing to limitations in artifact quality, generalization, or application. Evaluating self-improving agents requires measuring not only what they can do, but how effectively they become better through experience.
Figures & tables
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Research direction | Primary focus | Our emphasis |
|---|---|---|
| Reflection, memory, and skills ( Shinn et al., 2023 ; Wang et al., 2024 ) | Mechanisms for retaining and reusing experience. | Measuring how efficiently different models construct and exploit persistent artifacts. |
| Context and harness optimization ( Zhang et al., 2026c ; Agrawal et al., 2026 ) | Optimizing prompts, context, or agent implementations. | Evaluating a model that acts and improves its own future behavior under a common protocol. |
| Long-context evaluation ( Hsieh et al., 2024 ; Yen et al., 2025 ) | Retrieving, reasoning over, and following information in long contexts. | Measuring durable capability gains from self-constructed artifacts, including executable tools. |
| Sequential learning benchmarks ( Dou et al., 2025 ; He et al., 2026 ) | Measuring improvement and, in some settings, learning efficiency across interactions. | Combining cost-normalized held-out learning curves with artifact-level failure diagnostics. |
| Parametric adaptation ( Zweiger et al., 2025 ) | Persistent improvement through model-weight updates. | Isolating non-parametric adaptation through persistent agent artifacts. |
| Setting | Task and score | Opponent or difficulty | Games per checkpoint | Training feedback available to reflection | Offline annotation for failure analysis |
|---|---|---|---|---|---|
| Chess (Hard) | Standard chess; win/draw/loss scores . | Stockfish 16 at 20k nodes/move. Train: skills 1–8; held-out ID: 3–8; held-out OOD: 12. | 16 Train, 12 ID, 12 OOD; both colors. | After each training game, a separate skill-20 Stockfish evaluates every actor move. Centipawn loss gives ok , inaccuracy , mistake , or blunder . | The analysis normalizes the game, model, and sandbox event streams and carries the recorded Stockfish annotations into failure extraction; no second engine pass is added. |
| Chess (Easy) | The same board, rules, score, move limit, and annotation procedure. | Stockfish 16 at 20k nodes/move. Train and held-out ID: skills 2–5; held-out OOD: skills 8–10. | 16 Train, 16 ID, 12 OOD; both colors. | The same skill-20 counterfactual analysis and centipawn-loss severity labels as Hard. | The same trace normalization and recorded-annotation binding as Hard. |
| Go | Go, Chinese area scoring, positional superko, no suicide, komi 9.5; win/loss scores . | GNU Go levels 5–9 for Train and ID; level 10 for OOD; 75-action cap. | 20 Train, 20 ID, 12 OOD; both colors. | Deterministic descriptive diagnostics: captures, liberties, self-atari, avoidable passing while behind, and area-margin change; no expert severity label. | Frozen KataGo replays each position at 64 visits and refines candidates at 512 visits. Stable chosen-versus-best regret supplies mistake/blunder evidence; raw requests, responses, alternatives, perspective, and uncertain cases are recorded. |
| Hex | Deterministic Hex; connection win/loss scores . | Native alpha–beta levels 10k–1M nodes/move for Train and ID; 10M for OOD. | 40 Train, 40 ID, 32 OOD; both colors. | Rule-based missed immediate wins, preventable immediate losses, and own/opponent connection-distance changes; immediate tactical errors are blunder . | The frozen native engine compares legal actions at a common completed depth, refines likely errors at 1M nodes, and adjudicates threshold-near or high-impact cases at 10M. Exact, proven, stable, uncertain, and not-failure dispositions are preserved. |
| NetHack | NetHack Challenge v0; raw per-game maximum score, summarized by checkpoint mean, median, and maximum. | Lawful dwarf female Valkyrie on deterministic fresh seeds. The Codex and Claude Code agent interfaces; at most 50k primitive actions per game for all models; two-hour active play (details below). | 10 fresh games per lineage. | The reflector receives all ten verified public actor traces, terminal reasons, raw score, depth, experience, and action metrics, actor workspace evidence, and prior accepted inventories. No optimal-action oracle is used. |
| Lineage | Tool code (files / lines) | Tests | Instr. file lines | Tool calls per game (share) | Actions per game | Dug down | Mean score |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 6 / 2,713 | 125 | 419 | 622 (97%) | 3,585 | 5 | 9,015 |
| GPT-5.6 Sol | 1 / 2,089 | 58 | 176 | 218 (59%) | 1,211 | 0 | 417 |
| Claude Opus 4.8 | 5 / 1,341 | 58 | 413 | 10 (2%) | 716 | 2 | 933 |
| Claude Sonnet 5 | 4 / 1,683 | 81 | 467 | 914 (96%) | 955 | 1 | 402 |
| GPT-5.6 Luna | 0 / 0 | 0 | 146 | 0 (0%) | 880 | 0 | 239 |
| Claude Opus 5 | 13 / 8,237 | 505 | 452 | 293 (96%) | 1,704 | 1 | 1,919 |
| Stage | Input | What the prompt asks | Output |
|---|---|---|---|
| Failure identification | Game rules; for each decision the position, legal actions, chosen action, and engine evaluation; recorded failures and uncertain engine candidates. No artifacts. | Describe each recorded failure from direct evidence. Label each uncertain candidate; heuristic or unstable engine evidence alone is not a failure. Add failures only when same-decision evidence shows them, at most one optional failure per decision. | Failure list with names, descriptions, and cited events. |
| Failure verification | Same input plus the draft, marked as an untrusted proposal. | Re-check independently and return the complete corrected list: keep every recorded failure, relabel candidates, keep only directly evidenced additions, and drop duplicates. | Final failure list. |
| Failure–artifact relation | Verified failures and the records of their decisions; the full contents of every artifact available at that checkpoint; recorded artifact reads and invocations. | Judge whether an artifact covers the failure from its contents, not its name. Judge use separately, from evidence before or during the failure: an invocation and its end event for executables, access and application for text. | Per failure: no covering artifact, or covering artifact used or not used. |
| Relation verification | Same input plus the draft. | Check every failure against the complete artifact list, including artifacts the draft omitted. Producing the failing action does not by itself make an artifact covering. | Final coverage and use judgment. |
| Decision-level reuse | Every game-play decision, in batches of up to 20, with its verified failures and relation judgments; artifacts; recorded reads and invocations. | For each decision and artifact, judge relevance to choosing or executing the action, use from same-decision evidence, and outcome. An artifact that covers a failure at the decision is relevant. | Per pair: relevant, used, and successful or failure remained. |
| Reuse verification | Same input plus the draft. | Re-audit and return complete corrected rows. Availability, file listings, names, automatic loading, and evidence from other decisions do not establish use. | Final reuse judgments. |