Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or strategically revises the plan when it fails. We separate these two capabilities: execution-level capability, iterating within an established training strategy, and strategy-level capability, revising that strategy as experimental evidence accumulates. Analyzing 1,338 post-training trajectories of frontier agents, we find that agents reliably execute post-training but lock into a default strategy, which follows the agent rather than the task, and only 2.1% of transitions between adjacent training runs ever change strategy. We then test whether the agent lacks experience, reasoning, or the decision to switch. (1) Experience improves execution but not the strategy. (2) Additional reasoning compute yields front-loaded gains on easier tasks but refines, rather than revises, the committed strategy. (3) Human review before training changes which strategy the agent locks into, not whether it locks in, whereas a single mid-run instruction outperforms the agent's own continuation by up to 17.44 points under the same budget. In conclusion, what the agent lacks is the decision to reopen a committed strategy and try another one. Realizing RSI therefore calls for interaction protocols and training signals that make strategy revision an explicit, rewarded decision.
Figures & tables
Figure 1: Top: agents iterate reliably at the execution level yet rarely revise the strategy. Bottom: experience and reasoning improve execution, whereas human guidance redirects the strategy.
Figure 2: Benchmark scores of the submitted checkpoints averaged over four base models: agents reliably execute post-training pipelines and improve over the base model across seven benchmarks.
Agent (#Traj.)
Default Strategy ( κa )
Switch Rate ( ρa )
Claude Code (575)
Full SFT
71.9%
54/1,203 (4.5%)
Codex CLI (369)
PEFT
89.5%
15/943 (1.6%)
OpenCode (394)
Full SFT
66.4%
5/1,411 (0.4%)
Overall (1,338)
76.7%
74/3,557 (2.1%)
Table 1: Default-strategy concentration κa and switch rate ρa per agent (Appendix A ).
Figure 3: Time usage of representative runs. The autonomous baseline spends most of the budget in SFT, while the experience-driven agent alternates training with evaluation and reflection.
Benchmark
Aggregate
Setting
GSM8K
HumanEval
AIME 2025
Avg.
Gap Closed
Base model
10.84
5.48
0.00
5.44
0.0%
Official instruct model
88.70
66.46
33.33
62.83
100.0%
Opus 4.6 (Claude Code)
64.70 ± 9.6 (+53.9)
32.00 ± 10.4 (+26.5)
3.33 ± 0.0 (+3.3)
33.34
48.6%
GLM-5.2 (Claude Code)
49.51 ± 7.2 (+38.7)
44.51 ± 9.8 (+39.0)
3.33 ± 0.0 (+3.3)
32.45
47.1%
GPT-5.2 (Codex CLI)
43.44 ± 4.1 (+32.6)
13.41 ± 8.7 (+7.9)
0.00 ± 0.0 (0.0)
18.95
23.5%
Table 2: Benchmark scores (%) under the controlled setting with Qwen3-1.7B-Base as the base model, reported as mean ± one standard deviation over three independent runs.
Figure 4: Across evaluation cycles, the main agent implements every execution-level suggestion ( 22/22 ) but declines every suggestion that departs from its current strategy ( 0/21 ).
Figure 5: Average benchmark score against the main agent’s cumulative token consumption ( ∘ GSM8K, □ HumanEval, △ AIME 2025)
Figure 6: (a) Guidance at the initial decision: the AIME 2025 run peaks early; later iterations (shaded) never recover the peak. (b) Guidance at a mid-run decision: a single instruction to switch strategy outperforms the agent’s own continuation under the same remaining budget.
Figure 7: (a) RSI assumes a loop that closes globally in which the evaluation feeds back into the strategy. (b) The observed loops, however, only close at the execution level (repair & retry).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Algorithm
Executed evidence
Full-parameter SFT
Likelihood training that updates all model parameters.
Parameter-efficient fine-tuning
Likelihood training through Low-Rank Adaptation (LoRA), quantized LoRA (QLoRA), or another adapter.
Reinforcement learning
Sampled model outputs optimized with a reward signal, including Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO).
Preference optimization
Preferred/rejected response training, including Direct Preference Optimization (DPO)-style objectives.
Distillation
Training on supervision or outputs produced by a teacher model.
Appendix
Table 3: Training-algorithm labels and identification criteria.
Agent framework
Benchmark
Final metric ( n )
Initial strategy
Strategy changes
Claude Code
AIME 2025
0.049(33)
Full SFT: 19/31 ( 61.3% )
14/132 ( 10.6% )
ArenaHardWriting
0.135(34)
Full SFT: 27/30 ( 90.0% )
19/174 ( 10.9% )
BFCL
0.860(44)
Full SFT: 33/51 ( 64.7% )
0/171 ( 0.0% )
GPQA Main
0.279(36)
Full SFT: 17/27 ( 63.0% )
4/171 ( 2.3% )
GSM8K
0.556(38)
Full SFT: 24/31 ( 77.4% )
8/187 ( 4.3% )
HealthBench
0.281(36)
Full SFT: 22/28 ( 78.6% )
1/179 ( 0.6% )
Appendix
Table 4: Cross-framework strategy comparison per benchmark, pooled over the four base models. Final metrics are means over trained trajectories with a valid final score.
Changed dimension
Pairs
Rate
Training algorithm
35
0.98%
Data source
38
1.07%
Training stages
1
0.03%
Overall
74
2.08%
Appendix
Table 5: Decomposition of strategy changes among the 3,557 recognized adjacent experiment pairs.
Table 6: Trace-level case studies among the 16 trajectories with a training-algorithm change. The reported values come from the trajectory logs and retain each log’s original comparator and sample size.
Figure 8: Training dynamics under the experience-driven framework (best-performing run). Later experiments plateau or regress, while subsequent actions remain within the committed strategy.
Benchmark
Eval cycles
Execution sugg. (adopted)
Strategy sugg. (adopted)
GSM8K
11
6/6
0/7
HumanEval
14
11/11
0/11
AIME 2025
5
5/5
0/3
Total
30
22/22
0/21
Appendix
Table 7: Evaluator suggestions and their adoption (adopted/total). Counts are theme occurrences per evaluation cycle and match Figure 4 .
Figure 9: An annotated experience-driven trajectory on AIME 2025. The agent consults skills, maintains the experiment journal, and receives evaluator diagnoses throughout six training versions. All revisions remain at the execution level.
GSM8K
HumanEval
AIME 2025
Strategy
Score
Strategy
Score
Strategy
Score
Recorded continuation
SFT
47.76
SFT
52.0
SFT
0/30
Reconsider fork
SFT
48.0
SFT
31.7
SFT, LoRA
1/30
Guided branch
GRPO
65.20
RFT, GRPO
62.8
GRPO
2/30
Appendix
Table 8: Strategies and final scores after the branch point in the agent’s own recorded continuation, the fork with an instruction to reconsider (reconsider fork), and the guided branch. The recorded continuation and the reconsider fork both stay with SFT, whereas each guided branch moves to RL with GRPO and reaches the highest final score. Scores are accuracy (%) on GSM8K and HumanEval and solved problems out of 30 under pass@8 on AIME 2025.
Model
Benchmark
Evals
Diagnostic result
Trajectory diagnosis
Qwen3-1.7B
GSM8K
20
Best/final 0.587 @ 150
LoRA and full-SFT continuations both benefit from prompt-contract, EOS, and answer-only repairs.
Qwen3-1.7B
HumanEval
14
Best 0.273 @ 150 / 0.262 @ 164
MBPP, CodeSearchNet, and template repairs help, but 20–50-sample evaluations overstate the full-benchmark quality.
Qwen3-1.7B
AIME 2025
22
Best 1/30; final repeatedly 0/30
Many math-data and LoRA/SFT variants produce no stable gain; 1/30 is within evaluation variance.
Qwen3-4B
GSM8K
19
Base 0.507 @ 150; final 0.467 @ 150
After 13 successful training runs and 12 merges, the selected model remains worse than the base model.
Qwen3-4B
HumanEval
18
Base 0.420 @ 150; final 0.687 @ 150
MBPP-only data, a no- <think> template, checkpoint selection, and sampling configuration recover from several zero-scoring early runs.
Qwen3-4B
AIME 2025
11
Best 1/30; final 0/30
Seven successful QLoRA runs fail to improve over the noisy base result; intermediate outputs show tokenizer and generation corruption.
Appendix
Table 9: Trajectory-level summary of GPT-5.2 (Codex CLI) behavior at both model scales. Scores are in-run diagnostic pass@1 accuracies; “@ n ” gives the number of evaluated problems, and “Final” denotes the checkpoint exported by the agent.
Model
Benchmark
Evals
Diagnostic result
Trajectory diagnosis
Qwen3-1.7B
GSM8K
4
Base 0.200 @ 50; final 0.493 @ 150
Full SFT learns the answer contract; data cleaning and EOS repair further improve the result despite repeated out-of-memory failures.
Qwen3-1.7B
HumanEval
7
Base 0.116 @ 164; final 0.433 @ 164
Function-completion data and completion-only loss produce a large gain; a better run2 result is evaluated but not exported.
Qwen3-1.7B
AIME 2025
6
Base 0/5; final 0/30
Two full-SFT rounds, EOS repair, and rejection-filtered training complete successfully but yield no solved problem.
Qwen3-4B
GSM8K
24
Best 0.680 @ 200; final 0.654 @ 700
Three rejection-sampling rounds and four SFT stages improve the model; intermediate-checkpoint selection is consistently important.
Qwen3-4B
HumanEval
5
Base 0.300 @ 10; final 0.533 @ 150
Three SFT stages add code data and progressively lower the learning rate, improving the 150-sample result from 0.400 to 0.533.
Qwen3-4B
AIME 2025
5
Base 0/6; final 7/30
Correcting an unintended 2,048-token generation cap raises both trained runs to 7/30; run1 is retained as the final model.
Appendix
Table 10: Trajectory-level summary of GLM-5.2 (Claude Code) behavior at both model scales. Scores are in-run diagnostic pass@1 accuracies; notation as in Table 9 .
Benchmark
Record
Base
Reported final result
Trajectory diagnosis
GSM8K
A
0.1243 (1319)
0.3490±0.0119 (1319, three repeats)
Format-aligned data, assistant-only loss masking, and stop-token repair produce the clearest gain; continuation and checkpoint selection remain protocol-sensitive.
GSM8K
B
0.160 (50)
0.020 (full evaluation)
Training loss and token accuracy improve, but the full-evaluation score remains far below the 50-problem base estimate; the differing evaluation sizes limit the comparison.
HumanEval
C
0.060 (50)
0.2369±0.0302 (164, seven repeats)
The base and final scores use different sample sizes; the run terminates after an API failure, so the result is useful mainly as evidence of evaluation variance.
HumanEval
D
0.060 (150)
0.5427 (164, official path)
Several continuations and data-mixture changes produce intermediate results from approximately 0.400 to 0.555 ; the largest gain follows a change in the proportion of verified self-distillation data.
AIME 2025
E
0.0333 ( 1/30 )
0 ( 0/30 , repeated)
No fine-tuned checkpoint reproduces or exceeds the single solved problem in the base evaluation; long-output and truncation failures dominate the final trajectory.
AIME 2025
F
0 ( 0/30 )
0 official; 0.133 self-evaluation
The reported positive result comes from a non-official evaluation path and is not directly comparable with the official final score.
Appendix
Table 11: Trajectory-level summary of the Fable 5 (Claude Code) runs. Scores are reported with the evaluation protocol and sample size used in the supplementary logs.
RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training dynamics. This distinction highlights a potential flaw in fixed training schedules: by forcing all parameters along rigid paths, they fail to capture the dynamic exploration-exploitation tradeoffs that regularization must track. We uncover this through LLMZero, an agentic system that optimizes training trajectories via tree search by diagnosing pathologies at each checkpoint and proposing coordinated multi-parameter transitions. Across four diverse GRPO tasks, LLMZero discovers strategies that improve over the base model by 9% to 140% and over grid search by 6% to 15% (relative), consistently outperforming random search and a skill-based agent under a matched compute budget. The capacity--regularization asymmetry is consistent across all four tasks, offering a candidate design heuristic for multi-stage training.
Haoyang Fang, Wei Zhu, Boran Han +11
†LLMZero Project Core Team. · ∗Work done at Amazon.
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov decision process, which we term progress advantage -- log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This formulation makes the resulting signal annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, it consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.
Changdae Oh, Wendi Li, Seongheon Park +3
University of Wisconsin–Madison · Argonne National Laboratory
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to the evolving task state, but existing LLM TTT methods largely adapt once to a fixed input. We study continuous TTT in multi-turn agent episodes, where each update changes the policy that generates later training text. This creates a self-training loop that helps when new trajectory information appears, but can amplify drift when the agent gets stuck and repeatedly trains on similar text. We find that update-text repetition distinguishes these regimes and introduce Agentic Test-Time Training (aTTT), a token-level reweighting method that downweights the loss on tokens appearing in repeated n-grams from prior updates while leaving novel tokens fully weighted. To run such updates inside live episodes, we build a concurrent serving system using vLLM's runtime LoRA API, limiting overhead to 1.9× the no-TTT cost. aTTT improves success by up to 5.0 points on ALFWorld and 4.9 points on SWE-bench Lite. The gains concentrate where models already have task competence but drift over long trajectories, suggesting that aTTT mainly preserves existing competence rather than teaching new abilities.