What is Missing from AI Post-Training AI: An Empirical Analysis
Organizations: Tsinghua University · University of Electronic Science and Technology of China · Renmin University of China
Abstract
Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or strategically revises the plan when it fails. We separate these two capabilities: execution-level capability, iterating within an established training strategy, and strategy-level capability, revising that strategy as experimental evidence accumulates. Analyzing 1,338 post-training trajectories of frontier agents, we find that agents reliably execute post-training but lock into a default strategy, which follows the agent rather than the task, and only 2.1% of transitions between adjacent training runs ever change strategy. We then test whether the agent lacks experience, reasoning, or the decision to switch. (1) Experience improves execution but not the strategy. (2) Additional reasoning compute yields front-loaded gains on easier tasks but refines, rather than revises, the committed strategy. (3) Human review before training changes which strategy the agent locks into, not whether it locks in, whereas a single mid-run instruction outperforms the agent's own continuation by up to 17.44 points under the same budget. In conclusion, what the agent lacks is the decision to reopen a committed strategy and try another one. Realizing RSI therefore calls for interaction protocols and training signals that make strategy revision an explicit, rewarded decision.
Figures & tables
| Agent (#Traj.) | Default Strategy ( ) | Switch Rate ( ) | |
|---|---|---|---|
| Claude Code (575) | Full SFT | 71.9% | 54/1,203 (4.5%) |
| Codex CLI (369) | PEFT | 89.5% | 15/943 (1.6%) |
| OpenCode (394) | Full SFT | 66.4% | 5/1,411 (0.4%) |
| Overall (1,338) | 76.7% | 74/3,557 (2.1%) | |
| Benchmark | Aggregate | ||||
| Setting | GSM8K | HumanEval | AIME 2025 | Avg. | Gap Closed |
| Base model | 10.84 | 5.48 | 0.00 | 5.44 | 0.0% |
| Official instruct model | 88.70 | 66.46 | 33.33 | 62.83 | 100.0% |
| Opus 4.6 (Claude Code) | 64.70 9.6 (+53.9) | 32.00 10.4 (+26.5) | 3.33 0.0 (+3.3) | 33.34 | 48.6% |
| GLM-5.2 (Claude Code) | 49.51 7.2 (+38.7) | 44.51 9.8 (+39.0) | 3.33 0.0 (+3.3) | 32.45 | 47.1% |
| GPT-5.2 (Codex CLI) | 43.44 4.1 (+32.6) | 13.41 8.7 (+7.9) | 0.00 0.0 (0.0) | 18.95 | 23.5% |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Algorithm | Executed evidence |
|---|---|
| Full-parameter SFT | Likelihood training that updates all model parameters. |
| Parameter-efficient fine-tuning | Likelihood training through Low-Rank Adaptation (LoRA), quantized LoRA (QLoRA), or another adapter. |
| Reinforcement learning | Sampled model outputs optimized with a reward signal, including Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO). |
| Preference optimization | Preferred/rejected response training, including Direct Preference Optimization (DPO)-style objectives. |
| Distillation | Training on supervision or outputs produced by a teacher model. |
| Agent framework | Benchmark | Final metric ( ) | Initial strategy | Strategy changes |
|---|---|---|---|---|
| Claude Code | AIME 2025 | Full SFT: ( ) | ( ) | |
| ArenaHardWriting | Full SFT: ( ) | ( ) | ||
| BFCL | Full SFT: ( ) | ( ) | ||
| GPQA Main | Full SFT: ( ) | ( ) | ||
| GSM8K | Full SFT: ( ) | ( ) | ||
| HealthBench | Full SFT: ( ) | ( ) |
| Changed dimension | Pairs | Rate |
|---|---|---|
| Training algorithm | 35 | 0.98% |
| Data source | 38 | 1.07% |
| Training stages | 1 | 0.03% |
| Overall | 74 | 2.08% |
| Case | Setting | Change | Observation | Interpretation |
|---|---|---|---|---|
| 3fd3ea0b | ArenaHardWriting, SmolLM3-3B, Codex CLI | SFT DPO | Held-out 256-pair proxy: | Positive; pp |
| c92715d6 | ArenaHardWriting, Gemma-3-4B, Claude Code | SFT DPO | Benchmark win rate: | Positive; approximately |
| 50c64287 | ArenaHardWriting, Qwen3-4B, Codex CLI | SFT DPO | Same 8-prompt local proxy: | Positive; small sample |
| 860ceacd | AIME 2025, Qwen3-1.7B, Claude Code | SFT GRPO | 30-problem evaluation: (1/30) | Positive; high variance |
| 8226f786 | ArenaHardWriting, SmolLM3-3B, Codex CLI | SFT DPO | External proxy: ; internal DPO reward accuracy: | Counterexample; metric mismatch |
| Benchmark | Eval cycles | Execution sugg. (adopted) | Strategy sugg. (adopted) |
|---|---|---|---|
| GSM8K | 11 | 6/6 | 0/7 |
| HumanEval | 14 | 11/11 | 0/11 |
| AIME 2025 | 5 | 5/5 | 0/3 |
| Total | 30 | 22/22 | 0/21 |
| GSM8K | HumanEval | AIME 2025 | ||||
|---|---|---|---|---|---|---|
| Strategy | Score | Strategy | Score | Strategy | Score | |
| Recorded continuation | SFT | 47.76 | SFT | 52.0 | SFT | 0/30 |
| Reconsider fork | SFT | 48.0 | SFT | 31.7 | SFT, LoRA | 1/30 |
| Guided branch | GRPO | 65.20 | RFT, GRPO | 62.8 | GRPO | 2/30 |
| Model | Benchmark | Evals | Diagnostic result | Trajectory diagnosis |
|---|---|---|---|---|
| Qwen3-1.7B | GSM8K | 20 | Best/final 0.587 @ 150 | LoRA and full-SFT continuations both benefit from prompt-contract, EOS, and answer-only repairs. |
| Qwen3-1.7B | HumanEval | 14 | Best 0.273 @ 150 / 0.262 @ 164 | MBPP, CodeSearchNet, and template repairs help, but 20–50-sample evaluations overstate the full-benchmark quality. |
| Qwen3-1.7B | AIME 2025 | 22 | Best 1/30; final repeatedly 0/30 | Many math-data and LoRA/SFT variants produce no stable gain; 1/30 is within evaluation variance. |
| Qwen3-4B | GSM8K | 19 | Base 0.507 @ 150; final 0.467 @ 150 | After 13 successful training runs and 12 merges, the selected model remains worse than the base model. |
| Qwen3-4B | HumanEval | 18 | Base 0.420 @ 150; final 0.687 @ 150 | MBPP-only data, a no- <think> template, checkpoint selection, and sampling configuration recover from several zero-scoring early runs. |
| Qwen3-4B | AIME 2025 | 11 | Best 1/30; final 0/30 | Seven successful QLoRA runs fail to improve over the noisy base result; intermediate outputs show tokenizer and generation corruption. |
| Model | Benchmark | Evals | Diagnostic result | Trajectory diagnosis |
|---|---|---|---|---|
| Qwen3-1.7B | GSM8K | 4 | Base 0.200 @ 50; final 0.493 @ 150 | Full SFT learns the answer contract; data cleaning and EOS repair further improve the result despite repeated out-of-memory failures. |
| Qwen3-1.7B | HumanEval | 7 | Base 0.116 @ 164; final 0.433 @ 164 | Function-completion data and completion-only loss produce a large gain; a better run2 result is evaluated but not exported. |
| Qwen3-1.7B | AIME 2025 | 6 | Base 0/5; final 0/30 | Two full-SFT rounds, EOS repair, and rejection-filtered training complete successfully but yield no solved problem. |
| Qwen3-4B | GSM8K | 24 | Best 0.680 @ 200; final 0.654 @ 700 | Three rejection-sampling rounds and four SFT stages improve the model; intermediate-checkpoint selection is consistently important. |
| Qwen3-4B | HumanEval | 5 | Base 0.300 @ 10; final 0.533 @ 150 | Three SFT stages add code data and progressively lower the learning rate, improving the 150-sample result from 0.400 to 0.533. |
| Qwen3-4B | AIME 2025 | 5 | Base 0/6; final 7/30 | Correcting an unintended 2,048-token generation cap raises both trained runs to 7/30; run1 is retained as the final model. |
| Benchmark | Record | Base | Reported final result | Trajectory diagnosis |
|---|---|---|---|---|
| GSM8K | A | (1319) | (1319, three repeats) | Format-aligned data, assistant-only loss masking, and stop-token repair produce the clearest gain; continuation and checkpoint selection remain protocol-sensitive. |
| GSM8K | B | (50) | (full evaluation) | Training loss and token accuracy improve, but the full-evaluation score remains far below the 50-problem base estimate; the differing evaluation sizes limit the comparison. |
| HumanEval | C | (50) | (164, seven repeats) | The base and final scores use different sample sizes; the run terminates after an API failure, so the result is useful mainly as evidence of evaluation variance. |
| HumanEval | D | (150) | (164, official path) | Several continuations and data-mixture changes produce intermediate results from approximately to ; the largest gain follows a change in the proportion of verified self-distillation data. |
| AIME 2025 | E | ( ) | ( , repeated) | No fine-tuned checkpoint reproduces or exceeds the single solved problem in the base evaluation; long-output and truncation failures dominate the final trajectory. |
| AIME 2025 | F | ( ) | official; self-evaluation | The reported positive result comes from a non-official evaluation path and is not directly comparable with the official final score. |