Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Authors: Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, +2 more
Organizations: RPI · Microsoft Research · UC San Diego · KAIST · Mila
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
Figures & tables
Method
Trains actions
Online training
Trains reflections
Teacher/ demo
Verdict required
Reflection in context
Scalar reward
RLTF
✓
✓
✓
✓
✓
✓
✓
ERL
✓
✓
✓
×
✓
✓
✓
OPCD (experience)
✓
✓
×
∼
×
✓
×
STaR
✓
✓
×
×
✓
×
×
Early Experience SR
✓
×
✓
✓
×
✓
×
Reflexion / Self-Refine
×
—
×
×
×
✓
×
Table 1: ✓ yes; × no; ∼ partial: —: no weight training. Actions : directly trains task-solving outputs, including solutions within critiques; Online : fresh model generations during training; otherwise offline. Reflections : reflection/critique tokens are training targets. Teacher/demo : requires external teachers/demonstrations. Verdict : requires verifier correctness judgments, not just environment observations. In context : reflections used in context to affect actions. Reward : scalar reward required for training.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Input component
Retained characters
Task description
4,000
Thought, action, or observation within each step
600 each
Concatenated trajectory summary
24,000
Final patch
2,000
Test output
3,000
Appendix
Table 2: Character budgets for constructing the retrospection input. Budgets exclude inserted truncation markers and structural headings.
Hyperparameter
Value
Initial model
Qwen3.5-4B
Training tasks
767
Optimizer updates
Specified per experiment
Source attempts per update
64
Solver attempts per task selection
1
Retrospections per source
4
Appendix
Table 3: Reference online training and generation hyperparameters for ROFT. Experiment-specific settings take precedence. Batch sizes are nominal; targets rejected during preprocessing incur no loss.
Mean assistant tokens
Mean assistant turns
Rollouts
Baseline
Faster
Change
Baseline
Faster
Change
All
12,674
11,186
-11.7%
47.05
40.88
-13.1%
Solved
13,064
11,454
-12.3%
48.03
41.83
-12.9%
Unsolved
12,050
10,808
-10.3%
45.48
39.56
-13.0%
Appendix
Table 4: Mean coding-rollout lengths around update 10, pooled over batches 8–12. Percentage changes compare the faster-retrospection variant with the baseline. Each attempt receives equal weight.
Model
Solved
Rate
Δ vs. Uniform
SE (pp)
Uniform
324/640
50.63%
–
–
Thinking-down
318/640
49.69%
-0.94
2.28
Thinking-up
330/640
51.56%
+0.94
2.23
Evidence-up
347/640
54.22%
+3.59
2.26
Correction-up
348/640
54.38%
+3.75
2.12
Lesson-up
344/640
53.75%
+3.13
2.15
Appendix
Table 5: Repeated-evaluation results on the fixed 64 training problems. Rates average ten attempts per task. Changes relative to Uniform and their standard errors (SE) are in percentage points. SEs resample paired repetitions independently within each of the 64 fixed tasks.