Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Authors: Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, +2 more
Organizations: RPI · Microsoft Research · UC San Diego · KAIST · Mila
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
Figures & tables
Method
Trains actions
Online training
Trains reflections
Teacher/ demo
Verdict required
Reflection in context
Scalar reward
RLTF
✓
✓
✓
✓
✓
✓
✓
ERL
✓
✓
✓
×
✓
✓
✓
OPCD (experience)
✓
✓
×
∼
×
✓
×
STaR
✓
✓
×
×
✓
×
×
Early Experience SR
✓
×
✓
✓
×
✓
×
Reflexion / Self-Refine
×
—
×
×
×
✓
×
Table 1: ✓ yes; × no; ∼ partial: —: no weight training. Actions : directly trains task-solving outputs, including solutions within critiques; Online : fresh model generations during training; otherwise offline. Reflections : reflection/critique tokens are training targets. Teacher/demo : requires external teachers/demonstrations. Verdict : requires verifier correctness judgments, not just environment observations. In context : reflections used in context to affect actions. Reward : scalar reward required for training.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Input component
Retained characters
Task description
4,000
Thought, action, or observation within each step
600 each
Concatenated trajectory summary
24,000
Final patch
2,000
Test output
3,000
Appendix
Table 2: Character budgets for constructing the retrospection input. Budgets exclude inserted truncation markers and structural headings.
Hyperparameter
Value
Initial model
Qwen3.5-4B
Training tasks
767
Optimizer updates
Specified per experiment
Source attempts per update
64
Solver attempts per task selection
1
Retrospections per source
4
Appendix
Table 3: Reference online training and generation hyperparameters for ROFT. Experiment-specific settings take precedence. Batch sizes are nominal; targets rejected during preprocessing incur no loss.
Mean assistant tokens
Mean assistant turns
Rollouts
Baseline
Faster
Change
Baseline
Faster
Change
All
12,674
11,186
-11.7%
47.05
40.88
-13.1%
Solved
13,064
11,454
-12.3%
48.03
41.83
-12.9%
Unsolved
12,050
10,808
-10.3%
45.48
39.56
-13.0%
Appendix
Table 4: Mean coding-rollout lengths around update 10, pooled over batches 8–12. Percentage changes compare the faster-retrospection variant with the baseline. Each attempt receives equal weight.
Model
Solved
Rate
Δ vs. Uniform
SE (pp)
Uniform
324/640
50.63%
–
–
Thinking-down
318/640
49.69%
-0.94
2.28
Thinking-up
330/640
51.56%
+0.94
2.23
Evidence-up
347/640
54.22%
+3.59
2.26
Correction-up
348/640
54.38%
+3.75
2.12
Lesson-up
344/640
53.75%
+3.13
2.15
Appendix
Table 5: Repeated-evaluation results on the fixed 64 training problems. Rates average ten attempts per task. Changes relative to Uniform and their standard errors (SE) are in percentage points. SEs resample paired repetitions independently within each of the 64 fixed tasks.
LLM-based agents trained with reinforcement learning optimize step-wise action prediction but lack metacognitive awareness of task progress, inducing a gap that hinders long-horizon scaling. A pilot study reveals that online progress prompting hurts performance while retrospective demonstrations help, yet this capability cannot emerge from outcome-reward training alone. We present RePro, Retrospective Progress-Aware Training, a framework that trains agents to self-generate progress signals via a forward-then-reflect rollout paradigm: the agent executes actions online, then retrospectively reassesses its step-wise progress given the completed trajectory and known outcome. RePro initializes with a Retrospection Warmup that teaches reflection format from minimal external demonstrations, then further trains through RePro-PO with a composite reward that produces self-generated signals without continuous external supervision. Experiments on WebShop, ALFWorld, and Sokoban show that RePro enhances the Qwen family's performance, with up to 12% absolute success rate gains.
Xinbei Ma, Congmin Zheng, Jiyang Qiu +10
1Shanghai Jiao Tong University · 2OPPO Research Institute
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficial imitation? We study LMs trained to explain which features of their inputs influenced their behavior, using models' counterfactual behavior on modified inputs as supervision. Surprisingly, we find that LMs trained on fixed counterfactual explanations derived from earlier checkpoints of themselves, or even from behaviorally similar models in different families, frequently produce explanations more faithful to their own current behaviors than to those of their training targets. This "introspective" coupling between LM explanations and behaviors occurs when training explanations remain sufficiently correlated with current behaviors over the course of training, even as behaviors themselves shift. We also show that introspective coupling tracks behavior shifts: when explanation training is provided concurrently with other post-training objectives, explanations track those shifts without requiring updated supervision. This phenomenon appears in multiple tasks, including sycophancy and refusal, and is robust to label noise. Overall, our results show that even fixed datasets of counterfactual explanations can provide scalable and generalizable post-training signal for introspection.
We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing work focuses on evaluating faithfulness or using inference-time prompting frameworks to improve an LLM's self-explanation's tractability, these approaches do not provide a mechanism to directly optimize a model's parameters to generate faithful self-explanations. We bridge this gap by modifying existing faithfulness metrics into an RL training objective. We investigate (1) if models can be trained to accurately detect factors that affect their decisions, and (2) whether RL can directly optimize for the disclosure of these factors thereby improving LLM self-explanations' faithfulness. We experiment with two intervention types: random-word insertions and user-bias insertions, using a per-sample reward derived from the Phi-CCT correlation metric. RL fine-tuned Llama3.1-8B and Qwen3-8B show substantial improvements on the Phi-CCT faithfulness metric, with in-distribution scores rising from near-zero to as high as 0.664, and out-of-distribution scores reaching up to 0.691 on held-out tasks such as StrategyQA. Cross-intervention generalization is weaker but more interesting: a priori we would not expect a model trained only on random word insertions to generalize to user-bias phrases, yet Llama3.1-8B shows non-zero transfer in this direction. The reverse direction and Qwen3-8B do not replicate this, indicating model-dependent and setup-dependent effects we cannot yet explain. Lastly we analyze model behavior to rule out reward gaming behaviors that often plague RL training. Ultimately, we show that models can be trained to implicitly identify influential factors and disclose them, offering a scalable path toward reducing unfaithful reasoning in LLMs.
Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel +1
Centre for AI, Department of Computer Science University College London · Imperial College London