Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fails to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories themselves? To address this challenge, we introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes trajectory-derived evidence into nested shortcut trees. By consolidating redundant attempts, identifying resolved subtasks, and retaining useful steps alongside outstanding requirements, DENSE transforms noisy execution traces into structured and reusable task-solving feedback. To evaluate whether such feedback helps agents retry the same task, we design REFIT, which measures success-rate changes between the initial attempt and feedback-guided retries. Among feedback methods without external outcome supervision, DENSE achieves the highest strict pass rate across four agent models on Terminal-Bench 2.1, improving over initial attempts by 7.12-15.64 percentage points with 19.0-43.6% fewer agent tokens on retries. In addition, on hard tasks DENSE consistently outperforms self-reflection in cumulative pass rate across multiple feedback iterations on all four models, demonstrating its strong potential for continual agent self-improvement.
Figures & tables
Figure 1: Execution traces accumulate, while external supervision is costly and bounded. Can agents reuse experience from these traces without external rewards or correctness labels?
Figure 2: DENSE has two stages. Stage 1 identifies a semantic boundary for a consecutive prefix of queue Qk , diagnoses its actions, summarizes a parent node, and places it in Qk+1 to build the hierarchy. Stage 2 has three steps: A compresses sibling attempts into shortcuts retaining failed attempts, working paths, and evidence; B reconciles inherited issues using later recovery evidence; C summarizes completed branches and expands unfinished requirements with their evidence. The constructed CSV report example has a completed JSON summary and an unfinished PNG chart.
Figure 3: REFIT uses a shared run0 to construct feedback for each method’s separate run1, starting from initial environment state s0 and a fresh model context. Performance changes are measured against this common initial execution. Feedback generation is separated from post-hoc outcome evaluation.
Method
Feedback and construction
Role
Advisor
Advice based only on the task description, without the execution trajectory.
Task-prior control
Self-reflection
Raw trajectory for the next-round agent to summarize past experience and reflect on it.
Self-reflection baseline
All-at-once
One call analyzes the full trajectory and returns action-level assessments and rationales.
Global baseline
Step-by-step
One call per action assesses history through that action, excluding its result and later events.
Sequential baseline
DENSE (ours)
Nested subtasks compress redundant attempts while retaining key evidence and unresolved issues.
Proposed method
Verifier Feedback
The initial trajectory, its final reward, and verifier test output.
Privileged ref.
Table 1: Feedback supplied to the subsequent execution. Only Verifier Feedback (VF) uses post-hoc reward or verifier output. Advisor denotes the instruction-only advisor.
Method
MiniMax-M2.7
DeepSeek V4 Pro
GPT-5.5
Kimi K2.6
Baseline
44.57
51.44
68.54
39.33
Advisor
41.20 / -3.37
53.91 / +2.47
69.29 / +0.75
43.82 / +4.49
Self-reflection
38.58 / -5.99
63.79 / +12.35
72.28 / +3.75
46.07 / +6.74
All-at-once
44.57 / +0.00
62.14 / +10.70
73.03 / +4.49
47.94 / +8.61
Step-by-step
42.70 / -1.87
65.43 / +13.99
73.03 / +4.49
46.82 / +7.49
DENSE (ours)
52.43 / +7.87
67.08 / +15.64
75.66 / +7.12
50.94 / +11.61
Table 2: Strict pass rate (%) / change from baseline (pp), computed before rounding. Baseline aggregates the initial executions shared by all methods. Verifier Feedback (VF) † has privileged outcome information. Bold / underline mark the highest / second-highest rates per recipient, including Baseline and excluding VF; ties share a mark.
Figure 4: Strict pass rate versus observed recipient tokens per result for non-privileged conditions. Means use three repetitions and equal task weights: 89 tasks, or 81 for DeepSeek. Baseline is run0; other points count run1 execution only, excluding feedback generation and source run0. Axes vary by model; upper left is better. Token accounting follows Appendix A .
Method
MiniMax-M2.7
DeepSeek V4 Pro
GPT-5.5
Kimi K2.6
Baseline
1,402.2
2,075.0
319.1
545.3
Advisor
1,729.6 / +23.4%
2,224.4 / +7.2%
302.6 / -5.2%
602.2 / +10.4%
Self-reflection
1,198.3 / -14.5%
1,709.0 / -17.6%
230.7 / -27.7%
559.7 / +2.6%
All-at-once
1,353.9 / -3.4%
1,710.5 / -17.6%
211.4 / -33.7%
546.9 / +0.3%
Step-by-step
1,609.8 / +14.8%
1,500.9 / -27.7%
294.8 / -7.6%
587.8 / +7.8%
DENSE (ours)
1,136.0 / -19.0%
1,660.4 / -20.0%
179.9 / -43.6%
416.1 / -23.7%
Table 3: Observed recipient token use: mean execution tokens (thousands) / percentage change from the baseline row. Negative changes indicate reduced use. † Privileged reference. Bold / underline mark the lowest / second-lowest token use per recipient, excluding VF; ties share a mark.
Method
Token cost (K) ↓
Strict pass (%) ↑
run0
319.1
68.54
w/o S&R
228.6
72.28
w/o hierarchy
199.2
72.66
DENSE
179.9
75.66
Table 4: GPT-5.5 ablations: mean execution tokens (K) and strict pass (%).
Figure 5: Cumulative strict success on Terminal-Bench 2.1 hard tasks through three feedback rounds.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MiniMax-M2.7
DeepSeek V4 Pro
GPT-5.5
Kimi K2.6
Baseline
56.18 / 33.71
64.20 / 37.04
77.53 / 57.30
57.30 / 23.60
Advisor
59.55 / 20.22
70.37 / 37.04
87.64 / 52.81
60.67 / 30.34
Self-reflection
59.55 / 16.85
74.07 / 54.32
82.02 / 61.80
60.67 / 33.71
All-at-once
60.67 / 25.84
77.78 / 43.21
84.27 / 60.67
62.92 / 32.58
Step-by-step
53.93 / 29.21
79.01 / 50.62
85.39 / 61.80
61.80 / 33.71
DENSE (ours)
66.29 / 39.33
80.25 / 53.09
86.52 / 64.04
67.42 / 34.83
Appendix
Table 5: Task coverage and consistency: pass@3 / pass3 (%). MiniMax, GPT, and Kimi each use 89 tasks, and DeepSeek uses 81, with three results per task. † Privileged reference. Each metric is ranked separately: bold / underline mark the highest / second-highest values, excluding VF; ties share a mark.
Method
Execution ($)
Generator ($)
Total ()/\Delta$
Pass (%)
Baseline
1.8085
—
1.8085
68.54
Advisor
1.7067
0.0024
1.7091 / -5.5%
69.29
Self-reflection
1.2596
—
1.2596 / -30.4%
72.28
All-at-once
1.1565
0.0138
1.1704 / -35.3%
73.03
Step-by-step
1.5947
0.1207
1.7154 / -5.2%
73.03
DENSE (ours)
1.0075
0.1396
1.1470 / -36.6%
75.66
Appendix
Table 6: Mean estimated USD costs over 89 tasks with three repetitions each. Baseline is the initial GPT-5.5 execution; other rows show the GPT-5.5 rerun, Gemini feedback, and their sum. Δ is the sum’s percentage change from baseline, computed before rounding. Dashes mean no feedback generation. Bold / underline mark the best / second-best values: lower cost and higher pass rate are better; feedback is ranked among generators.
Visible component
Interpretation
Root summary , state , and completion-scope
Overall historical progress and the scope of the completion judgment.
Nested task blocks
Subtask boundaries, summaries, and completion states.
shortcut , outcome , and source
A compressed group of attempts, its local result, and source nodes.
dead-end / working-path
Unproductive attempts and useful procedures or intermediate results.
open-issue / evidence
A remaining problem and the historical evidence supporting it.
key-action , tool-call , and observation
Selected executable details and their observed consequences.
Appendix
Table 7: Reading the recipient-visible DENSE artifact. Field names follow the emitted report; the explanation describes their role rather than adding information to the recipient’s input.
Figure 6: Structure and rendering for largest-eigenval , Kimi K2.6, repeat 2. This is a simplified view of the actual artifact, with source IDs retained and node descriptions shortened. Colors and explicit state labels distinguish completed content from unresolved content. The right panel paraphrases the rendered organization; Figure C.2 provides the original wording.
Arm
Extractor input and unit
Information delivered to the recipient
Advisor
Task instruction only; task-level advice.
General strategy, risks, and a completion checklist.
Self-reflection
Serialized source trajectory; no externally generated analysis.
Raw actions and observations for the next-round agent to summarize and reflect on.
All-at-once
One joint analysis of the full trajectory.
Action-level assessments and rationales alongside the trajectory.
Step-by-step
One decision-time prefix per target action, excluding its result and later events.
Action-level assessments with a “Known at the time” basis. The delivered table also contains the source observations.
DENSE (ours)
Nested subtasks, grouped attempts, and issue reconciliation.
Task states, reusable paths, unresolved issues, and selected evidence.
Appendix
Table 8: Input views and information processing in the five compared conditions. The Step-by-step feedback in these cases identifies its view as causal_prefix ; the restriction applies to each scoring call.
Arm / source
Retained feedback span
Information expressed
Advisor Strategy 5
“Format the discovered matches (one per line) and write them to /app/recovered_passwords.txt.”
Explicit delivery requirement, without the discovered fragments.
Self-reflection Observation 22
PASSWORD=8XDP5Q2RT9Z … K7VB3BV4WW54
Both fragments are available; no extractor adds a synthesis.
All-at-once Reasoning 29
“Utilizing byte offsets to locate strings is a high-quality forensic approach that helps pinpoint the exact location of the target data.”
A positive assessment of locating byte offsets within the search.
Step-by-step Reasoning 29
“Using grep -boa to find offsets is the correct forensic step to then examine the raw bytes around that location and reconstruct the full 23-character password.”
Prefix evidence motivates further byte inspection to reconstruct the string.
DENSE Root issue, n26
“The identified password ’8XDP5Q2RT9ZK7VB3BV4WW54’ must be written to /app/recovered_passwords.txt.”
A concrete assembled candidate is bound to the outstanding deliverable.
Appendix
Table 9: Feedback excerpts for the same password-recovery run0, with code formatting normalized. OA row numbers refer to the feedback’s action–observation table. The interpretation column is our analysis, not additional recipient input.
Searches for a full contiguous match (3); carves and inspects archives; attempts a loopback device (25).
25
No
0
All-at-once
Scans known image patterns (2); inspects archives and offsets; prints the assembled candidate and length 23 (26).
26
No
0
Step-by-step
Searches full patterns (1–3); inspects fragments, extracts archives, and parses the central directory (35).
35
No
0
DENSE (ours)
Lists the environment (1); rechecks fragments with strings (2); writes the candidate and reads it back (3).
3
Yes
1
Appendix
Table 10: Observed behavior in password-recovery , Kimi K2.6, repeat 0. Call indices count tool invocations in recorded order; a call can contain several shell operations, and parallel calls are counted separately. “File” means the required output exists at verification.
Method
Retained
Regressed
Repaired
Net gain
Successes
Self-reflection
83
36
20
−16
103
DENSE (ours)
110
9
30
+21
140
VF
82
37
38
+1
120
Appendix
Table 11: MiniMax outcome transitions relative to the shared source run0. Retained and regressed partition the 119 successful sources. Repaired counts successes among the 148 unsuccessful sources. Net gain is repaired minus regressed. Each method has 267 subsequent executions.
Strict pass (0/1)
Tool calls
Condition
0
1
2
0
1
2
Baseline
1
1
1
5
4
5
Self-reflection
1
1
1
4
4
5
DENSE (ours)
1
1
1
4
4
10
VF
0
1
0
0
4
0
Appendix
Table 12: All three repetitions of log-summary-date-ranges on MiniMax. Tool calls are counted from actual transcript records, excluding historical commands inside prompts and tool-like text in responses. Run0 is the source reference. Repetition indices follow the recorded 0, 1, 2 convention.
Task / recipient
Advisor
Self-reflection
All-at-once
Step-by-step
DENSE (ours)
password-recovery Kimi K2.6
0.0 (0/3)
0.0 (0/3)
0.0 (0/3)
0.0 (0/3)
66.7 (2/3)
largest-eigenval Kimi K2.6
0.0 (0/3)
33.3 (1/3)
33.3 (1/3)
33.3 (1/3)
100.0 (3/3)
gpt2-codegolf GPT-5.5
33.3 (1/3)
0.0 (0/3)
0.0 (0/3)
33.3 (1/3)
66.7 (2/3)
Appendix
Table 13: Three-repeat outcomes. Entries report strict pass rate (%), with strict passes out of three in parentheses. The process analysis uses password-recovery repeat 0, largest-eigenval repeat 2, and gpt2-codegolf repeat 0. Within each case, bold / underline mark the highest / second-highest rates, including ties.
Prompt
Role
Visible input
P E.1
Recipient wrapper
Generated feedback, followed by the original task
P E.2
Task-only Advisor
Original task instruction only
P E.2
Global action scoring
Task, complete OA sequence, and action indices
P E.2
Prefix action scoring
Task and prefix through the target action; its result is hidden
P E.3
Subtask boundaries
Root task, current level, and consecutively numbered queue
P E.3
Local action scoring
Root task, newly formed subtask, target action and observation
Appendix
Table 14: Prompt roles and input views. The recipient wrapper is static text; the remaining rows identify analysis-model roles.
Figure 8: Strict pass rates across all seven conditions, using the outcomes in Table 2 . Strict pass requires reward exactly one; rates average three repetitions within each task and then weight tasks equally. MiniMax, GPT, and Kimi use 89 tasks each, and DeepSeek uses 81. Shared zero-based axes show absolute performance on the same percentage scale. DENSE leads the tested non-privileged methods for every recipient. VF uses additional outcome information and exceeds DENSE on GPT and Kimi, but not on MiniMax or DeepSeek, so it serves as a privileged reference rather than a guaranteed upper bound.
Figure 9: Paired strict pass change from the initial execution, using Table 2 . Each value averages the run1-minus-run0 binary success difference over matched repetitions and equally weighted tasks, expressed in percentage points rather than relative percentages. Zero marks unchanged performance; negative bars indicate degradation. Bars show methods without post-hoc verifier information, and gray dashed lines show VF. DENSE gains 7.87, 15.64, 7.12, and 11.61 pp on MiniMax, DeepSeek, GPT, and Kimi, respectively. On MiniMax, it is the only non-privileged feedback method with a positive change.
Figure 10: Task coverage ( pass@3 ) and three-result consistency ( pass3 ), computed from each task’s three binary outcomes as in Equation 6 and Table 5 . The blue bar counts tasks solved at least once; the gold bar requires all three attempts to pass. Their difference is the fraction of tasks solved in exactly one or two attempts. All applicable tasks, including those never solved, remain in the denominators: 89 for MiniMax, GPT, and Kimi, and 81 for DeepSeek. DENSE improves both metrics over Baseline on all four recipients; differing method rankings distinguish broad task coverage from repeatable success.
Figure 11: Observed recipient-token change, computed from Table 3 as 100(Tg/T0−1) , where each mean first averages repetitions within tasks and then weights tasks equally. This is a ratio of aggregate means, not an average of task-specific percentage changes. Tokens include input, output, and cache reads and writes; feedback generation is excluded, and incomplete logs contribute recorded completed calls only (Appendix A.3 ). Negative bars indicate lower execution use. DENSE reduces it by 19.0–43.6% across recipients, while Step-by-step achieves the largest reduction on DeepSeek.
Figure 12: Estimated USD per result for GPT-5.5, using the recorded input/output usage and fixed model-specific prices in Appendix B.2 . Costs average three repetitions within each of 89 tasks, then weight tasks equally. Method-colored segments denote recipient execution; gray segments denote Gemini feedback generation. Baseline and the dashed line show run0; other bars combine run1 and feedback without adding the source execution. DENSE totals 1.15,including0.14 for feedback, versus $1.81 for run0 (Table 6 ). The 36.6% reduction includes feedback generation under this pricing configuration; these standard base-rate estimates do not represent historical invoices.
Figure 13: Strict pass rate versus observed execution plus feedback-generation tokens per result. Each point aggregates one model–method condition over three repetitions per task with equal task weights; VF is excluded. Horizontal values sum recipient and attributed generator tokens in millions, while vertical values match Table 2 . Source run0 is not added to run1. Upper-left positions combine higher success with fewer tokens, but panel scales differ. DENSE has the highest pass rate in every panel, while its combined token use exceeds Baseline for all four recipients. Including generation thus exposes overhead absent from the execution-only comparison in Figure 4 ; token totals do not apply the model-specific dollar prices used in Figure 12 .
Figure 14: Task coverage ( pass@3 ) versus observed tokens for three results per task. Coverage is the fraction of applicable tasks with at least one strict success (Table 5 ); never-solved tasks remain in the denominator. Horizontal values sum execution and attributed feedback tokens across all three results before averaging over tasks, giving three times the per-result values in Figure 13 . DENSE has the highest coverage among the plotted conditions on MiniMax, DeepSeek, and Kimi. On GPT, Advisor reaches slightly higher coverage (87.64% versus 86.52%), illustrating that coverage alone does not capture how consistently a method solves those tasks.
Figure 15: Three-result consistency ( pass3 ) versus observed tokens for three results per task. Horizontal values and per-model horizontal axes match Figure 14 ; only the success criterion changes. A task contributes to the numerator only when all three attempts receive reward exactly one, with all applicable tasks retained in the denominator (Table 5 ). DENSE leads the plotted conditions on MiniMax, GPT, and Kimi. On DeepSeek, Self-reflection reaches 54.32% versus DENSE’s 53.09%, with fewer total tokens. Reading this plot alongside task coverage separates occasional success from success maintained across all three measured attempts.
Figure 16: MiniMax-M2.7: all 89 tasks and seven conditions, three results per cell. Each cell reports strict pass rate (%) over three repetitions.
Figure 17: DeepSeek V4 Pro: 81 applicable tasks and eight visual tasks marked N/A. Each cell reports strict pass rate (%) over three repetitions.
Figure 18: GPT-5.5: all 89 tasks and seven conditions, three results per cell. Each cell reports strict pass rate (%) over three repetitions.
Figure 19: Kimi K2.6: all 89 tasks and seven conditions, three results per cell. Each cell reports strict pass rate (%) over three repetitions.
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
Jonathan Light, Christopher Zhang Cui, Jeonghye Kim +7
On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored. In this work, we instantiate Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information. We identify that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal uninformative. We term this phenomenon decoding collapse, a failure mode that can be missed by existing evaluation metrics. To understand its underlying cause, we show that although the self-teacher achieves stronger performance, learning remains inherently unstable due to inconsistent supervision signals. We further decompose this inconsistency into model inconsistency and prompt inconsistency, and show that the latter can significantly degrade the quality of the supervision signal, limiting the effectiveness of self-teacher learning. To mitigate this inconsistency, we introduce an exponential moving average (EMA) teacher to stabilize the self-teacher and provide more consistent supervision signals. Although the EMA teacher requires a warm-up phase during which performance may temporarily regress, it ultimately improves model performance by providing more stable supervision.
Fan Yang, Rui Meng, Yuxin Wen
Chapman University, Orange, CA, USA · Lawrence Berkeley National Laboratory, Berkeley, CA, USA
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level action-output signals, or by using feedback-conditioned self-distillation. However, generating feedback at every turn is inefficient when many intermediate turns are already successful or neutral, and applying feedback at a fixed or misaligned turn often fails to supervise the actions that contributed to the failure. To bridge this gap, we propose HINT-SD, a targeted self-distillation framework that uses full-trajectory hindsight to select failure-relevant actions and applies feedback-conditioned distillation only on targeted action spans. Experiments on BFCL v3 and AppWorld show that our method improves over the dense per-turn feedback baseline by up to 18.80 percent while achieving 2.26× lower time per training step, suggesting that selecting where to distill is a key factor for both effective and efficient long-horizon agent training.