After the Fix: Transfer of Corrected Agent Experience
Organizations: International Digital Economy Academy (IDEA)
Abstract
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox's Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full's 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full's 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX's accepted execution reaches 52% versus its summary's 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.
Figures & tables
| Format | Episode content given to B |
|---|---|
| None | Shared state channel only; no additional episode text. |
| Full | Chronological actions, observations, and any failure/feedback/repair, subject to trace limits. |
| Skill | Abstract procedures, constraints, checks, and error–correction rules. |
| Hybrid | ThinkingBox: history summary plus four recent raw assistant segments. APEX: the same Skill plus grounded facts and provenance. |
| Summary | C-only; at most 350 requested words on accepted results, entities, and decisive actions; no earlier failures/feedback. |
| Execution (Exec) | Identical Summary plus the final accepted segment’s actions, tool calls, and observations; not the entire repair history. |
| Condition | ThinkingBox | State | Text |
|---|---|---|---|
| B-only | 42 | 44 | 47 |
| U-None / C-None | 37 / 36 | 45 / 44 | 45 / 47 |
| U-Full / C-Full | 23 / 67 | 42 / 41 | 45 / 45 |
| U-Skill / C-Skill | 35 / 64 | 39 / 42 | 44 / 45 |
| U-Hybrid / C-Hybrid | 28 / 60 | 41 / 45 | 48 / 40 |
| Summary / Execution | 56 / 67 | 40 / 44 | 40 / 52 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Regime | B’s starting environment | A-derived information |
|---|---|---|
| ThinkingBox | Independent standard B sandbox | Terminal-state description and the specified textual representation; no live A database restoration |
| APEX-state | Source environment state plus the specified text; None supplies the state channel alone | |
| APEX-text | Text alone; None supplies terminal observations, with the same branch prefix in Full/Skill/Hybrid |
| Contrast | Held fixed | What the contrast measures |
|---|---|---|
| C–U, same representation | Target and evaluation | The complete correction-associated experience package |
| C–B-only | Target and evaluation | Benefit beyond independent target execution |
| Execution–Summary | Exact summary, framing, starting state | Adding the accepted segment, including its content and length |
| Full–Execution | Target and evaluation | Alternative whole handoffs, not history alone |
| APEX-text–APEX-state | Pair identities and target rubric | Implemented handoff regimes, with state and textual-construction differences |
| Component | ThinkingBox | APEX-state |
|---|---|---|
| Acting model | Environment-resolved DeepSeek alias; adapter default deepseek-v4-flash | DeepSeek V4.1 Flash (multimodal), API deepseek-flash |
| Agent sampling | Thinking enabled; explicit reasoning_effort=high ; temperature omitted | Thinking enabled; no explicit reasoning effort; temperature 0.2; no supplied seed |
| Agent budgets | 32,768 output; 400 interactions; 30 user turns; 1,200 seconds | 8,192 output; 100 steps; 7,200 seconds |
| Runtime user | Temperature 0.2; 4,096 output; cannot terminate agent conversation | No separate conversational user during B |
| Correction review | Temperature 0.2; 2,048 output; trace and state differences | Temperature 0; 8,192 output; artifacts and criteria |
| Memory compressor | Temperature 0; 8,192 output; thinking disabled | deepseek-v4-pro ; temperature 0; 8,192 output |
| Condition | ThinkingBox | APEX-state A/Q | APEX-text A/Q |
|---|---|---|---|
| B-only | 42 | 44 / 55.54 | 47 / 57.40 |
| U-None | 37 | 45 / 54.79 | 45 / 53.50 |
| U-Full | 23 | 42 / 54.00 | 45 / 55.83 |
| U-Skill | 35 | 39 / 50.89 | 44 / 54.29 |
| U-Hybrid | 28 | 41 / 53.11 | 48 / 56.80 |
| C-None | 36 | 44 / 55.38 | 47 / 56.10 |
| Regime | Cmp. | G/L | Gap | Cluster CI | Raw | ||
|---|---|---|---|---|---|---|---|
| TB | C-F / U-F | 46/2 | +44 | [25.3, 56.3] | |||
| TB | C-S / U-S | 39/10 | +29 | [12.7, 39.8] | 0.0003 | 0.0008 | |
| TB | C-H / U-H | 36/4 | +32 | [19.1, 44.0] | |||
| TB | C-F / B | 31/6 | +25 | [13.1, 37.7] | 0.0004 | 0.0009 | |
| TB | C-S / B | 30/8 | +22 | [8.0, 39.0] | 0.0005 | 0.0038 | 0.0094 |
| TB | C-H / B | 24/6 | +18 | [6.7, 32.5] | 0.0014 | 0.0100 | 0.0272 |
| Regime | Hypothesis | Gap | Cluster CI | Group | ||
|---|---|---|---|---|---|---|
| TB | H1a-full | +44 | [25.3, 56.2] | 0.0054 | 0.0439 | 0.1235 |
| TB | H1a-skill | +29 | [12.8, 40.0] | 0.0259 | 0.1812 | 0.4917 |
| TB | H1a-hybrid | +32 | [19.3, 44.0] | 0.0049 | 0.0439 | 0.1172 |
| TB | H1b-full | +25 | [12.9, 37.6] | 0.0078 | 0.0703 | 0.1719 |
| TB | H1b-skill | +22 | [8.0, 39.0] | 0.0203 | 0.1418 | 0.4053 |
| TB | H1b-hybrid | +18 | [6.5, 32.5] | 0.0137 | 0.1094 | 0.2871 |
| Regime | Comparison | G/L | Gap | Cluster CI | Target | Group |
|---|---|---|---|---|---|---|
| TB | U-F–U-N | 4/18 | -14 | [-27.2, -4.3] | 0.1433 | 0.7734 |
| TB | U-S–U-N | 12/14 | -2 | [-10.8, 4.8] | 1.0000 | 1.0000 |
| TB | U-H–U-N | 9/18 | -9 | [-25.4, 3.3] | 1.0000 | 1.0000 |
| TB | U-S–U-F | 16/4 | +12 | [2.0, 22.9] | 0.3782 | 1.0000 |
| TB | U-H–U-F | 10/5 | +5 | [-4.3, 12.9] | 1.0000 | 1.0000 |
| TB | U-H–U-S | 8/15 | -7 | [-20.7, 3.6] | 1.0000 | 1.0000 |
| Regime | Format contrast | Gap difference | Cluster CI | Group |
|---|---|---|---|---|
| TB | Full–None | +45 | [27.8, 61.9] | 0.0352 |
| TB | Skill–None | +30 | [11.2, 48.6] | 0.3945 |
| TB | Hybrid–None | +33 | [14.9, 53.9] | 0.0623 |
| TB | Skill–Full | -15 | [-28.7, -1.4] | 1.0000 |
| TB | Hybrid–Full | -12 | [-22.6, 2.4] | 1.0000 |
| TB | Hybrid–Skill | +3 | [-11.3, 21.4] | 1.0000 |
| ThinkingBox | APEX-state | APEX-text | |||||||
| B/U/C | F | S | H | F | S | H | F | S | H |
| 000 | 26 | 22 | 31 | 43 | 40 | 42 | 38 | 37 | 38 |
| 001 | 24 | 21 | 16 | 4 | 7 | 6 | 5 | 5 | 2 |
| 010 | 1 | 6 | 3 | 3 | 4 | 3 | 4 | 7 | 8 |
| 011 | 7 | 9 | 8 | 6 | 5 | 5 | 6 | 4 | 5 |
| 100 | 5 | 4 | 5 | 7 | 6 | 5 | 7 | 5 | 9 |
| Relationship | Pairs | Shared basis | What changes in B |
|---|---|---|---|
| Same-family variant | 42 | Same computation, model, or legal decision test | LBO debt/rate assumptions; a different insured claim |
| Shared intermediate procedure | 26 | Identifiable reusable subprocedure; distinct end operation | Comparable filtering for peer benchmarking versus LBO inputs |
| Shared evidence | 20 | Common case materials or model; distinct operations | Manual drafting versus compliance comparison |
| Context only | 12 | Same project or domain; no specific shared procedure or material established | Board-tenure comparison versus CAPM estimation |
| Relationship | Count | Full | Skill | Hybrid | E–S |
|---|---|---|---|---|---|
| Same-family | 42 | -4.8 / +0.0 | +7.1 / -7.1 | -2.4 / -9.5 | +2.4 / +14.3 |
| Shared procedure | 26 | +0.0 / +0.0 | +0.0 / +19.2 | +19.2 / -3.8 | +3.8 / +3.8 |
| Shared evidence | 20 | -5.0 / +5.0 | -5.0 / -5.0 | -5.0 / -5.0 | +15.0 / +15.0 |
| Context only | 12 | +16.7 / -8.3 | +8.3 / +0.0 | +8.3 / -16.7 | -8.3 / +16.7 |
| Domain | Count | Full | Skill | Hybrid | E–S |
|---|---|---|---|---|---|
| Banking | 48 | -10.4 / -6.2 | +0.0 / -6.2 | +0.0 / -18.8 | +14.6 / +10.4 |
| Law | 23 | +4.3 / +0.0 | +17.4 / +30.4 | +17.4 / +4.3 | -8.7 / +21.7 |
| Consulting | 29 | +10.3 / +10.3 | -3.4 / -10.3 | +0.0 / +0.0 | -3.4 / +6.9 |
| Improved | Regressed | |||||
|---|---|---|---|---|---|---|
| Regime | All 3 | All 3 | ||||
| ThinkingBox | 61 | 42 | 18 | 15 | 1 | 0 |
| APEX-state | 29 | 5 | 1 | 22 | 6 | 1 |
| APEX-text | 22 | 6 | 1 | 28 | 7 | 1 |
| Contrast | State gains | Text gains | Gain in both | Loss in both | G to L | L to G |
|---|---|---|---|---|---|---|
| Full | 8 | 10 | 2 | 2 | 0 | 0 |
| Skill | 15 | 14 | 5 | 5 | 3 | 2 |
| Hybrid | 12 | 5 | 2 | 4 | 4 | 0 |
| Execution–Summary | 12 | 14 | 2 | 0 | 0 | 4 |
| Contrast | State | Text | Difference | World interval |
|---|---|---|---|---|
| C-Full–U-Full | -1 | 0 | +1 | |
| C-Skill–U-Skill | +3 | +1 | -2 | |
| C-Hybrid–U-Hybrid | +4 | -8 | -12 | |
| Execution–Summary | +4 | +12 | +8 | |
| Full–Execution | -3 | -7 | -4 |
| Comparison | Before | After | Removed | Then pass | Still fail |
|---|---|---|---|---|---|
| U-Full to C-Full | 55 | 2 | 54 | 37 | 17 |
| U-Skill to C-Skill | 40 | 7 | 39 | 28 | 11 |
| U-Hybrid to C-Hybrid | 48 | 3 | 48 | 30 | 18 |
| Summary to Execution | 13 | 1 | 13 | 8 | 5 |
| Regime | Comparison | Clusters | Minimum | Maximum |
|---|---|---|---|---|
| ThinkingBox | C-F / U-F | 15 | +40.66 | +48.42 |
| ThinkingBox | C-S / U-S | 15 | +25.27 | +31.96 |
| ThinkingBox | C-H / U-H | 15 | +27.78 | +34.74 |
| ThinkingBox | E / S | 15 | +8.64 | +13.33 |
| APEX-state | C-F / U-F | 16 | -2.22 | +1.08 |
| APEX-state | C-S / U-S | 16 | +1.05 | +4.30 |
| Requested quantity | Summary | Execution | Reference |
|---|---|---|---|
| Current stake | 4,935.9 | 5,499.7 | 5,499.7 |
| PV of revised cash flows | 4,389.4 | 4,790.1 | 4,790.1 |
| Discounted terminal value | 17,822.9 | 19,959.1 | 19,959.0 |
| Revised stake | 4,442.5 | 4,949.8 | 4,949.8 |
| Percentage loss | 10.0% | 10.0% | 10.0% |
| Criteria satisfied | 1/5 | 5/5 | – |
| Regime | Summary | Execution | Full | G/L | E–S | |
|---|---|---|---|---|---|---|
| ThinkingBox | 56 | 67 | 67 | 15/4 | +11 | 0.0961 |
| APEX-state | 40 | 44 | 41 | 12/8 | +4 | 1.0000 |
| APEX-text | 40 | 52 | 45 | 14/2 | +12 | 0.0251 |
| Regime | Condition | Accepted | Input M | Calls | Input/call k | Tools |
|---|---|---|---|---|---|---|
| ThinkingBox | B-only | 42 | 0.232 | 9.48 | 24.4 | 14.96 |
| ThinkingBox | Summary | 56 | 0.264 | 9.48 | 27.9 | 17.09 |
| ThinkingBox | Execution | 67 | 0.353 | 9.27 | 38.1 | 18.34 |
| ThinkingBox | C-Full | 67 | 0.434 | 8.85 | 49.0 | 17.07 |
| ThinkingBox | C-Skill | 64 | 0.286 | 9.80 | 29.2 | 17.61 |
| ThinkingBox | C-Hybrid | 60 | 0.318 | 9.82 | 32.4 | 17.44 |
| Regime / set | Completed | Budget failed | Error | Total |
|---|---|---|---|---|
| State: original nine conditions | 895 | 5 | 0 | 900 |
| State: Summary | 92 | 8 | 0 | 100 |
| State: Execution | 88 | 10 | 2 | 100 |
| Text: all eleven conditions | 1,072 | 28 | 0 | 1,100 |
| Study | How experience transfers | Main comparison axis |
|---|---|---|
| ExpeL ( Zhao et al., 2024 ) | Knowledge extraction and experience retrieval | Accumulation and transfer |
| CLIN ( Majumder et al., 2024 ) | Causal abstractions and meta-memory | Adaptation and generalization |
| AWM ( Wang et al., 2025 ) | Induced reusable workflows | Offline/online and cross-domain |
| ReasoningBank ( Ouyang et al., 2026 ) | Strategies from success and failure | Memory; exploration via MaTTS |
| Feng et al. (2026) | Task- or subtask-induced skills | Granularity text/code |
| SkillsBench ( Li et al., 2026a ) | Curated or generated skills | Skill source, domain, and utility |