MemTrace: State-Consistent Memory for Long-Horizon Coding Agents
Organizations: Shanghai Jiao Tong University · MemTensor (Shanghai) Technology Co., Ltd. · Theseus Lab
Abstract
As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these challenges through techniques like larger context windows, compression, retrieval, or repository representations, but often fail to reconstruct a consistent task state after a context refresh or verify whether recalled evidence remains valid. Thus, we introduce MemTrace, a provenance-aware memory system that preserves execution history and aligns its reuse with the evolving task (e.g., iterative cross-file repair) and repository state. MemTrace stores history as immutable Memory Traces anchored to key information (e.g., files, symbols, tests), and organizes their execution order and dependencies in a Memory Trace Graph. When context is constrained, working memory retains only compact Memory Anchors, from which the agent can reconstruct the latest execution state and locate evidence relevant to its next action. Before restoring historical evidence, MemTrace checks its validity against the current repository state and retrieves only what the next action requires. Across three complementary long-horizon coding benchmarks, MemTrace consistently outperforms all fully evaluated baselines under the same backbone and harness, improving DeepSWE pass@1 by 21.2 points, SWE-EVO Resolved Rate by 4.4 points, and SWE-Milestone Score by 17.8 points under Codex CLI.
Figures & tables
| Method | Context management | Execution state | Repository alignment | ||
| Recoverable history | Structured task state | Revision history | Code graph | Repo-aware evidence reuse | |
| Observation Masking | |||||
| MemGPT | |||||
| Context as a Tool | |||||
| Self-GC | |||||
| Scroll | |||||
| DeepSWE | SWE-EVO | SWE-Milestone | ||||||
| Method | pass@1 | Resolved | Fix | Apply | Score | Prec. | Recall | Resolved |
| Harness: Codex CLI | ||||||||
| Native | 35.4 (32.4) | 47.8 | 53.4 | 97.8 | 28.6 (30.7) | 26.9 (26.4) | 41.9 (84.5) | 7.1 ( 8.3 ) |
| Obs. Masking | 34.5 (38.2) | 65.2 | 72.1 | 97.8 | 24.4 (30.0) | 24.2 (26.0) | 37.1 (84.9) | 9.2 ( 8.3 ) |
| RepoGraph | – ( 41.2 ) | 58.7 | 68.8 | 97.8 | – ( 45.2 ) | – ( 36.7 ) | – ( 91.5 ) | – ( 8.3 ) |
| MemTrace | 56.6 (67.6) | 69.6 | 70.6 | 100.0 | 46.4 (46.1) | 43.6 (37.8) | 60.9 (91.7) | 16.3 (16.7) |
| Variant | Component | SWE-EVO Resolved Rate (%) | DeepSWE pass@1 (%) | ||
|---|---|---|---|---|---|
| Trace Store | MTG | RSG | |||
| Native | – | – | – | 47.8 | 35.4 |
| Trace Store Only | ✓ | – | – | 65.2 | 32.1 |
| MemTrace w/o MTG | ✓ | – | ✓ | 65.2 | 34.5 |
| MemTrace w/o RSG | ✓ | ✓ | – | 69.6 | 44.2 |
| MemTrace | ✓ | ✓ | ✓ | 69.6 | 56.6 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| DeepSWE | SWE-EVO | SWE-Milestone | ||
|---|---|---|---|---|
| Language | Tasks | Tasks | Itineraries | Graded milestones |
| Python | 34 | 48 | 1 | 12 |
| TypeScript | 35 | 0 | 1 | 18 |
| Go | 34 | 0 | 2 | 32 |
| JavaScript | 5 | 0 | 0 | 0 |
| Rust | 5 | 0 | 2 | 24 |
| DeepSWE | SWE-EVO | ||||
|---|---|---|---|---|---|
| Harness | Method | All (113) | Python (34) | Non-Python (79) | All (46) |
| Codex CLI | Native | 40 | 11 | 29 | 22 |
| Obs. Masking | 39 | 13 | 26 | 30 | |
| RepoGraph | – | 14 | – | 27 | |
| MemTrace | 64 | 23 | 41 | 32 | |
| mini-swe-agent | Native | 55 | 16 | 39 | 24 |
| Quantity | Count |
|---|---|
| Event-instrumented trajectories | 85 |
| Trajectories with eligible file-bound Memory Traces | 84 |
| Eligible time-aligned, file-bound Memory Traces | 9,226 |
| Memory Traces with a later bound-file change | 5,261 |
| Repository revision events | 2,897 |
| Semantic evidence invalidations | 16,197 |
| Variables | Spearman [95% CI] | Pearson | |
|---|---|---|---|
| Revisions vs. exposure | 84 | 0.394 | |
| Revisions vs. invalidations | 85 | 0.881 |
| Event-step age | [95% CI] | Memory Traces at risk |
|---|---|---|
| 97 | 6,380 | |
| 1,001 | 3,654 | |
| 2,012 | 2,404 |
| Group | Native responses | Median | Tasks |
|---|---|---|---|
| Short | 10–197 | 145 | 37 |
| Medium | 198–290 | 248 | 38 |
| Long | 291–526 | 343 | 38 |
| Metric | Short | Medium | Long |
|---|---|---|---|
| Stored Memory Traces | 79 [61,112] | 102.5 [74.25,160.25] | 129 [92.5,218.5] |
| Repository revisions | 21 [17,27] | 30 [20,37] | 41 [26,49] |
| Evidence invalidations | 122 [66,177] | 142 [79,231] | 223 [142.5,350.5] |
| Context Refreshes | 2 [1,2] | 2 [2,4] | 3 [2,4.5] |
| Exposure (%) | 54.4 [50.0,64.4] | 49.7 [36.5,60.7] | 64.4 [50.0,75.5] |
| Metric | Spearman [95% CI] | Long–Short [95% CI] | |
|---|---|---|---|
| Stored Memory Traces | 113 | ||
| Repository revisions | 85 | ||
| Evidence invalidations | 85 | ||
| Context Refreshes | 85 | ||
| Exposure (pp) | 84 |
| Metric | Responses | Tool calls | Output tokens |
|---|---|---|---|
| Stored Memory Traces | |||
| Repository revisions | |||
| Evidence invalidations | |||
| Context Refreshes | |||
| Exposure |
| Step | Record | Observed information |
|---|---|---|
| 1,061 | Memory Trace sealing | is sealed; its manifest includes the compiler parser-test file. |
| 17,152 | Repository revision | The bound file is modified; the revision records 15 evidence invalidations. |
| 17,986 | Context Refresh | A new context epoch is activated after model observation. |
| 18,680 | Semantic update | Two semantic updates are accepted for M003. |
| 21,392 | Runtime verification | The comparison test command exits with code 0; criterion M003.C001 is marked successful. |