Organizations: Brown University · Purdue University · Carnegie Mellon University · Harvard University · Princeton University · Pennsylvania State University · Independent Researcher · Utah State University · Massachusetts Institute of Technology
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more consistent decisions and less redundant exploration, but constructing high-quality memory without model training remains challenging. We introduce CONTRAMEM, a source-flexible, training-free framework for self-evolving procedural memory that treats same-task outcome variation as supervision: differences in correctness, efficiency, recovery, and failure modes expose outcome-relevant procedural distinctions, distilled into a compact bank of app-level Function Cards and task-level Skill Cards that evolves through localized curation rather than append-only accumulation or whole-bank rewriting. On held-out GAIA2/ARE computer-use tasks, CONTRAMEM more than doubles the success rate across the three source-model targets (26.2% to 55.3%), with consistent per-model gains (GPT-5.5: 27.5 to 61.0; Claude Sonnet 4.6: 28.0 to 52.5; DeepSeek V4 Pro: 23.0 to 52.5). The same bank transfers unchanged to the unseen Qwen3.7 Plus (18.5 to 35.5), indicating transferable procedural knowledge rather than model-specific behavior. The same construction carries over unchanged to AppWorld, beating both no memory and its own single-source self-memory variant for all three mid-tier agents on both public test splits. Under a matched trajectory budget, heterogeneous multi-model trajectories yield stronger memory than self- or same-model multi-rollout memory: the margin comes from contrastive behavioral diversity, not stronger source agents or more sampling.
Figures & tables
Figure 1: Held-out success averaged over three source targets (GAIA2/ARE) and target-macro TGC/SGC (AppWorld): ContraMem beats self memory on every metric.
Figure 2: Overview of ContraMem : heterogeneous same-task trajectories (A) are compared by the Reflector to isolate outcome-relevant procedural differences (B); the Curator uses these signals to refine Function and Skill Cards in the current memory bank (C), from which a compact task-relevant subset guides a single target agent at runtime (D).
Target
Method
Exec.
Search
Ambig.
Adapt.
Time
Overall
Source target models
GPT-5.5
No memory
47.5
52.5
12.5
17.5
7.5
27.5
Self memory
72.5
92.5
40.0
35.0
12.5
50.5 (+23.0)
ContraMem
80.0
100.0
52.5
55.0
17.5
61.0 (+33.5)
Claude Sonnet
No memory
40.0
57.5
17.5
22.5
2.5
28.0
Self memory
60.0
80.0
52.5
52.5
10.0
51.0 (+23.0)
Table 1: Held-out success rates (%) on GAIA2/ARE (40 tasks per ability and target). Self memory applies the same pipeline to only the target’s own reference trajectories; ContraMem uses the shared three-model bank. Qwen3.7 Plus is unseen during construction, hence no self-memory row. Bold: best per column; parentheses: absolute gain over no memory.
Test-Normal
Test-Challenge
Target
Method
TGC
SGC
Δ Base
TGC
SGC
Δ Base
DeepSeek-V3.1
No memory
51.2
30.4
–
56.8
38.1
–
Self memory
56.0
35.7
+4.8
64.3
46.8
+7.4
ContraMem
76.2
58.9
+25.0
70.0
56.1
+13.2
Qwen3.6-Flash
No memory
63.7
42.9
–
46.0
31.7
–
Self memory
73.2
58.9
+9.5
50.8
37.4
+4.8
Table 2: AppWorld task (TGC) and scenario (SGC) goal completion (%) on both public test splits. Self memory applies the pipeline to each target’s own train-split trajectories (self-curated); ContraMem uses the shared three-model bank; all memories are frozen. Bold: best. Δ Base: TGC gain over no memory.
Memory variant
Exec.
Search
Ambig.
Macro
No memory
47.5
52.5
12.5
37.5
GPT-5.5 self memory
72.5
92.5
40.0
68.3
GPT-5.5 multi-rollout
75.0
92.5
42.5
70.0
Function Cards only
67.5
85.0
30.0
60.8
Skill Cards only
60.0
100.0
40.0
66.7
ContraMem (three models)
80.0
100.0
52.5
77.5
Table 3: GPT-5.5 held-out source and component ablations. Multi-rollout uses three GPT-5.5 trajectories, matching ContraMem ’s count; single-card rows use the shared bank and frozen retrieval path. Macro averages the three abilities.
Figure 3: ContraMem versus trajectory- and playbook-based memory on GPT-5.5 held-out tasks. All methods share the same source pool and scenarios. Exact values: Appendix Table A5.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7
Ability
Avg No
Avg Ours
Δ
Exec
30.3
31.1
+0.8
Search
55.6
40.2
-15.5
Ambig
29.5
24.3
-5.2
Adapt
19.0
20.7
+1.6
Time
70.9
61.0
-9.9
Macro
41.1
35.5
-5.6
Appendix
Table A2: Agent-event efficiency on GAIA2/ARE held-out tasks, paired by scenario. Negative Δ means that ContraMem executes the task with fewer agent-level events.
Figure 9
Ability
Raw Cards
Curated Cards
Card Reduction
Append-Only Success
Curated Success
Search
74
62
16.2%
39/40 (97.5%)
40/40 (100.0%)
Ambiguity
69
47
31.9%
19/40 (47.5%)
21/40 (52.5%)
Execution
75
59
21.3%
26/40 (65.0%)
32/40 (80.0%)
Adaptability
55
37
32.7%
20/40 (50.0%)
22/40 (55.0%)
Total
273
205
24.9%
104/160 (65.0%)
115/160 (71.9%)
Appendix
Table A4: Curator ablation on GPT-5.5 held-out tasks. Append-only commits every Reflector delta as Add , skipping local consolidation; ContraMem uses the Curator to merge duplicates, narrow triggers, reject weak deltas, and keep compact transferable cards. The ablation covers the four primary non-temporal abilities for which append-only banks were constructed.
Method
Exec.
Search
Ambig.
Macro
No memory
47.5
52.5
12.5
37.5
Raw retrieval
45.0
75.0
17.5
45.8
AWM
42.5
85.0
15.0
47.5
ACE
50.0
70.0
12.5
44.2
ContraMem
80.0
100.0
52.5
77.5
Appendix
Table A5: Exact GPT-5.5 held-out success rates underlying the main paper’s baseline figure. Macro is the unweighted mean over Execution, Search, and Ambiguity. Each cell contains the same 40 held-out scenarios.
Table 12Figure 13
Figure B1: Runtime renderings of one Function Card (blue) and one Skill Card (green) exactly as the target agent receives them. The Function Card records a callable tool contract with destructive-write guards; the Skill Card records the decision boundary, completion ledger, and contrastive evidence distilled from same-task trajectory contrast. Field schemas appear in Appendix B.1.
Figure B2: From disagreement to transfer. GPT-5.5 and DeepSeek guess an underspecified shopping variant and fail, whereas Claude completes the safe branch and asks for the missing attribute. The distilled Skill Card retains this decision boundary and guides GPT-5.5 to complete the unambiguous branch and request clarification on a new task.
Long-horizon autonomous agents require memory systems to retain historical information, track evolving states, and reuse relevant knowledge beyond finite context windows. Existing agentic memory systems typically follow a memory construction-retrieval (MCR) pipeline, but often adapt mainly the memory bank while keeping the surrounding pipeline fixed after deployment. This fixed-pipeline design struggles to handle heterogeneous task-specific failure modes and can become misaligned with memory banks that evolve in scale and structure over time. To address these limitations, we propose MemPro, a system-level evolution framework that treats the entire MCR pipeline as an evolvable program rather than adapting only the memory bank or prompt text. MemPro maintains a version tree of runnable memory-system implementations, where an Evolving Agent iteratively selects promising versions, diagnoses recurring failures, and creates improved child versions through failure-mode-guided edit-debug refinement. Experiments on LongMemEval, LoCoMo, HotpotQA, and NarrativeQA show that MemPro consistently outperforms strong static and prompt-level evolving baselines within a few iterations, continues to improve with evolution, and achieves a favorable performance-cost trade-off. Code is available at https://github.com/wanghai673/MemPro.
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones. The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. Experiments show that procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. We further find that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer. These results provide practical guidance for building, evaluating, and deploying procedural memory systems in production agent platforms.
Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge. Prior memory approaches aim to resolve the situation, but mainly focus on storing factual information. Recent work on procedural memory improves task reuse, yet often reduces to replaying past successes without addressing failure cases or online scalability. We introduce a unified and automatic memory framework that integrates semantic, episodic, and procedural memory in a bi-level design combining short-term and long-term stores. A multi-agent architecture with actor, memory, and critic agents enables automatic memory generation, reward annotation, and adaptive retrieval. Long-term memory is managed through reward-based evaluation, merging, and pruning, ensuring scalability and continual improvement. Experiments across various environments show that our approach improves robustness and success on long multi-turn tasks compared to existing baselines. This work highlights the importance of comprehensive, adaptive memory for advancing LLM-based agents.