Organizations: Brown University · Purdue University · Carnegie Mellon University · Harvard University · Princeton University · Pennsylvania State University · Independent Researcher · Utah State University · Massachusetts Institute of Technology
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state, tool semantics, or task progress. Procedural memory promises more consistent decisions and less redundant exploration, but constructing high-quality memory without model training remains challenging. We introduce CONTRAMEM, a source-flexible, training-free framework for self-evolving procedural memory that treats same-task outcome variation as supervision: differences in correctness, efficiency, recovery, and failure modes expose outcome-relevant procedural distinctions, distilled into a compact bank of app-level Function Cards and task-level Skill Cards that evolves through localized curation rather than append-only accumulation or whole-bank rewriting. On held-out GAIA2/ARE computer-use tasks, CONTRAMEM more than doubles the success rate across the three source-model targets (26.2% to 55.3%), with consistent per-model gains (GPT-5.5: 27.5 to 61.0; Claude Sonnet 4.6: 28.0 to 52.5; DeepSeek V4 Pro: 23.0 to 52.5). The same bank transfers unchanged to the unseen Qwen3.7 Plus (18.5 to 35.5), indicating transferable procedural knowledge rather than model-specific behavior. The same construction carries over unchanged to AppWorld, beating both no memory and its own single-source self-memory variant for all three mid-tier agents on both public test splits. Under a matched trajectory budget, heterogeneous multi-model trajectories yield stronger memory than self- or same-model multi-rollout memory: the margin comes from contrastive behavioral diversity, not stronger source agents or more sampling.
Figures & tables
Figure 1: Held-out success averaged over three source targets (GAIA2/ARE) and target-macro TGC/SGC (AppWorld): ContraMem beats self memory on every metric.
Figure 2: Overview of ContraMem : heterogeneous same-task trajectories (A) are compared by the Reflector to isolate outcome-relevant procedural differences (B); the Curator uses these signals to refine Function and Skill Cards in the current memory bank (C), from which a compact task-relevant subset guides a single target agent at runtime (D).
Target
Method
Exec.
Search
Ambig.
Adapt.
Time
Overall
Source target models
GPT-5.5
No memory
47.5
52.5
12.5
17.5
7.5
27.5
Self memory
72.5
92.5
40.0
35.0
12.5
50.5 (+23.0)
ContraMem
80.0
100.0
52.5
55.0
17.5
61.0 (+33.5)
Claude Sonnet
No memory
40.0
57.5
17.5
22.5
2.5
28.0
Self memory
60.0
80.0
52.5
52.5
10.0
51.0 (+23.0)
Table 1: Held-out success rates (%) on GAIA2/ARE (40 tasks per ability and target). Self memory applies the same pipeline to only the target’s own reference trajectories; ContraMem uses the shared three-model bank. Qwen3.7 Plus is unseen during construction, hence no self-memory row. Bold: best per column; parentheses: absolute gain over no memory.
Test-Normal
Test-Challenge
Target
Method
TGC
SGC
Δ Base
TGC
SGC
Δ Base
DeepSeek-V3.1
No memory
51.2
30.4
–
56.8
38.1
–
Self memory
56.0
35.7
+4.8
64.3
46.8
+7.4
ContraMem
76.2
58.9
+25.0
70.0
56.1
+13.2
Qwen3.6-Flash
No memory
63.7
42.9
–
46.0
31.7
–
Self memory
73.2
58.9
+9.5
50.8
37.4
+4.8
Table 2: AppWorld task (TGC) and scenario (SGC) goal completion (%) on both public test splits. Self memory applies the pipeline to each target’s own train-split trajectories (self-curated); ContraMem uses the shared three-model bank; all memories are frozen. Bold: best. Δ Base: TGC gain over no memory.
Memory variant
Exec.
Search
Ambig.
Macro
No memory
47.5
52.5
12.5
37.5
GPT-5.5 self memory
72.5
92.5
40.0
68.3
GPT-5.5 multi-rollout
75.0
92.5
42.5
70.0
Function Cards only
67.5
85.0
30.0
60.8
Skill Cards only
60.0
100.0
40.0
66.7
ContraMem (three models)
80.0
100.0
52.5
77.5
Table 3: GPT-5.5 held-out source and component ablations. Multi-rollout uses three GPT-5.5 trajectories, matching ContraMem ’s count; single-card rows use the shared bank and frozen retrieval path. Macro averages the three abilities.
Figure 3: ContraMem versus trajectory- and playbook-based memory on GPT-5.5 held-out tasks. All methods share the same source pool and scenarios. Exact values: Appendix Table A5.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7
Ability
Avg No
Avg Ours
Δ
Exec
30.3
31.1
+0.8
Search
55.6
40.2
-15.5
Ambig
29.5
24.3
-5.2
Adapt
19.0
20.7
+1.6
Time
70.9
61.0
-9.9
Macro
41.1
35.5
-5.6
Appendix
Table A2: Agent-event efficiency on GAIA2/ARE held-out tasks, paired by scenario. Negative Δ means that ContraMem executes the task with fewer agent-level events.
Figure 9
Ability
Raw Cards
Curated Cards
Card Reduction
Append-Only Success
Curated Success
Search
74
62
16.2%
39/40 (97.5%)
40/40 (100.0%)
Ambiguity
69
47
31.9%
19/40 (47.5%)
21/40 (52.5%)
Execution
75
59
21.3%
26/40 (65.0%)
32/40 (80.0%)
Adaptability
55
37
32.7%
20/40 (50.0%)
22/40 (55.0%)
Total
273
205
24.9%
104/160 (65.0%)
115/160 (71.9%)
Appendix
Table A4: Curator ablation on GPT-5.5 held-out tasks. Append-only commits every Reflector delta as Add , skipping local consolidation; ContraMem uses the Curator to merge duplicates, narrow triggers, reject weak deltas, and keep compact transferable cards. The ablation covers the four primary non-temporal abilities for which append-only banks were constructed.
Method
Exec.
Search
Ambig.
Macro
No memory
47.5
52.5
12.5
37.5
Raw retrieval
45.0
75.0
17.5
45.8
AWM
42.5
85.0
15.0
47.5
ACE
50.0
70.0
12.5
44.2
ContraMem
80.0
100.0
52.5
77.5
Appendix
Table A5: Exact GPT-5.5 held-out success rates underlying the main paper’s baseline figure. Macro is the unweighted mean over Execution, Search, and Ambiguity. Each cell contains the same 40 held-out scenarios.
Table 12Figure 13
Figure B1: Runtime renderings of one Function Card (blue) and one Skill Card (green) exactly as the target agent receives them. The Function Card records a callable tool contract with destructive-write guards; the Skill Card records the decision boundary, completion ledger, and contrastive evidence distilled from same-task trajectory contrast. Field schemas appear in Appendix B.1.
Figure B2: From disagreement to transfer. GPT-5.5 and DeepSeek guess an underspecified shopping variant and fail, whereas Claude completes the safe branch and asks for the missing attribute. The distilled Skill Card retains this decision boundary and guides GPT-5.5 to complete the unambiguous branch and request clarification on a new task.