Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
Figures & tables
Figure 1: Overview of AMBER. At each step, the agent jointly reasons, acts, and generates a memory entry that is appended to a persistent memory bank provided as context to future steps. See text for details.
Figure 2: Rollout architecture. Each task is expanded into a group of G rollouts on a leased website VM, and each rollout gets its own Ray-managed browser session and website container. Red arrows denote HTTP connections. Solid arrows show per-step traffic between the inference server, rollout manager, and environment server, and dashed arrows show per-episode VM lease and release.
Model
Shopping
Admin
Reddit
Map
Multi-site
Average
Closed Models/Zero-Shot with AMBER
GPT-5-mini
28.68
38.10
29.41
25.64
27.78
30.71 ± 1.11
Haiku 4.5
44.96
62.86
52.94
25.64
38.89
46.72 ± 3.24
Qwen 3.5 9B
27.05
28.00
10.72
14.10
16.67
21.82 ± 2.97
Qwen 3.5 27B
37.21
43.43
42.35
17.69
40.00
35.75 ± 4.19
Qwen 3.5 9B
Table 1: Success rate (%) on WebArena Lite, averaged over five evaluation runs.
Table 4Table 5
k
1
2
3
4
5
% trajectories
100.0
25.6
7.5
4.5
0.8
Table 5: Percentage of trajectories retaining environment feedback in memory after k steps.
Figure 4: Error recovery on a Maps task. The environment supports fill but not clear . Both agents issue clear , see the error, and record it in memory. MemAgent’s next overwrite erases the record, and it issues clear repeatedly. AMBER’s appended entry persists, and it never repeats the action. See Figure 7 for the detailed trajectory.
Figure 5: Recursive-compression loss. The task asks for reviewers who mention good fingerprint resistance, and the two matching reviews never share a viewport. Both agents record Rachel after the first scroll. The overwrite agent’s rewrite after it sees T. Gannon drops Rachel, and it answers with T. Gannon only. AMBER’s original entry persists alongside the new one, and it answers correctly with both reviewers.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
System
Evaluation environment
Observation input
Retained info.
Retention
Training
Go-Browse
WebArena
Accessibility tree
Reasoning history
Accumulate
SFT
WebAgent-R1
WebArena
Simplified HTML
Reasoning history
Accumulate
SFT + RL
MEM1
WebShop, multi-hop QA
Custom text
Explicit memory
Recurrent overwrite
RL
MemAgent
Long-document QA
Text segments
Explicit memory
Recurrent overwrite
RL
AMBER (ours)
WebArena
Accessibility tree
Explicit memory
Append-only
SFT + RL
Appendix
Table 6: Representative combinations of evaluation environment, observation input, retained context, and training. Recurrent overwrite memory has so far been validated on single-template or non-interactive settings, and not on diverse web navigation. The table organizes design choices; it does not imply that results are directly comparable across different tasks and environments.
Figure 6: Asynchronous Rollouts: Each rollout is a multi-step episode with each step accessing the environment and inference asynchronously.
Figure 7: Task 33. Both agents issue an invalid clear() at step 4 and record the error at step 5. (a) AMBER keeps that entry in memory through step 13, never issues clear() again, and answers at step 13. (b) MemAgent erases the error at step 6 and issues clear() again at step 12; its reasoning at step 13 notes that clear() is not allowed, but its memory does not record it, and it tries clear() on another field at step 15.
Figure 8: Task 56. Both agents issue an invalid clear() at step 5 and record the error at step 6. AMBER keeps that entry in context, never retries clear() , and answers at step 9. MemAgent retries clear() at step 7 anyway, then overwrites the entry with one that says only that the field was cleared, dropping the error; it issues clear() eight times in total before answering at step 24.
Figure 9: Task 75. (a) AMBER writes the rejected clear() into memory at step 8 and switches to fill() , never issuing clear() again. (b) MemAgent records the rejected clear() as done and, although its reasoning notes the failure, never writes it down; it calls clear() seven times in total.
Figure 10: Recursive-compression loss on task 202. Both agents record that order #000000136 (May 23) is the most recent canceled order. The overwrite agent’s later rewrite keeps only older canceled orders, and it answers with an April order. AMBER’s original entry persists, and it answers correctly. Text is abridged for legibility.
Figure 11: Recall from older memory on task 720 (AMBER). At step 4 the author page shows all three of CameronKelsey’s posts, and m4 records them, including the third (Little White Salmon River). By step 7 the agent has moved to the EarthPorn listing, where that post is not on the page, and the latest entry, m6, no longer names it, yet the step-7 reasoning quotes its title, which can only come from m4 or m5; an overwrite agent would have only m6 in context. The post is not named again until a search surfaces it at step 17, after which the agent upvotes it and succeeds.
Violation
Condition
LOSSY_COMPRESSION
Useful fact or progress dropped from memory
NON_CUMULATIVE_MEMORY
Still-useful earlier memory not carried forward
SUPERSEDED_INFORMATION
Obsolete claim kept alongside its correction
UNSUPPORTED_ACTION_OUTCOME
Outcome asserted before it is observable
FUTURE_PLAN_IN_MEMORY
Near-term plan written into memory
PENDING_ACTION_IN_MEMORY
Pending action written into memory
Appendix
Table 7: Violation taxonomy used by the judge. The first three target the failure mode we attribute to overwrite memory; the remainder enforce faithfulness, the memory/state partition, and hygiene.