Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
Figures & tables
Method
Mean traj. ↑
Final checkpoint ↑
BWT ↑
Accepted
Harmful
Static
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0.000 [0.000, 0.000]
0.00
0.00
Frozen-compute
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0.000 [0.000, 0.000]
0.00
0.00
Latest-only
0.703 [0.678, 0.725]
0.720 [0.681, 0.755]
-0.253 [-0.300, -0.204]
12.00
6.25
Self-judge
0.736 [0.696, 0.774]
0.722 [0.641, 0.802]
-0.140 [-0.220, -0.059]
3.38
1.75
Replay
0.707 [0.676, 0.738]
0.681 [0.621, 0.745]
-0.240 [-0.297, -0.191]
7.88
3.88
ORC
0.713 [0.653, 0.761]
0.707 [0.645, 0.760]
0.000 [0.000, 0.000]
1.00
0.75
Table 1: Main ProcStream-RSI results over eight paired streams. Brackets are seed-bootstrap 95% intervals. Harmful accepts count accepted rounds whose fixed hidden checkpoint score decreased.
Contrast
Mean difference [bootstrap 95%]
p
Holm p
ORC - Replay mean trajectory
0.006 [-0.045, 0.048]
0.8438
1.0000
ORC - Latest-only mean trajectory
0.010 [-0.036, 0.051]
0.7188
1.0000
ORC - Batch-ORC final
-0.068 [-0.123, -0.004]
0.0938
0.2812
Table 2: Prespecified paired contrasts. Intervals are descriptive percentile seed-bootstrap 95% intervals. Paired sign-flip sensitivity p -values require symmetry of seed-level differences; Holm adjustment covers the three rows.
Method
HumanEval+ pass rate
GPT-OSS hidden pass rate
Static
0.914 [0.898, 0.930]
0.941 [0.907, 0.974]
Latest-only
0.918 [0.891, 0.941]
0.911 [0.886, 0.937]
Replay
0.922 [0.895, 0.945]
0.916 [0.878, 0.951]
ORC
0.895 [0.867, 0.922]
0.951 [0.926, 0.968]
Batch-ORC
0.914 [0.898, 0.930]
0.941 [0.907, 0.974]
Scoped retrieval †
—
0.960 [0.930, 0.987]
Table 3: Terminal-policy transfer. Values are means with seed-bootstrap 95% intervals. HumanEval+ uses the fixed 32-task sample; GPT-OSS uses each stream’s 18-task sealed final bank. † Family-scoped retrieval over the same archived completions.
Deployment
Mean traj. ↑
Final checkpoint ↑
Accepted
Harmful
Static
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0
0
Global ORC
0.713 [0.653, 0.762]
0.707 [0.645, 0.760]
8
6
Scoped retrieval
0.816 [0.782, 0.852]
0.819 [0.786, 0.855]
8
0
Table 4: Same accepted skills, different retrieval scope on the eight main streams. Scoped retrieval applies each ORC rule only to its originating family and uses the initial skill elsewhere.
Deployment
Accepted
Harmful
Checkpoint t=1
Non-current Δ
Global ORC
5
3
0.761 [0.709, 0.802]
-0.047 [-0.120, 0.005]
Scoped ORC
5
0
0.802 [0.779, 0.826]
0.000 [0.000, 0.000]
Table 5: Balanced randomized-entry replication (18 streams; two per first family). Brackets show stream-bootstrap 95% intervals.
Method
Trajectory
Final checkpoint
Accepts/stream
Streams ≥2
Harmful/accepted
Global- ORC
0.785 [0.757, 0.812]
0.789 [0.758, 0.817]
0.44
2/27
6/12
Scoped- ORC
0.848 [0.825, 0.871]
0.900 [0.871, 0.927]
2.33
19/27
0/63
Table 6: Full 12-round randomized-order extension (27 paired streams). Brackets are stream-bootstrap 95% intervals; harmful updates reduce the next hidden checkpoint.
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.
Large language model reasoning leaves no trace once it is done. The steps of a chain of thought disappear when the context window closes, a pruned search branch is just gone, and memory buffers cannot be diffed, merged, or audited. Code, infrastructure, and experiments are all version-controlled. Reasoning is not. GitOfThoughts stores an agent's reasoning tree as a git repository. Every scored thought becomes a commit, scores become notes, outcomes become tags, and retrieval is just git log over the agent's own history. We use this to test something simple. Does giving an agent memory from past problems actually make it more accurate? We tried five memory stores (none, a markdown file, a vector database, a graph, and git) across two benchmarks, two model sizes, and several pre-registered repeat experiments. The answer, on new problems, is no, including one promising early result that did not hold up when we repeated it. Memory only helps once the problem being solved is nearly identical to something already in memory (cosine similarity above about 0.8); below that, it does nothing. In other words, the model is finding the answer rather than learning the method. Even a model 4.5x larger still cannot pull a reusable method out of a worked example; it just gets better at spotting near-copies. The only thing that reliably helped on new problems was generating several answers and picking the most common one (self-consistency). So the case for using git as the memory store is not that it retrieves better. It is that it gives auditability, history, and the ability to merge two agents' memories, at no cost to accuracy.
Self-improving personal agents now write profiles, memories, and reusable skills that carry over from one chat to the next. Prior work asks whether user pressure bends a model's next answer. Yet these agents can also write the user's claim down, so a later chat may read it back as trusted context. We call this persistent sycophancy. We introduce the Personal Agent Sycophancy Benchmark, PASB, with 1,600 tasks run on two real agents, Hermes-Agent and OpenClaw, across twelve models. Each task isolates a first chat containing the claim from a neutral follow-up chat, so any carryover must pass through a note the agent chose to write. Our analysis shows that downstream failure, meaning how often later answers side with the claim, reason from it, treat it as fact, or stretch it, reaches 71.9% when the follow-up chat can read a saved claim, against 45.0% when the claim stays in the first chat. Writing also edits the claim, as agents save it as a stable preference, a background fact, or a reusable procedure in 51.4% of runs. A saved claim still shapes answers in a different domain. Among the mitigations we test, explicit memory editing helps most, cutting downstream failure to 32.7% for Hermes-Agent and 54.5% for OpenClaw on same-domain follow-ups. PASB shows that self-improving agents must govern what they write down before it governs what they say. Our benchmark is available at https://github.com/henrymao2004/agent-sycophancy.
Xutao Mao, Liangjie Zhao, Leyao Wang +6
City University of Hong Kong · Adelaide University · Yale University +3