Persistent memory allows an LLM agent to carry experience across conversations, but it also turns a local reasoning mistake into a durable one. During deliberation, an agent may consider a plan, simulate a tool result, report another speaker's belief, and then reject all of them. If memory retains only the resulting sentences, those once-useful possibilities can later return as facts. The record is neither fabricated nor irrelevant; it has simply been detached from the context in which it was valid. We identify this missing context as \emph{discourse ownership}: the world, branch, or speaker that licenses a proposition. Our first finding is counterintuitive. Language models already carry a causally active signal for ownership, yet conventional memory interfaces discard it when they convert reasoning into records. We introduce CASK (Causally Anchored Scoping Keys), a commit rule that preserves this signal so that shared-world facts enter durable memory while provisional content remains available only within its original scope. Our second finding is that the most obvious way to preserve the signal---storing the discovered internal coordinates---is unreliable because equivalent representations need not keep the same coordinates. CASK instead preserves the stable relations that express ownership. Controlled long-conversation conflicts and tool-agent traces show that this design improves memory admission and prevents provisional content from contaminating later answers while complementing runtime provenance. The resulting commit boundary lets agents explore more possibilities without granting every intermediate sentence authority over future behavior.
Figures & tables
Figure 1 : Reasoning creates propositions that belong to different worlds. The agent considers Paris but rejects that branch, while committing to Sydney. Cask converts the model’s ownership-sensitive relational geometry into an explicit commit edge: Sydney enters actual-world memory, whereas Paris remains available only inside its scoped world. A later query may use committed evidence without turning the discarded branch into a fact.
Swapped state
Δp [95% CI]
Flip (%)
Ownership subspace, opposite scope
0.715[0.656,0.775]
75.8
Complement, opposite scope
0.048[0.030,0.069]
6.6
Random rank-16, norm matched
0.004[0.003,0.005]
0.6
Ownership subspace, same scope
−0.034[−0.046,−0.021]
5.9
Table 1 : Frozen-gate mediator intervention on 3B. Δp is signed toward donor scope; intervals cluster by conversation.
Representation
Qwen2.5-3B
Qwen2.5-1.5B
TF–IDF
50.21
50.21
Basis coordinates
65.64
74.07
Full hidden state
72.12
77.47
Output logits
78.19
62.65
Causal Gram
61.42
71.71
Causal Gram + logits
79.53
70.06
Table 2 : Complete source-isolated attachment comparison (%). All models use the same conversation folds and fold-local alternative sources. Cask uses its default symmetric threshold τ=0.5 and is best across every evaluated representation family on both frozen backbones.
Figure 2 : Controlled evidence for scope-sensitive memory. (a) The Cask advantage is concentrated in scope-only, not payload-only, attachment. Error bars on the random row denote one standard deviation across 20 decompositions. (b) At τ=0.5 , Cask moves toward higher recall and lower contamination relative to full-Gram. Under a development-calibrated 95% actual-admission constraint, output logits define the stronger high-recall boundary. (c) The safety gain reduces adoption of the injected answer, while ordinary reader utility remains a separate trade-off.
Method
Att.
R@5
Cont.
MRR
EM
F1
Alt. ↓
Full Gram+logits
81.28
43.42
7.61
31.08
24.69
35.43
1.03
Cask ( τ=0.5 )
83.33
44.86
7.41
32.35
26.75
37.31
1.03
Flat memory
–
–
–
–
29.63
40.19
4.53
Table 3 : Answer-conflicting memory. Attachment and retrieval/reader metrics are percentages except MRR. Lower is better for contamination and alternative-answer adoption.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Rank 8
Rank 16
Rank 32
Block 18
99.2/.636
100.0/.305
98.4/.312
Block 20
98.4/.788
97.7/.716
99.2/.683
Block 22
100.0/.776
100.0/.801
100.0/.869
Appendix
Table 4 : Development sensitivity. Each cell is attachment accuracy (%) / aligned Δp under the ownership-subspace swap.
Commit policy
BAcc
Actual
Scoped ↑
Output logits
82.29
81.52
83.05
Complement Gram+logits
86.37
85.87
86.87
TF–IDF
90.40
81.52
99.28
Cask fallback
91.10
89.13
93.08
Cask + 50% provenance
95.02
93.48
96.56
Runtime provenance (100%)
100.00
100.00
100.00
Appendix
Table 5 : Generated tool-agent traces (%). BAcc is balanced accuracy; the 50% hybrid averages five metadata-retention masks.
LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.
Yiqi Wang, Jiaqi Liu, Jiaqi Zhang +4
University of Southern Queensland · Southern University of Science and Technology · Jiangsu University +2
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.
Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions. Existing memory systems often store observations as isolated records, summaries, or indexed fragments, which makes evidence aggregation, fact revision, and memory maintenance difficult. We propose Infini Memory, a maintainable text-based persistent memory architecture that treats agent memory as topic-structured documents. Each topic document serves as a semantic unit for collecting related evidence, preserving metadata, and revising facts over time. New observations are first staged in a buffer and periodically consolidated into coherent textual contexts. At inference time, an agentic retrieval procedure lets the LLM read memory through iterative tool calls rather than a single retrieval step. On MemoryAgentBench, Infini Memory achieves 64.7% overall score. Ablations show that topic-structured maintenance and iterative evidence inspection improve complementary aspects of long-term memory use.
Suozhao Ji, Baodong Wu, Zehao Wang +8
1Infinigence AI · 2Tsinghua University · 3Shanghai Jiaotong University