Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never established. We characterize this problem through three coupled requirements: (1) Evidence scope; (2) Compositional validity; (3) Admission reliability. We therefore ask whether the interaction history available at write time supports what enters persistent memory. We introduce DerivAudit, a framework for auditing whether a memory is actually supported by the history available when it was written. The audit separates three questions: whether supporting evidence lies beyond writer-provided citations, whether the composed memory introduces unsupported meaning, and how write-time admission decisions affect later memory use. Across two natural memory corpora, audits using broader pre-write history recover support for nearly 60% of memories that appear unsupported from citations alone, while 17-21% remain unsupported after expansion. Yet broader evidence does not by itself make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion alone worsens it on two backbones.
Figures & tables
Figure 1: Persistent memory as derived state. (a) Interaction history is compressed into persistent memory that can later be reused as agent state. (b) This write step has two distinct failure modes: incomplete citations can make a valid memory appear unsupported, while supported historical pieces can be combined into a memory whose full meaning is not supported.
Work
Agent Memory
Evidence Beyond Attached Provenance
Compositional Semantics
Write-time Admission
Downstream Memory Effects
Long-term memory and reliability
LoCoMo / LongMemEval a
✓
–
–
–
–
HaluMem ( Chen et al., 2026 )
✓
–
–
–
✓
Eywa ( Joshi, 2026 )
✓
–
–
✓
–
ConsistencyGate ( Zhang and Li, 2026 )
✓
–
–
✓
–
MemTxn ( Cui et al., 2026 )
✓
–
–
✓
–
Table 1: Comparison with closely related work. DerivAudit studies whether interaction history supports what is written into persistent memory, including evidence beyond attached provenance, compositional meaning, and the consequences of admission decisions.
Figure 2: DerivAudit audits memory derivation from pre-write history to later use. The framework organizes the analysis around three questions: where a memory’s historical support resides, whether that support licenses the full meaning introduced during composition, and how alternative admission decisions affect the memory available to later tasks.
Model
Verification Setting
History Expansion
Composition Check
Unsupported Admission ↓
Valid Retention ↑
Repaired Retention ↑
Qwen
Citation-only
–
–
0.846
0.904
0.821
+ Expanded history
✓
–
0.808
0.901
0.938
+ Obligation-aware check
✓
✓
0.769
0.907
0.929
Gemma †
Citation-only
–
–
0.667
0.837
0.580
+ Expanded history
✓
–
0.769
0.971
0.955
+ Obligation-aware check
✓
✓
0.590
0.939
0.920
Table 2: Memory admission under three write-time verification settings on 391 decidable natural writes. Expanded history adds evidence from pre-write interactions beyond writer-provided citations; composition-aware checking additionally verifies meaning introduced when information is composed into memory. Lower unsupported admission and higher retention are better. A no-gate policy accepts every write and is omitted.
Model
View Type
Verification View
Unsupported Detection U↑
Multi-span Retention ↑
Valid Rejection ↓
Qwen
Established
Holistic judge ( Luo et al., 2023 )
0.608
0.846
0.088
Atomic claims ( Min et al., 2023 )
0.165
0.744
0.121
Predicate–argument QA ( Cattan et al., 2025 )
0.392
0.564
0.115
Composition-aware
Compact relations
0.595
0.949
0.082
Composition graph
0.582
0.872
0.093
Per-obligation
0.291
0.897
0.093
Table 3: Compositional verification on the 512-record controlled suite. Established and composition-aware verification views are evaluated on the same candidate memories and evidence in a shared scoring harness. U denotes unsupported-relation detection, and Multi denotes retention of valid memories requiring joint support from multiple passages.
Figure 3: Evidence scope and compositional verification. (a) Historical evidence recovers support beyond writer-supplied citations. (b–c) Distributed plausible premises reduce valid–invalid separation and increase acceptance of unsupported memory compositions relative to matched irrelevant history.
Figure 4: From persistent memory to downstream behavior. (a) Semantic qualifications drift during repeated consolidation. (b) Distorted and missing memories both increase downstream failure. (c) Natural retrieval links invalid memory to misinformation and valid memory to correct answers.
Figure 5: A natural failure of memory derivation. The writer compresses future-oriented plans into an ongoing activity and combines separately supported facts into a relation that the pre-write history does not establish.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Representation
Unsupported Detection ↑
Multi-span Retention ↑
Observation
Qwen
Raw memory
0.608
0.846
holistic baseline
Correct composition graph
0.468
0.923
no stable gain
Isomorphic graph
0.354
0.974
rendering-sensitive
Shuffled graph
0.430
0.846
graph form alone does not help
Length-matched filler
0.544
0.846
length alone does not explain effect
Compact relations
0.595
0.949
strong retention / detection balance
Appendix
Table 4: Representation robustness on fixed candidate memories and evidence. Values are unsupported-relation detection / multi-span valid retention.
Oracle upstream inputs
Setting
Model
Full
Atom
Graph
Metric
Gold decomposition + gold evidence
Qwen
0.292
0.555
0.394
unsupported detection
Equal-budget prompt search
Interface
Qwen
Llama
Observation
Default
Tuned
Default
Tuned
Holistic
12.8
46.2
55.4
61.7
interface gap narrows
Appendix
Table 5: Protocol controls. Correct upstream inputs, prompt search, alternative aggregation, and additional verification compute change the operating point but do not provide a universal solution.
Model
Median Δ log-odds k=2
Median Δ log-odds k=4
Records moving down ( k=4 )
Acceptance k=0
Acceptance k=2
Acceptance k=4
Qwen
−5.7
−7.7
98.7%
0.994
0.987
0.981
Gemma †
−43.3
−45.9
100.0%
0.933
0.121
0.026
Llama †
–
−20.0
99.0%
0.974
–
0.355
Appendix
Table 6: Accumulated verification noise. Supported memories receive k additional obligations independently verified as true; the candidate memory is unchanged, so any score decrease measures noise from checking more correct content, not detection of a defect.
Remote-evidence interventions (69 families, five conditions)
Interface
Accept intact ↑
Still accept after support deletion ↓
Wrong flip after redundant deletion ↓
False admit after applicable correction ↓
Holistic
0.986
0.116
0.000
0.014
Obligation-aware
1.000
0.000
0.000
0.000
Predicate–argument QA
1.000
0.000
0.000
0.000
Mechanism control: plausible premises versus token-matched irrelevant content
Interface
Accept with plausible premises
Accept with irrelevant content
Paired Δ log-odds [95% CI]
Appendix
Table 7: Remote-history and mechanism controls (family-paired, Qwen verifier). Top: when the manipulation is inside the verifier’s input, the verifier is competent—deleted support is noticed, redundant-path deletion never flips a decision, and a temporally applicable correction is honoured. Bottom: the same invalid record is accepted when its plausible premises are present but rejected when they are replaced by token-matched irrelevant history, isolating plausible-premise composition bias as the mechanism of the distributed-evidence paradox.
Pre-write history says
Memory says
Mutation
Cited verdict
Expanded verdict
“I’ve been working on this car, doing engine swaps and suspension modifications. Now I’m learning about body modifications.”
“…transforming it with engine swaps, suspension modifications, and body modifications .”
prospective → performed
unsup.
unsup.
“What instrument are you playing?” — “I’m learning how to play the violin now…” / “How long have you been playing the piano again?” — “I’ve been playing for about four months .”
“Tim has been learning to play the violin for about four months …”
duration re-bound across entities
supp.
unsup.
“I’m still just learning how to draw, but I love expressing myself through writing .” / “I’ve been a bit frustrated lately with my new phone .”
“Sam has recently taken up drawing as a new form of self-expression …despite occasional frustration with his progress .”
attribute and cause re-bound
unsup.
unsup.
“I scored a deal to continue collaboration with Frank Ocean!”
“Calvin expressed excitement about his new collaboration with Frank Ocean…”
ongoing → new (refuted)
contra.
contra.
Appendix
Table 8: Additional unedited natural writes exhibiting derivation failures. Each row quotes the pre-write history and the resulting memory; verdicts are the adjudicated labels under the citation-only and expanded-history audits. The second case is the study’s canonical illustration that citation matching and history-level validity are different targets: its cited spans pass the provenance check while the expanded audit finds the meaning unsupported.
Evaluation
Agent-memory question
Scale
Primary comparison
Natural paired audit
Is a memory unsupported, or is writer provenance incomplete?
400 writes
Cited evidence → expanded pre-write evidence
Natural admission replay
Which supported and unsupported memories survive write-time verification?
391 writes
Citation-only / expanded history / composition-aware
Matched distribution
Can a verifier distinguish licensed memory from plausible composition when support is distributed?
70 families × 4
Local/distributed × valid/invalid
Semantic verification
Which representation best preserves relations asserted by a memory?
512 records
Prior paradigms vs. derivation-aware interfaces
Remote-history controls
Are failures caused by ignored history or by plausible but non-licensing premises?
LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.
Yiqi Wang, Jiaqi Liu, Jiaqi Zhang +4
University of Southern Queensland · Southern University of Science and Technology · Jiangsu University +2
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.
Yufeng Li, Shuxin Li, Zhenhua Xu +7
East China Normal University, Shanghai, China · Nanyang Technological University, Singapore · Zhejiang University, Hangzhou, China +3
LLM agents increasingly maintain long-term memory of user facts across sessions. Yet such memory is usually evaluated by aggregating accuracy over question rows or episodes. Because this approach scores question rows independently, even when several questions probe the same fact, it cannot show how that fact behaves as conditions change. We introduce MemTrace, a benchmark whose unit of measurement is the knowledge point: a single typed fact about the user, rather than an individual question. MemTrace probes each fact along three controlled dimensions: memory age, defined by how many sessions ago the fact appeared in the history; question type, covering current state, earlier state, and trajectory of change; and evidence condition, covering present, missing, and contradicted-by-false-premise settings. Evaluating 13 memory-system configurations across four paradigms, we find that similar pooled accuracy hides different failures: recovering a fact's current and earlier states does not imply tracking how it changed, and safe abstention does not imply correcting a false premise. The dominant bottleneck is evidence use, not retrieval: when systems fail, the evidence was retrievable 10 times more often than it was missing. These results suggest that improving long-term memory requires better use of reachable evidence, not simply more storage or retrieval.
Xianxuan Long, Zhikai Chen, Shenglai Zeng +3
Michigan State University · Case Western Reserve University