Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents
Organizations: New York University
Abstract
Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never established. We characterize this problem through three coupled requirements: (1) Evidence scope; (2) Compositional validity; (3) Admission reliability. We therefore ask whether the interaction history available at write time supports what enters persistent memory. We introduce DerivAudit, a framework for auditing whether a memory is actually supported by the history available when it was written. The audit separates three questions: whether supporting evidence lies beyond writer-provided citations, whether the composed memory introduces unsupported meaning, and how write-time admission decisions affect later memory use. Across two natural memory corpora, audits using broader pre-write history recover support for nearly 60% of memories that appear unsupported from citations alone, while 17-21% remain unsupported after expansion. Yet broader evidence does not by itself make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion alone worsens it on two backbones.
Figures & tables
| Work | Agent Memory | Evidence Beyond Attached Provenance | Compositional Semantics | Write-time Admission | Downstream Memory Effects |
| Long-term memory and reliability | |||||
| LoCoMo / LongMemEval a | – | – | – | – | |
| HaluMem ( Chen et al., 2026 ) | – | – | – | ||
| Eywa ( Joshi, 2026 ) | – | – | – | ||
| ConsistencyGate ( Zhang and Li, 2026 ) | – | – | – | ||
| MemTxn ( Cui et al., 2026 ) | – | – | – | ||
| Model | Verification Setting | History Expansion | Composition Check | Unsupported Admission | Valid Retention | Repaired Retention |
| Qwen | Citation-only | – | – | 0.846 | 0.904 | 0.821 |
| + Expanded history | ✓ | – | 0.808 | 0.901 | 0.938 | |
| + Obligation-aware check | ✓ | ✓ | 0.769 | 0.907 | 0.929 | |
| Gemma † | Citation-only | – | – | 0.667 | 0.837 | 0.580 |
| + Expanded history | ✓ | – | 0.769 | 0.971 | 0.955 | |
| + Obligation-aware check | ✓ | ✓ | 0.590 | 0.939 | 0.920 |
| Model | View Type | Verification View | Unsupported Detection | Multi-span Retention | Valid Rejection |
| Qwen | Established | Holistic judge ( Luo et al., 2023 ) | 0.608 | 0.846 | 0.088 |
| Atomic claims ( Min et al., 2023 ) | 0.165 | 0.744 | 0.121 | ||
| Predicate–argument QA ( Cattan et al., 2025 ) | 0.392 | 0.564 | 0.115 | ||
| Composition-aware | Compact relations | 0.595 | 0.949 | 0.082 | |
| Composition graph | 0.582 | 0.872 | 0.093 | ||
| Per-obligation | 0.291 | 0.897 | 0.093 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Representation | Unsupported Detection | Multi-span Retention | Observation |
| Qwen | Raw memory | 0.608 | 0.846 | holistic baseline |
| Correct composition graph | 0.468 | 0.923 | no stable gain | |
| Isomorphic graph | 0.354 | 0.974 | rendering-sensitive | |
| Shuffled graph | 0.430 | 0.846 | graph form alone does not help | |
| Length-matched filler | 0.544 | 0.846 | length alone does not explain effect | |
| Compact relations | 0.595 | 0.949 | strong retention / detection balance |
| Oracle upstream inputs | |||||
| Setting | Model | Full | Atom | Graph | Metric |
| Gold decomposition + gold evidence | Qwen | 0.292 | 0.555 | 0.394 | unsupported detection |
| Equal-budget prompt search | |||||
| Interface | Qwen | Llama | Observation | ||
| Default | Tuned | Default | Tuned | ||
| Holistic | 12.8 | 46.2 | 55.4 | 61.7 | interface gap narrows |
| Model | Median log-odds | Median log-odds | Records moving down ( ) | Acceptance | Acceptance | Acceptance |
| Qwen | 98.7% | 0.994 | 0.987 | 0.981 | ||
| Gemma † | 100.0% | 0.933 | 0.121 | 0.026 | ||
| Llama † | – | 99.0% | 0.974 | – | 0.355 |
| Remote-evidence interventions (69 families, five conditions) | ||||
| Interface | Accept intact | Still accept after support deletion | Wrong flip after redundant deletion | False admit after applicable correction |
| Holistic | 0.986 | 0.116 | 0.000 | 0.014 |
| Obligation-aware | 1.000 | 0.000 | 0.000 | 0.000 |
| Predicate–argument QA | 1.000 | 0.000 | 0.000 | 0.000 |
| Mechanism control: plausible premises versus token-matched irrelevant content | ||||
| Interface | Accept with plausible premises | Accept with irrelevant content | Paired log-odds [95% CI] | |
| Pre-write history says | Memory says | Mutation | Cited verdict | Expanded verdict |
| “I’ve been working on this car, doing engine swaps and suspension modifications. Now I’m learning about body modifications.” | “…transforming it with engine swaps, suspension modifications, and body modifications .” | prospective performed | unsup. | unsup. |
| “What instrument are you playing?” — “I’m learning how to play the violin now…” / “How long have you been playing the piano again?” — “I’ve been playing for about four months .” | “Tim has been learning to play the violin for about four months …” | duration re-bound across entities | supp. | unsup. |
| “I’m still just learning how to draw, but I love expressing myself through writing .” / “I’ve been a bit frustrated lately with my new phone .” | “Sam has recently taken up drawing as a new form of self-expression …despite occasional frustration with his progress .” | attribute and cause re-bound | unsup. | unsup. |
| “I scored a deal to continue collaboration with Frank Ocean!” | “Calvin expressed excitement about his new collaboration with Frank Ocean…” | ongoing new (refuted) | contra. | contra. |
| Evaluation | Agent-memory question | Scale | Primary comparison |
| Natural paired audit | Is a memory unsupported, or is writer provenance incomplete? | 400 writes | Cited evidence expanded pre-write evidence |
| Natural admission replay | Which supported and unsupported memories survive write-time verification? | 391 writes | Citation-only / expanded history / composition-aware |
| Matched distribution | Can a verifier distinguish licensed memory from plausible composition when support is distributed? | 70 families 4 | Local/distributed valid/invalid |
| Semantic verification | Which representation best preserves relations asserted by a memory? | 512 records | Prior paradigms vs. derivation-aware interfaces |
| Remote-history controls | Are failures caused by ignored history or by plausible but non-licensing premises? | Matched interventions | Necessary-path removal / redundant-path removal / correction |
| Oracle diagnostics | Do errors remain with correct evidence and decomposition? | 235 records | Gold evidence + gold decomposition |