cs.AIOct 5, 2026

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

Authors: Zi Wang, Xingqiao Wang, Emmanuel Addai, Devika Ambekar, Xiaowei Xu

Organizations: University of Arkansas at Little Rock

Abstract

Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

    Oct 7, 2026Yufeng Li, Shuxin Li, Zhenhua Xu +7

  2. MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

    Jun 15, 2026Xianxuan Long, Zhikai Chen, Shenglai Zeng +3Long-Term MemoryTraces

  3. Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

    Jul 27, 2026Ruizhe Li, Mingxuan Du, Benfeng Xu +1Agentic MemoryPrecision Recall