The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory
Organizations: University of Arkansas at Little Rock
Abstract
Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.
Figures & tables
| Question | Population and comparison | Primary outcomes |
|---|---|---|
| RQ1: Support | RHELM/MemOps queries; global, namespace, random, pre/post-filter, routed, and released-field arms; metadata stress | Recall, feasibility, bounds, coverage, constrained loss, typed exposure, work |
| RQ2: Verification | 96 public-development queries with fixed top-20 pools plus controlled scenarios; fields versus text decisions | ROC-AUC, violation precision/recall, anchor false denial, consistency, overflip |
| RQ3: Enforcement | paired benchmark-native reader cases plus controlled pairs per reader; no reader pooling | Answer correctness/quality, non-answer action, protected/stale disclosure, cell effects, selectivity |
| Method | Visible information | Recall | Feasible | Similarity evals. | ||
|---|---|---|---|---|---|---|
| Global dense | text | .432 | .237 | [.187,.195] | .798 | 90,122 |
| Global recency+dense | text + released order | .452 | .256 | [.198,.206] | .784 | 90,122 |
| Namespace dense | text + trusted namespace | .533 | .311 | [.118,.126] | .713 | 1,551 |
| Namespace global | +.101 | +.074 | — | .085 | 58.1 fewer | |
| Support intervention | Recall | Feasible | Similarity evals. | |
|---|---|---|---|---|
| Global dense | .432 | .237 | .798 | 90,122 |
| Random same-size | .061 | .024 | .978 | 1,540 |
| Namespace pre-filter | .533 | .311 | .713 | 1,551 |
| Global post-filter ( ) | .531 | .311 | .713 | 90,122 |
| Anchor-preserving same-size oracle | .840 | .747 | .390 | 1,551 |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Support | Known non-usable | Coverage | Lower | Upper |
|---|---|---|---|---|
| Global dense | 0.575 | 0.538 | 0.308 | 0.770 |
| Namespace dense | 0.348 | 0.355 | 0.134 | 0.778 |
| Axis family | Route | Recall | Feasible | Any violation | ||
|---|---|---|---|---|---|---|
| Scope + policy + lifecycle | Global | .432 | .237 | [.187,.195] | .798 | .528 |
| Scope + policy + lifecycle | Namespace | .533 | .311 | [.118,.126] | .713 | .396 |
| Scope + lifecycle | Global | .432 | .237 | [.102,.111] | .787 | .356 |
| Scope + lifecycle | Namespace | .533 | .311 | [.017,.025] | .693 | .134 |
| Route | Known risk | Coverage | Any violation | |
|---|---|---|---|---|
| Global dense | .1893 | .9918 | [.1873,.1955] | .5285 |
| Namespace dense | .1122 | .9913 | [.1105,.1191] | .3721 |
| Axis | Exact agreement | Uncertain | Evaluable | Reviewer–source | |
|---|---|---|---|---|---|
| Relevance | 0.8937 (185/207) | 0 | 207 | 0.8135 | .797–.802 |
| Scope | 0.9952 (206/207) | 0 | 207 | 0.9808 | .995–1.000 |
| Lifecycle state | 0.8599 (178/207) | 23 | 184 | 0.6703 | .841–.947 |
| Prohibited evidence | 0.8841 (183/207) | 0 | 207 | 0.0829 | .729–.826 |
| Legacy usable composite (derived) | 0.9710 (201/207) | 0 | 207 | 0.9405 | .836–.855 |
| Work | Invalidity target | Primary unit / stages | Matched-recall bounds | Paired exposure estimand |
|---|---|---|---|---|
| GateMem [ Ren et al., 2026 ] | Principal access, updates, forgetting | Episode/checkpoint utility and leakage | Not primary | Not primary |
| MemOps [ Hao et al., 2026 ] | Operation target, scope, state transition | Structured operation trace and probes | Not primary | Not primary |
| STALE [ Chao et al., 2026 ] | Implicit invalidation of earlier state | State-resolution and answer probes | Not primary | Not primary |
| A-TMA [ Shi et al., 2026 ] | Current, historical, and transition state | Bank, retrieval, and answer diagnostics | Not primary | Not primary |
| MemConflict [ Tao et al., 2026 ] | Temporal, factual, and contextual conflicts | Retrieval, ranking, and final answers | Not primary | Not primary |
| MemGate [ Zhang et al., 2026b ] | Query-conditioned contextual appropriateness | Learned admission gate, threat, and utility | Not primary | Not primary |
| Model | ROC-AUC | Violation precision | Violation recall | Anchor false deny |
|---|---|---|---|---|
| GPT-5.6 Sol | ||||
| Gemini 3.6 Flash |
| Reader | answer accuracy | answer quality | non-answer rate |
|---|---|---|---|
| DeepSeek V4 Pro | +0.053 [+0.020, +0.084] | +0.043 [+0.023, +0.062] | -0.048 [-0.068, -0.031] |
| Gemini 3.6 Flash | +0.068 [+0.047, +0.090] | +0.050 [+0.036, +0.065] | -0.041 [-0.057, -0.025] |
| GPT-5.6 Luna † | +0.066 [+0.039, +0.096] | +0.048 [+0.031, +0.066] | -0.038 [-0.059, -0.015] |
| Reader | Source | answer accuracy | non-answer rate |
|---|---|---|---|
| DeepSeek V4 Pro | RHELM | +.021 [ .037,+.075] | .044 [ .075, .016] |
| MemOps | +.084 [+.052,+.116] | .053 [ .077, .030] | |
| Gemini 3.6 Flash | RHELM | +.059 [+.032,+.094] | .042 [ .068, .022] |
| MemOps | +.076 [+.045,+.106] | .039 [ .061, .018] | |
| GPT-5.6 Luna † | RHELM | +.046 [+.004,+.094] | .059 [ .096, .016] |
| MemOps | +.086 [+.053,+.121] | .017 [ .034, .001] |
| Reader | answer accuracy | answer quality | non-answer rate |
|---|---|---|---|
| DeepSeek V4 Pro | +.062 [+.033,+.092] | +.050 [+.033,+.068] | .050 [ .069, .032] |
| Gemini 3.6 Flash | +.070 [+.048,+.094] | +.052 [+.037,+.067] | .040 [ .057, .024] |
| GPT-5.6 Luna † | +.072 [+.045,+.101] | +.053 [+.036,+.071] | .032 [ .049, .014] |
| Arm | Selected setting |
|---|---|
| Global BM25 | |
| Global dense | Parameter-free |
| Global BM25+dense RRF | |
| Global recency dense | |
| Namespace dense | Parameter-free |
| Query-agnostic current-only | Remove stale and superseded records for every query |
| Channel | Reference | Last dominant | First failure |
|---|---|---|---|
| Namespace false deny | Global dense | .10 | .20 |
| Namespace missing | Global dense | .10 | .20 |
| Namespace source swap | Global dense | .10 | .20 |
| Namespace false allow | Global dense | .50 | Not observed |
| Policy false deny | Clean namespace | .00 | .02 |
| Policy false allow | Clean namespace | .30 | .40 |
| Arm | Recall | Feasible | Penalized non-usable upper risk | Candidates |
|---|---|---|---|---|
| Global BM25 | 0.448 | 0.284 | 0.913 | 90122 |
| Global dense | 0.717 | 0.539 | 0.873 | 90122 |
| Global BM25+dense RRF | 0.691 | 0.503 | 0.883 | 90122 |
| Global recency dense | 0.728 | 0.550 | 0.864 | 90122 |
| Namespace dense | 0.866 | 0.762 | 0.837 | 1551 |
| Query-agnostic current-only | 0.824 | 0.672 | 0.845 | 1549 |
| Support | Evidence recall | Wrong-scope leakage | Non-usable fraction |
|---|---|---|---|
| Global dense | 1.0 | 0.025 | 0.1189 |
| Namespace dense | 1.0 | 0.000 | 0.0824 |
| Layer | Status | Relation to earlier evidence |
|---|---|---|
| Retrieval setting selection | Prespecified development procedure | Selected before evaluation with historical loss; no test retuning. |
| Frozen v1 evaluation | Held-out evaluation of the frozen v1 route family | Preserves the original non-usable status semantics and all nine arms. |
| Admissibility v2 score | Post-hoc correction; frozen routes | Replaces by in scoring only; no reranking or retuning. |
| Released-governance v2 | Post-hoc untuned attribution; frozen routes | Applies policy independently of intent and reports lifecycle separately. |
| Two-reader analysis | Prespecified paired execution | Cases, routes, prompts, and two readers frozen before outcome scoring. |
| GPT-5.6 Luna reader | Gated sequential replication | Repeats the three primary routes after the two-reader continuation criterion. |