Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Organizations: Department of Computer Science Illinois Institute of Technology Chicago, IL
Abstract
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.
Figures & tables
| Design | AUC | Lift | Gold loss |
|---|---|---|---|
| Store-level rand. | |||
| Exposure (CMP) |
| Support | AUC gap | Median | |
|---|---|---|---|
| exposure only | |||
| Variant | AUC | Destroyed | Gold loss |
|---|---|---|---|
| CMP (full) | |||
| – one-sided abstention | |||
| – exposure design | |||
| – both | |||
| two-sided abstention | |||
| two-sided exposure |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Description |
|---|---|
| Structural causal model (Def. 1 ) | |
| Memory state at time | |
| Memory state excluding , | |
| Universe of possible memories | |
| Retrieved subset at time | |
| Memory operation at time |
| Analysis | Benchmark | Draws/mem. | ||||
|---|---|---|---|---|---|---|
| Identification, single query | LongMemEval | |||||
| Two-sided exposure | LongMemEval | |||||
| Identification, pooled | LoCoMo | |||||
| Identification, per query | LoCoMo |
| Nomination rule | Signal used | Lift over chance |
| Lowest access frequency | Access counts | – |
| Lowest cosine similarity | Query similarity | |
| Recency-weighted blend | Both | – |
| Newest-first | Insertion order | – |
| Utility, store-level estimates | Estimated | |
| Utility, exposure estimates | Estimated |
| Benchmark | Pairs | Store | AUC store | AUC exposure | |
|---|---|---|---|---|---|
| LongMemEval (independent) | |||||
| LongMemEval (all suspect) | — | ||||
| LoCoMo (pooled estimand) | |||||
| LoCoMo (per-query estimand) | — | ||||
| Mem0 | [ , ] | — | — |
| Population | Exposure | Cosine | Diff. | |||
|---|---|---|---|---|---|---|
| Independent | ||||||
| All suspect | ||||||
| All |
| Variant | AUC | Evicted | Destroyed | Gold loss |
|---|---|---|---|---|
| Exposure one-sided gate | ||||
| Exposure, ungated | ||||
| Store-level one-sided gate | ||||
| Store-level, ungated | ||||
| Exposure two-sided gate | ||||
| Store-level two-sided gate |
| Selector | Type | Headroom recovered |
|---|---|---|
| Five bi-encoders (MiniLM to gemini-embedding-2) | Pointwise | at chance on contradiction pairs |
| Cross-encoder reranker (BGE-reranker-v2-m3) | Pointwise | LongMemEval, – HotpotQA, – MuSiQue |
| Maximal marginal relevance | Set-aware | to |
| Coverage-greedy (ceiling) | Set-aware |
| Predictor | Deletion | Insertion |
|---|---|---|
| Pooled estimate | ||
| Nearest estimation query | ||
| Similarity-weighted mean | ||
| Maximum over estimation queries | ||
| Ground-truth evidence labels (control) | — |