When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory
Organizations: Meta AI
Abstract
Persistent agent memory is only as reliable as its retention decision: an assertion weakly supported by its source can be stored and later reused as established fact. We study whether the retention decision should be governed by a confidence bar conditioned on the semantic category of the assertion rather than by a single global threshold, retaining well-evidenced categories liberally while abstaining more aggressively where inference is unreliable. We evaluate this in a deployed cold-start memory pipeline on 100 synthetic personas. The empirical evaluation is motivated by a sharp reliability asymmetry: across 4{,}715 candidate assertions, only 77.9% of value and belief assertions are supported by their source, versus 96.2% for all other categories. A global confidence threshold cannot separate these: it either admits unsupported value claims or discards well-evidenced ones. Conditioning the threshold on category resolves the tradeoff. In repeated held-out evaluation, a stricter bar on values alone reduces unsupported retentions from 6.2% to 4.0% (an relative reduction, modest but consistent across folds) and, as corroborating evidence, preserves an estimated 13 percentage points more coverage (95% CI 9.8--16.0) than a global threshold at comparable retention. Our results suggest that reliable retention depends on the type of assertion, not on confidence alone, and that a category-conditioned threshold can act as a simple, effective form of selective prediction at the write boundary.
Figures & tables
| Policy | Faith. (%) | Est. cov. (%) | Retained (%) |
|---|---|---|---|
| Write all | 92.24 | 85.27 | 100.0 |
| Global | |||
| Values-targeted | |||
| Targeted global | [ , ] | [ , ] | [ , ] |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Judge configuration | All | Values | Non-values |
|---|---|---|---|
| Gemini 2.5 Pro ( ) | 0.935 [0.914, 0.954] | 0.840 [0.768, 0.898] | 0.959 [0.939, 0.977] |
| Gemini 2.5 Pro ( ) | 0.928 [0.904, 0.951] | 0.845 [0.773, 0.908] | 0.949 [0.927, 0.968] |
| GPT-5 ( ) | 0.894 [0.874, 0.915] | 0.624 [0.540, 0.712] | 0.963 [0.952, 0.974] |