Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We implement two inference-time designs without retraining: factor-compiled admission (FC), which assesses whole memory entries, and permission-semantic admission (PS), which decomposes entries into typed units; both translate adjudicated attributes into eligibility decisions via deterministic policies. We evaluate on a four-backbone development suite and an external benchmark with four tasks of 300 samples each. Relative to verbatim injection, FC and PS reduce pooled judge-assessed failure rates on the external benchmark by 6.7 and 8.8 percentage points (p = 2.7e-7 and 4.1e-12), and development-set cross-domain leakage falls by up to 29.5 percentage points. A query-conditioned gating baseline shows no significant change in objective-fact failure or pooled failure. Under matched admission budgets, PS outperforms random and relevance-based selection on external objective-fact judgment after Holm correction. Holding presentation fixed, tightening admission cuts cross-domain failure by a further 17.5 percentage points (p = 1.6e-4); in contrast, no comparison between two renderings of identical adjudicated outputs survives multiple-comparison correction. Both designs increase personalization failures, and PS misses the preregistered improvement and personalization-preservation criteria. These results support evaluating admission and presentation separately: selection quality provides task-specific safety gains, while preserving beneficial memory use remains unresolved.
Figures & tables
Figure 1. Governance decomposed into two levers over one adjudication. Admission (upper half): a single batched adjudication assigns attributes to retrieved memories in situ; two admission designs (FC, PS) compile the adjudicated attributes to decide what enters the context, and excluded material is physically absent. Presentation (lower half): the same admitted verdicts are rendered under either sectioned form (PN vs. CH) —the verdict-shared comparison of Section 5.3 varies only this stage.
Backbone
Dim
Bare
PS+PN
p
PS+CH
p
MiniMax-M3
syc
31.1
27.0
.40
29.5
.83
MiniMax-M3
ben
16.5
18.7
.82
16.5
1.0
MiniMax-M3
xd
9.5
4.0
.052
4.0
.027
Qwen3.7-Max
syc
50.0
36.1
.0009
41.8
.053
Qwen3.7-Max
ben
0.0
0.0
1.0
0.0
1.0
Qwen3.7-Max
xd
44.0
14.5
<10−4
18.0
<10−4
Table 1. Development surface: failure rate (%) under verbatim injection (Bare) vs. PS+PN and PS+CH; paired exact McNemar, raw p . Dimensions: syc = sycophancy, ben = beneficial use, xd = cross-domain leakage. Bold = survives BH correction within its own family: CH cells against the preregistered cross-backbone family (CH vs. Bare on syc/xd, four backbones); PN cells against a separate exploratory family.
Task
Bare
Gate
p
FC
p
PS
p
ofj
45.2
45.5
1.0
23.7
1.4×10−13
20.4
7.6×10−19
csc
33.3
22.7
.0004
27.0
.042
25.3
.0049
mec
91.7
85.7
.0003
85.0
8.8×10−5
84.0
3.4×10−5
pm
14.7
22.7
.0012 †
22.3
.0022 †
20.0
.0365 †
pooled ( n=1,199 )
46.2
44.1
.08
39.5
2.7×10−7
37.4
4.1×10−12
Table 2. External benchmark: judged failure rate (%) by task, n=300 per cell (one ofj cell 299 after pairwise exclusion). Paired exact McNemar vs. Bare. Gate = our reimplementation of the query-conditioned gating policy of MemGate ( Zhang et al., 2026a ) on the same substrate (memories admitted by query-conditioned relevance, no adjudication). pm failure = appropriate stored preference not applied (lower is better). † marks a failure-rate increase relative to Bare (Section 5.5 ). Task abbreviations: ofj = objective fact judgment, csc = contextual scope control, mec = memory–evidence conflict, pm = personalized memory use.
Task
Metric
PN
CH
p
ofj
accuracy
.737
.777
.043
ofj
misconception
.220
.190
.12
csc
accuracy
.170
.160
.76
csc
scope violation
.253
.237
.60
mec
misled
.833
.823
.68
mec
accuracy
.003
.007
1.0
Table 3. Verdict-shared rendering experiment: identical admission decisions, two sectioned renderings (PN vs. CH). Paired exact McNemar; n=300 per row.
Task
NoMem
PS+CH
pm: accuracy / pref. used
1.0 / 2.0
60.3 / 80.7
csc: accuracy / violation
0.0 / 0.0
16.0 / 23.7
mec: misled
0.0
82.3
Table 4. No-memory ablation compared with the PS+CH arm. Values are percentages. Removing memory sharply reduces personalization accuracy and preference use, while yielding zero observed scope violations and misleading responses; this ablation removes both beneficial and conflicting content (Section 5.4 ).