Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores (Δ) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6% to 18.3% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.
Figures & tables
Figure 1: User preferences remain valid across situations, but their applicability changes with context. The same request may call for applying a preference in one situation and setting it aside in another. PairPref evaluates whether assistants make this distinction across paired situations.
Benchmark
Main evaluation target
Preference following
Preference suppression
Selection
Generation
Long-term memory
LoCoMo
Conversational recall
—
—
—
—
LongMemEval
Long-term memory
—
—
—
Preference following
PrefEval
Preference following
—
—
—
PersonaMem
Evolving user profiles
Table 1: Capabilities assessed by preference and memory benchmarks.
Figure 2: PairPref construction and evaluation. Controlled pairs pass three validation stages and support selection and free generation on the same situations.
Selection
Free generation
Model
Use + ↑
Use - ↓
Selectivity ↑
Pair success ↑
Retained pairs
Use both ↓
Pair success ↑
Claude Opus 5
76.4
11.5
64.9 [-1pt] [58.2–70.6]
66.4
1,067
88.7
9.3
Kimi K3
76.0
18.3
57.6 [-1pt] [50.8–63.6]
59.5
1,091
83.6
10.6
GPT-6 Astra
71.2
14.7
56.6 [-1pt] [49.8–62.6]
57.9
1,068
71.0
17.9
GLM-5.3
75.1
18.4
56.6 [-1pt] [50.3–61.9]
58.4
1,056
90.8
5.7
Gemini 3.5 Flash
78.9
22.9
56.0 [-1pt] [49.1–62.5]
57.9
1,065
75.1
18.3
Table 2: Main results on PairPref.
Figure 3: Multiple-choice Δ versus free-answer always-use on the same pairs (Table 2 ). Contextual separation in selection coexists with frequent preference use in both generated answers.
A. Selection: selectivity
B. Generation: preference uptake
Full input
No situation
Full input
No memory
Model
Δ
Δ
Model
Use +
Use -
Use +
Use -
Gemini 3.5
56.0
0.0
Gemini 3.5
78.8
64.7
11.1
7.5
Kimi K3
57.6
-0.2
Kimi K3
84.4
77.6
15.1
13.1
DeepSeek V4
53.5
-0.5
DeepSeek V4
84.1
73.4
11.7
9.3
GLM-5.3
56.6
-0.7
GLM-5.3
87.5
84.2
16.3
15.8
Table 3: Situation and memory removal.
Figure 4: Two controls on multiple choice. (a) Removing memory preserves positive-side but reduces negative-side selection. (b) Negative-side use remains similar across memory presentations.
Full pool (16)
Smaller pool (5)
Model
Use + ↑
Use - ↓
Pair success ↑
Use + ↑
Use - ↓
Pair success ↑
Gemini 3.5 Flash
78.9
22.9
57.9
81.5
25.8
57.8
DeepSeek V4 Pro
76.4
23.0
55.2
78.6
27.8
53.1
Kimi K3
76.0
18.3
59.5
79.7
20.7
61.3
GLM-5.3
75.1
18.4
58.4
77.7
21.0
59.0
GPT-6 Astra
71.2
14.7
57.9
73.3
11.7
62.5
Table 4: Effect of memory-pool size on selection performance.
Figure 5: Prompt wording. Bars are negative-side multiple choice (left axis). The line shows free-answer always-use (right).
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Situation category
Pairs
Companions
205
Occasion
163
Time constraint
136
Beneficiary shift
296
Role shift
54
Resource constraint
110
Appendix
Table 5: Situation coverage in PairPref.
Benchmark
Comparison or variation
Boundary assessment
Usefulness assessment
PrefEval
Explicit or implicit preferences; conversation length
As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.
Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typically rely on heuristic retrieval signals or always-on memory usage, failing to account for the changing and potentially inconsistent nature of user preferences. In this work, we propose a unified framework for memory access and selection based on changing preferences. We formulate personalized memory retrieval as identifying which historical turns provide evidence about a user's latent preference state, rather than relying on surface-level semantic similarity. To this end, we quantify the utility of each memory turn using a Bayes factor, defined as the improvement in the model's likelihood of the reference response when the turn is included in context. This provides a principled measure of evidence strength and a unified signal for both memory access and selection. By framing memory retrieval as utility estimation, the model learns to identify salient turns and regulate memory usage based on expected utility. Experiments on four heterogeneous memory benchmarks show that our approach outperforms existing embedding-based retrieval on long-context, preference-intensive tasks where modeling changing preferences is essential, while remaining competitive in low-density regimes where semantic similarity suffices.
User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this problem as long-horizon personalization and observe that progress on it is limited by data availability and measurement, with no existing resource providing both naturalistic long-horizon interactions and the ground-truth provenance needed to diagnose why models fail. We introduce a data generator that produces conversations from a structured mental state graph, yielding ground-truth provenance for every preference change across 6-month timelines, and from it construct HorizonBench, a benchmark of 4,245 items from 360 simulated users with 6-month conversation histories averaging ~4,300 turns and ~163K tokens. HorizonBench provides a testbed for long-context modeling, memory-augmented architectures, theory-of-mind reasoning, and user modeling. Across 25 frontier models, the best model reaches 52.8% and most score at or below the 20% chance baseline. When these models err on evolved preferences, over a third of the time they select the user's originally stated value without tracking the updated user state. This belief-update failure persists across context lengths and expression explicitness levels, identifying state-tracking capability as the primary bottleneck for long-horizon personalization.