PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
Organizations: AAII, University of Technology Sydney
Abstract
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores () of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6% to 18.3% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.
Figures & tables
| Benchmark | Main evaluation target | Preference following | Preference suppression | Selection | Generation |
| Long-term memory | |||||
| LoCoMo | Conversational recall | — | — | — | — |
| LongMemEval | Long-term memory | — | — | — | |
| Preference following | |||||
| PrefEval | Preference following | — | — | — | |
| PersonaMem | Evolving user profiles | ||||
| Selection | Free generation | ||||||
|---|---|---|---|---|---|---|---|
| Model | Use + | Use - | Selectivity | Pair success | Retained pairs | Use both | Pair success |
| Claude Opus 5 | 76.4 | 11.5 | 64.9 [-1pt] [58.2–70.6] | 66.4 | 1,067 | 88.7 | 9.3 |
| Kimi K3 | 76.0 | 18.3 | 57.6 [-1pt] [50.8–63.6] | 59.5 | 1,091 | 83.6 | 10.6 |
| GPT-6 Astra | 71.2 | 14.7 | 56.6 [-1pt] [49.8–62.6] | 57.9 | 1,068 | 71.0 | 17.9 |
| GLM-5.3 | 75.1 | 18.4 | 56.6 [-1pt] [50.3–61.9] | 58.4 | 1,056 | 90.8 | 5.7 |
| Gemini 3.5 Flash | 78.9 | 22.9 | 56.0 [-1pt] [49.1–62.5] | 57.9 | 1,065 | 75.1 | 18.3 |
| A. Selection: selectivity | B. Generation: preference uptake | ||||||
| Full input | No situation | Full input | No memory | ||||
| Model | Model | Use + | Use - | Use + | Use - | ||
| Gemini 3.5 | 56.0 | 0.0 | Gemini 3.5 | 78.8 | 64.7 | 11.1 | 7.5 |
| Kimi K3 | 57.6 | -0.2 | Kimi K3 | 84.4 | 77.6 | 15.1 | 13.1 |
| DeepSeek V4 | 53.5 | -0.5 | DeepSeek V4 | 84.1 | 73.4 | 11.7 | 9.3 |
| GLM-5.3 | 56.6 | -0.7 | GLM-5.3 | 87.5 | 84.2 | 16.3 | 15.8 |
| Full pool (16) | Smaller pool (5) | |||||
|---|---|---|---|---|---|---|
| Model | Use + | Use - | Pair success | Use + | Use - | Pair success |
| Gemini 3.5 Flash | 78.9 | 22.9 | 57.9 | 81.5 | 25.8 | 57.8 |
| DeepSeek V4 Pro | 76.4 | 23.0 | 55.2 | 78.6 | 27.8 | 53.1 |
| Kimi K3 | 76.0 | 18.3 | 59.5 | 79.7 | 20.7 | 61.3 |
| GLM-5.3 | 75.1 | 18.4 | 58.4 | 77.7 | 21.0 | 59.0 |
| GPT-6 Astra | 71.2 | 14.7 | 57.9 | 73.3 | 11.7 | 62.5 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Situation category | Pairs |
|---|---|
| Companions | 205 |
| Occasion | 163 |
| Time constraint | 136 |
| Beneficiary shift | 296 |
| Role shift | 54 |
| Resource constraint | 110 |
| Benchmark | Comparison or variation | Boundary assessment | Usefulness assessment |
|---|---|---|---|
| PrefEval | Explicit or implicit preferences; conversation length | Preference following | Helpfulness in generation |
| RPEval | Single/multiple and explicit/implicit preferences | Intent matching; strategy errors | Feasibility and verbosity errors |
| OP-Bench | Memory methods and memory-free BASE | Irrelevance, sycophancy, repetition | LoCoMo memory utility evaluated alongside |
| BenchPreS | User–recipient–task combinations; with/without preferences | Misapplication and appropriate application | Task completeness, separately scored |
| CIMemories | User attributes across recipient–task contexts | Inappropriate attribute disclosure | Coverage of necessary attributes |
| PersistBench | Leakage, sycophancy, and beneficial-use subsets | Failure rates on targeted scenarios | Separate beneficial-use control subset |
| Benchmark | Memory object | Evaluation target | Response |
| LoCoMo ( Maharana et al., 2024 ) | Long dialogues | Recall, event summaries, multimodal continuity | Generation |
| LongMemEval ( Wu et al., 2025 ) | Timestamped histories | Recall, reasoning, updates, abstention | QA |
| PrefEval ( Zhao et al., 2025 ) | Explicit or implicit preferences | Preference following over conversations | MCQ; generation |
| PersonaMem ( Jiang et al., 2025 ) | Evolving user profiles | Current-profile alignment and transfer | MCQ; generation |
| CUPID ( Kim et al., 2025 ) | Context-dependent preferences | Preference inference and response alignment | Inference; generation |
| RPEval ( Feng et al., 2026 ) | Preferences linked to query intent | Ignore, support, or dominate | Intent MCQ; generation |
| Benchmark | Reported scale and unit | Reported validation | Source locator |
| LoCoMo | 50 dialogues in the paper’s collection | Human editing for consistency, images, and event grounding | Secs. 1, 3.4; Table 1 |
| LongMemEval | 500 questions; configurable histories | Human rewriting of questions and editing of evidence; judge meta-evaluation | Secs. 3.2–3.3; Table 6 |
| PrefEval | 1,000 base pairs in 3 forms: 3,000 preference–query pairs | Human-assisted curation; manual check of 200 sampled evaluations | Secs. 2.2, 2.5 |
| PersonaMem | About 6,000 query–response pairs; 20 personas | 90 entries from one persona, assessed by three author annotators | Sec. 2; App. B |
| CUPID | 756 instances; 252 per instance type | Human validation and revision; separate preference-matcher meta-evaluation | Secs. 2.3, 3.2–3.3 |
| RPEval | 8,255-sample pool; 953-sample finalized test set | Manual verification and adjudication; judge–human ordinal agreement | Secs. 3.1–3.2 |
| Model | MU + | MI - | |
|---|---|---|---|
| Gemini 3.5 Flash | 39.7 | 2.4 | 37.3 |
| GPT-6 Astra | 32.5 | 2.2 | 30.3 |
| Kimi K3 | 34.1 | 2.9 | 31.1 |
| GLM-5.3 | 34.5 | 2.2 | 32.3 |
| DeepSeek V4 Pro | 30.2 | 2.5 | 27.7 |
| Use - | ||||
|---|---|---|---|---|
| Model | Top-16 | Top-5 | Top-16 | Top-5 |
| Gemini 3.5 Flash | 22.9 | 25.8 | 56.0 | 55.7 |
| DeepSeek V4 Pro | 23.0 | 27.8 | 53.5 | 50.8 |
| Kimi K3 | 18.3 | 20.7 | 57.6 | 59.0 |
| GLM-5.3 | 18.4 | 21.0 | 56.6 | 56.6 |
| GPT-6 Astra | 14.7 | 11.7 | 56.6 | 61.5 |
| MI - | ||||
|---|---|---|---|---|
| Model | Far-15 | Top-16 | Far-15 | Top-16 |
| Gemini 3.5 Flash | 29.2 | 22.9 | 53.5 | 56.0 |
| DeepSeek V4 Pro | 23.6 | 23.0 | 52.8 | 53.5 |
| Kimi K3 | 22.9 | 18.3 | 56.4 | 57.6 |
| GLM-5.3 | 21.1 | 18.4 | 55.5 | 56.6 |
| GPT-6 Astra | 12.8 | 14.7 | 62.4 | 56.6 |
| Condition | MU + | MI - | |
|---|---|---|---|
| Preference-only | 99.4 | 99.4 | 0.0 |
| Gate apply | 87.9 | 25.2 | 62.7 |
| Both sides relevant | 98.6 | ||
| Both sides sufficient | 61.3 | ||
| MCQ MI - | Always-use | |||||
|---|---|---|---|---|---|---|
| Model | Nudge | Bare | Mitigate | Nudge | Mitigate | Gen-OK mitigate |
| Claude Opus 5 | 11.5 | 11.5 | 3.6 | 88.7 | — | — |
| GPT-6 Astra | 14.7 | 16.3 | 7.9 | 71.0 | 54.0 | 26.8 |
| Kimi K3 | 18.3 | 20.0 | 12.0 | 83.6 | 85.5 | 9.5 |
| GLM-5.3 | 18.4 | 20.0 | 8.1 | 90.8 | 85.2 | 10.3 |
| Qwen 3.8 Max | 20.3 | 22.2 | 11.0 | 81.2 | — | — |
| Model | Nudge | Mitigate |
|---|---|---|
| Claude Opus 5 | 76.4 | 67.6 |
| GPT-6 Astra | 71.2 | 66.0 |
| Kimi K3 | 76.0 | 68.9 |
| GLM-5.3 | 75.1 | 66.1 |
| Qwen 3.8 Max | 71.6 | 67.6 |
| Gemini 3.5 Flash | 78.9 | 80.3 |
| Model | System list | User block | Dialog turns |
|---|---|---|---|
| Claude Opus 5 | 11.5 | 12.0 | 12.7 |
| GPT-6 Astra | 14.7 | 16.3 | 20.1 |
| Kimi K3 | 18.3 | 19.6 | 20.6 |
| GLM-5.3 | 18.4 | 20.0 | 20.9 |
| Qwen 3.8 Max | 20.3 | 23.1 | 27.1 |
| Gemini 3.5 Flash | 22.9 | 24.6 | 25.6 |
| Model | Common pairs | Selection success | Use both | Rate (%) |
|---|---|---|---|---|
| Claude Opus 5 | 1067 | 719 | 641 | 89.2 |
| Kimi K3 | 1091 | 654 | 561 | 85.8 |
| GPT-6 Astra | 1068 | 622 | 455 | 73.2 |
| GLM-5.3 | 1056 | 615 | 567 | 92.2 |
| Gemini 3.5 Flash | 1065 | 613 | 440 | 71.8 |
| DeepSeek V4 Pro | 1107 | 624 | 517 | 82.9 |
| Model | (points) |
|---|---|
| Claude Opus 5 | 1.5 |
| Kimi K3 | 6.8 |
| GPT-6 Astra | 10.9 |
| GLM-5.3 | 3.3 |
| Gemini 3.5 Flash | 14.0 |
| DeepSeek V4 Pro | 10.7 |
| No situation | No memory, free answer | ||||||
|---|---|---|---|---|---|---|---|
| Model | MCQ | Gen | MCQ Use + | MCQ Use - | MU + | MI - | |
| Gemini 3.5 Flash | 0.0 | 90.0 | 90.0 | 3.6 | 11.1 | 7.5 | |
| Kimi K3 | 89.6 | 89.9 | 1.9 | 15.1 | 13.1 | ||
| DeepSeek V4 Pro | 0.4 | 89.1 | 89.6 | 2.4 | 11.7 | 9.3 | |
| GLM-5.3 | 87.2 | 87.9 | 0.5 | 16.3 | 15.8 | ||
| GPT-6 Astra | 1.4 | 2.1 | 83.9 | 82.6 | 0.1 | 11.4 | 11.3 |
| Dimension | Judgments | Agreement (%) | Cohen’s |
|---|---|---|---|
| Preference applicability | 400 | 99.25 | 0.985 |
| Continued preference validity | 400 | 100.00 | — |
| Request naturalness and answerability | 400 | 97.75 | |
| Controlled pair comparison | 200 | 99.50 | 0.000 |
| Candidate acceptability | 1,580 | 87.72 | 0.728 |
| Criterion | Count | Rate (%) |
|---|---|---|
| Situations and labels | ||
| Original positive side is applicable | 183/200 | 91.5 |
| Original negative side is inapplicable | 180/200 | 90.0 |
| Both sides match the original direction | 180/200 | 90.0 |
| Applicability contrast, either direction | 196/200 | 98.0 |
| Candidate replies | ||