cs.AISep 28, 2026

PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use

Authors: Mingfei Lu, Mengjia Wu, Yi Zhang

Organizations: AAII, University of Technology Sydney

Abstract

Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores (ΔΔ) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6% to 18.3% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Know It, Act on It: Investigating Memory Utilization in LLM Personalization

    Jul 31, 2026Zhaoxin Feng, Jianfei Ma, Emmanuele ChersoniLarge Language Model PersonalizationPeak Memory

  2. Memory Retrieval for Changing Preferences

    Jun 2, 2026Yuehan Qin, Li Li, Linxin Song +4Interaction HistoryPreference Alignment Learning

  3. HorizonBench: Long-Horizon Personalization with Evolving Preferences

    Apr 19, 2026Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9Preference Alignment LearningInteraction History