cs.CLSep 28, 2026

Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs

Authors: Haeun Jang, Yonghyun Jun, Hwanhee Lee

Organizations: Department of Artificial Intelligence, Chung-Ang University

Abstract

Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings

    May 20, 2026Philipp Spohn, Leander Girrbach, Zeynep AkataLarge Language Model PersonalizationPersonalization

  2. POPI: Personalizing LLMs via Optimized Natural Language Preference Inference

    Oct 17, 2025Yizhuo Chen, Xin Liu, Ruijie Wang +7Large Language Model PersonalizationPersonalization

  3. Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation

    Sep 28, 2026Jiashen Ren, Wenlin Zhang, Bohan Zhang +5Artificial Intelligence PersonasPersona Consistency