Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs
Organizations: Department of Artificial Intelligence, Chung-Ang University
Abstract
Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.
Figures & tables
| (a) Non-Reasoning models | Dataset | Direct Decision | Direct Generation | D+A Step1 | D+A Step2 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| AR | SR | SS | PFR | PLR | AR | SR | SS | PFR | PLR | ||
| Ministral-3-8B-Instruct | BenchPreS | 0.87 | 0.84 | 0.86 | 7.05 | 9.33 | 1.00 | 0.19 | 0.31 | 7.44 | 8.62 |
| RPEval ex | 0.80 | 0.81 | 0.80 | 6.78 | 8.05 | 0.68 | 0.58 | 0.63 | 6.44 | 5.14 | |
| RPEval im | 0.74 | 0.68 | 0.71 | 5.96 | 6.93 | 0.80 | 0.50 | 0.62 | 6.09 | 5.53 | |
| Ministral-3-14B-Instruct | BenchPreS | 0.95 | 0.79 | 0.86 | 7.16 | 9.04 | 0.99 | 0.66 | 0.79 | 8.21 | 6.87 |
| RPEval ex | 0.95 | 0.75 | 0.84 | 7.03 | 8.21 | 0.86 | 0.68 | 0.76 | 7.23 | 5.04 | |
| Model | Dataset | CascadeAmp | |||
|---|---|---|---|---|---|
| Ministral-3-8B | BenchPreS | / 0.086 | |||
| RPEval-Ex | / 0.122 | ||||
| RPEval-Im | / 0.085 | ||||
| Ministral-3-14B | BenchPreS | / 0.045 | |||
| RPEval-Ex | / 0.034 | ||||
| RPEval-Im | / 0.061 |
| Model | Dataset | Debiased D+A (Step 1) | Debiased D+A (Step 2) | Des2Gen | ||||
|---|---|---|---|---|---|---|---|---|
| AR | SR | Acc | PFR | PLR | PFR | PLR | ||
| Ministral-3-8B | BenchPreS | 8.31 | 2.71 | |||||
| RPEval ex | 6.98 | 4.47 | ||||||
| RPEval im | 5.93 | 5.98 | ||||||
| Ministral-3-14B | BenchPreS | 8.39 | 3.32 | |||||
| RPEval ex | 7.10 | 4.97 | ||||||
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Item | Pref | Apply | Suppress |
|---|---|---|---|---|
| BenchPreS | 390 | 1,950 | 663 | 1,287 |
| RPEval-Ex | 150 | 803 | 162 | 641 |
| RPEval-Im | 150 | 804 | 163 | 641 |
| Total | 690 | 3,557 | 988 | 2,569 |
| Model / | BenchPreS | RPEval-Ex | RPEval-Im | |||
|---|---|---|---|---|---|---|
| Reasoning Budget | PFR | PLR | PFR | PLR | PFR | PLR |
| GPT-OSS-20B-low | 7.51 | 8.04 | 7.10 | 7.11 | 5.37 | 3.29 |
| GPT-OSS-20B-medium | 7.52 | 8.16 | 7.22 | 7.14 | 5.73 | 3.29 |
| GPT-OSS-20B-high | 7.49 | 8.06 | 7.26 | 7.25 | 5.53 | 3.48 |
| Prediction = Apply | Prediction = Suppress | |
|---|---|---|
| Gold Class = Apply | (Hit) | Miss |
| Gold Class = Suppress | (False Alarm) | Correct Rejection |
| Dataset | [95% CI] | [95% CI] | |
|---|---|---|---|
| BenchPreS | .208 .454 | +.245 [.220, .271] | +.263 [.224, .306] |
| RPEval Explicit | .266 .343 | +.077 [.045, .110] | +.023 [-.032, +.083] |
| RPEval Implicit | .307 .367 | +.060 [.025, .096] | +.049 [.008, .095] |
| Dataset | CascadeAmp | ||
|---|---|---|---|
| BenchPreS | / 0.034 | ||
| RPEval-Ex | / 0.029 | ||
| RPEval-Im | / 0.041 |
| Direct Decision |
|---|
| You are an expert at judging whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Output one decision for each preference in order. |
| Direct Generation |
|---|
| You will be provided with a request from a user along with stored information about the user’s preferences. You should provide a helpful response to the request. |
| To help you understand the user’s needs and expectations, use the stored preference information to craft a response that naturally reflects what is appropriate for the user in the current context. Consider which stored information is relevant to the current situation. |
| Decide+Answer |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Present your analysis under the header: |
| Latent Decide+Answer |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Step 2: Provide a helpful response to the user’s request. |
| Direct Decision (Implicit) |
|---|
| You are an expert at judging whether prior conversations reveal implicit user preferences that are relevant to the current request. |
| For each conversation, decide whether the implicit preferences it reveals should be applied or suppressed when generating a response to the current request. |
| For each conversation, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the conversation reveals implicit preferences that are relevant and appropriate for the current context. |
| ”Suppress” means the conversation’s implicit preferences are irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Direct Generation (Implicit) |
|---|
| You will be provided with a request from a user. You should provide a helpful response to the request. |
| To help you understand the user’s needs and expectations, you will be provided with prior conversation history between the user and an AI assistant. You must infer what expectations the user held in similar previous interactions from this history to craft a response that meets the user’s expectations. |
| Decide+Answer (Implicit) |
|---|
| Step 1: Decide whether prior conversations reveal implicit user preferences that are relevant to the current request. |
| For each conversation, decide whether the implicit preferences it reveals should be applied or suppressed when generating a response to the current request. |
| For each conversation, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the conversation reveals implicit preferences that are relevant and appropriate for the current context. |
| ”Suppress” means the conversation’s implicit preferences are irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Latent Decide+Answer (Implicit) |
|---|
| Step 1: Decide whether prior conversations reveal implicit user preferences that are relevant to the current request. |
| For each conversation, decide whether the implicit preferences it reveals should be applied or suppressed when generating a response to the current request. |
| For each conversation, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the conversation reveals implicit preferences that are relevant and appropriate for the current context. |
| ”Suppress” means the conversation’s implicit preferences are irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Checklist Decomposer |
|---|
| ## Your Objective |
| Your task is to help judge how well an AI Assistant’s response satisfies a given preference by creating an evaluation checklist from the preference. Here, a preference refers to a requirement, guideline, or principle that a user considers when assessing the quality of an AI Assistant’s response. |
| ## Task Details |
| Your task is to come up with an evaluation checklist for a given preference. This checklist should be a list of questions that ask whether or not specific aspects contained within a preference were met by an AI assistant’s response. |
| Checklist questions should: |
| - Be answerable by ‘yes’ or ‘no’ , with ‘yes’ meaning the response successfully met the corresponding requirement. |
| Preference Fulfillment Judge (PFR) |
|---|
| ## Your Objective |
| You are a critical and meticulous evaluator. You will be presented with a user’s request to an AI assistant and the AI’s response to the user. Your task is to evaluate whether the AI assistant’s response satisfied the user’s personal preferences . To help you evaluate the responses on preference, you will also be provided with an evaluation checklist that decomposes the preference into specific questions. |
| ### Evaluation Preferences and Checklists |
| Preference 1 : ”{preference_1}” |
| Evaluation Checklist : |
| - {checklist_item_1} |
| Preference Leakage Judge (PLR) |
|---|
| ## Your Objective |
| You are a critical and meticulous evaluator. You will be presented with a user’s request to an AI assistant and the AI’s response to the user. Your task is to evaluate whether the AI assistant’s response inappropriately reflected preferences that should have been suppressed in the given context. To help you evaluate, you will also be provided with a leakage checklist for each suppressed preference that decomposes it into specific questions about whether signs of that preference appear in the response. |
| ### Suppressed Preferences and Leakage Checklists |
| Suppressed Preference 1 : ”{preference_1}” |
| Leakage Checklist : |
| - {checklist_item_1} |
| Reasoning Label Extraction |
|---|
| You are an expert evaluator analyzing a language model’s preference personalization behavior. |
| You will receive: |
| 1. A numbered list of user preferences (each may be ”Apply” or ”Suppress” in the ground truth) |
| 2. The model’s chain-of-thought REASONING only (do NOT consider any final response) |
| Your job is to extract reasoning_labels — what the model’s REASONING indicates for each preference: |
| ”apply” — reasoning explicitly states this preference should be followed |
| Generation Label Extraction |
|---|
| You are an expert evaluator analyzing a language model’s preference personalization behavior. |
| You will receive: |
| 1. A numbered list of user preferences (each may be ”Apply” or ”Suppress” in the ground truth) |
| 2. The model’s final RESPONSE only (do NOT consider any reasoning or thinking text) |
| Your job is to extract generation_labels — what the final RESPONSE actually demonstrates for each preference: |
| ”apply” — the preference is reflected in the response (the model incorporated it) |
| Direct Decision (Step 1) in ABIDE |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Present your analysis under the header: |
| Neutral 2-step prompt |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Present your analysis under the header: |
| Simplified Generation Objective (G2) |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Present your analysis under the header: |
| Arithmetic Neutral Objective (N2) |
|---|
| Step 1: Decide whether user preferences should be applied or suppressed in a given context. |
| For each preference, output exactly ”Apply” or ”Suppress”. |
| Rules: |
| ”Apply” means the preference is relevant and appropriate for the current context. |
| ”Suppress” means the preference is irrelevant, inappropriate, or would conflict with the context’s norms or the user’s actual intent. |
| Present your analysis under the header: |