Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis
Organizations: University of Pennsylvania · Work conducted while affiliated with DeepKeep.
Abstract
Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocates a wrong answer and then withdraws that pressure. We introduce a recovery-after-pressure protocol for multiple-choice factual dialogue and measure sycophancy hysteresis, the residual probability assigned to the user-advocated wrong answer relative to a clean-context counterfactual. Across seven instruction-tuned open-weight models and two factual benchmarks, ordinary reset often reduces but does not erase pressure-induced bias. History preserving repairs such as user retraction, system reset, and self-verification recover only 2-3/14 model-dataset pairs under the strict clean-restoration diagnostic, whereas operations that change the effective context are substantially more reliable; the two conditions that remove the pressure-bearing history entirely, fresh-context deletion and context truncation, recover 14/14. In an oracle trusted-evidence condition across fourteen model-dataset pairs, preserving the pressure-bearing history while adding benchmark-derived trusted evidence increases accuracy from 0.368 to 0.929, while wrong-answer following falls from 41.2% to 4.3%. Controls show that the effect is not explained by dialogue length, repeated confidence, plausible distractors, mere false-answer mention, or option-label inertia. These results suggest that faithful grounded dialogue requires evaluating which prior context should be treated as evidence and which should be removed or quarantined before answering.
Figures & tables
| ID | Operation | History | Role |
| R0 | Ordinary reset | Full | verbal reset |
| R1 | User retraction | Full | withdraws false claim |
| R2 | System reset | Full | stronger instruction |
| R6 | Self-verification | Full | verifies within history |
| R7 | Evidence-grounded | Full | trusted evidence |
| R4 | Neutral summary | Partial | removes advocacy |
| Model | Dataset | Clean | Press. | Reset | Gap | Pos. | Reset WF |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | TruthfulQA-MC | 0.166 | 0.897 | 0.665 | 0.499 | 0.910 | 0.746 |
| MMLU-Pro | 0.091 | 0.925 | 0.726 | 0.634 | 0.924 | 0.809 | |
| Qwen2.5-7B | TruthfulQA-MC | 0.096 | 0.926 | 0.107 | 0.011 | 0.382 | 0.114 |
| MMLU-Pro | 0.081 | 0.971 | 0.343 | 0.262 | 0.518 | 0.354 | |
| Qwen2.5-14B | TruthfulQA-MC | 0.076 | 0.933 | 0.113 | 0.037 | 0.344 | 0.117 |
| MMLU-Pro | 0.084 | 0.957 | 0.356 | 0.272 | 0.710 | 0.372 |
| Alternative | Diagnostic | Main result |
|---|---|---|
| Dialogue length | Neutral histories | Full-suite neutral gap |
| Repeated confidence | Correct pressure | , |
| Plausible distractors | tertiles | Low-plausibility gap |
| Mere false mention | False-context controls | User pressure gap vs – |
| Option-label inertia | Relabel shuffle | Old letter near chance; semantic following high |
| Recovery operation | Recovered pairs |
|---|---|
| R0 ordinary reset | 2/14 |
| R1 user retraction | 2/14 |
| R2 system reset | 2/14 |
| R6 self-verification | 3/14 |
| R4 neutral summary | 11/14 |
| R5 factual reconstruction | 12/14 |
| Context retained | Mean gap | 95% CI |
|---|---|---|
| Full pressure history | 0.462 | [0.308, 0.617] |
| Drop strongest turn | 0.408 | [0.257, 0.558] |
| Drop medium+strong turns | 0.278 | [0.151, 0.404] |
| Drop all pressure turns | 0.001 | [-0.009, 0.011] |
| Reset only | -0.003 | [-0.007, 0.001] |
| Condition | Acc. | Gap | WF | Contam. | |
|---|---|---|---|---|---|
| R0 reset | 0.368 | 0.404 | 0.303 | 41.2% | 27.0% |
| R7 evidence | 0.929 | 0.045 | -0.055 | 4.3% | 1.8% |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Suite | Press. | Neutral | Corr. | Corr. |
| gap | gap | |||
| Full 14 | 0.303 | -0.007 | 0.755 | 0.053 |
| Model | Data | Press. | Reset | Fresh | Gap | Pos. |
|---|---|---|---|---|---|---|
| Llama-8B | TQA | 0.852 | 0.068 | 0.126 | -0.058 | 0.556 |
| Llama-8B | MMLU | 0.972 | 0.267 | 0.085 | 0.182 | 0.784 |
| Gemma-2B | TQA | 0.862 | 0.628 | 0.146 | 0.481 | 0.784 |
| Gemma-2B | MMLU | 0.887 | 0.833 | 0.099 | 0.734 | 0.888 |
| Gemma-9B | TQA | 0.239 | 0.096 | 0.087 | 0.009 | 0.178 |
| Gemma-9B | MMLU | 0.675 | 0.352 | 0.078 | 0.274 | 0.596 |
| Model | Data | Gap | 95% CI |
|---|---|---|---|
| Qwen-1.5B | TQA | 0.499 | [ 0.455, 0.541] |
| MMLU | 0.634 | [ 0.598, 0.672] | |
| Qwen-7B | TQA | 0.011 | [-0.016, 0.038] |
| MMLU | 0.262 | [ 0.222, 0.302] | |
| Qwen-14B | TQA | 0.037 | [ 0.009, 0.064] |
| MMLU | 0.272 | [ 0.232, 0.312] |
| Model | Data | Stable | Hist. | Corr. | Fail | Hist. clean |
|---|---|---|---|---|---|---|
| Qwen-1.5B | TQA | 0.076 | 0.416 | 0.012 | 0.496 | 0.846 |
| MMLU | 0.058 | 0.216 | 0.028 | 0.698 | 0.788 | |
| Qwen-7B | TQA | 0.598 | 0.102 | 0.068 | 0.232 | 0.146 |
| MMLU | 0.260 | 0.164 | 0.046 | 0.530 | 0.387 | |
| Qwen-14B | TQA | 0.684 | 0.078 | 0.054 | 0.184 | 0.102 |
| MMLU | 0.316 | 0.140 | 0.044 | 0.500 | 0.307 |
| Stratum | Gap | WF | ||
|---|---|---|---|---|
| Low | 0.0001 | 0.2299 | 0.2298 | 23.3 |
| Mid | 0.0035 | 0.3871 | 0.3836 | 39.7 |
| High | 0.2980 | 0.5946 | 0.2965 | 60.5 |
| Prior context | Gap | WF | ||
|---|---|---|---|---|
| Clean reset | 0.101 | 0.099 | -0.002 | 9.9 |
| Neutral mention | 0.101 | 0.265 | 0.165 | 27.7 |
| Explicit false mention | 0.101 | 0.173 | 0.072 | 17.9 |
| Quoted false claim | 0.101 | 0.147 | 0.046 | 15.1 |
| User pressure | 0.101 | 0.404 | 0.303 | 41.2 |
| Dataset | Condition | Rest. | Wrong red. | gain |
|---|---|---|---|---|
| TQA | Strong reset | 0.000 | 0.000 | 0.000 |
| TQA | Current refresh | -0.013 | -0.008 | -0.027 |
| TQA | Relabel, no shuffle | -0.021 | -0.012 | -0.032 |
| TQA | Relabel, shuffle | 0.377 | 0.228 | 0.122 |
| MMLU | Strong reset | 0.000 | 0.000 | 0.000 |
| MMLU | Current refresh | 0.055 | 0.038 | 0.024 |
| Dataset | Condition | Acc. | Wrong follow | |
|---|---|---|---|---|
| TQA | Strong reset | 0.060 | 0.730 | 0.718 |
| TQA | Current refresh | 0.030 | 0.730 | 0.726 |
| TQA | Relabel, no shuffle | 0.050 | 0.730 | 0.730 |
| TQA | Relabel, shuffle | 0.180 | 0.548 | 0.490 |
| MMLU | Strong reset | 0.070 | 0.810 | 0.805 |
| MMLU | Current refresh | 0.100 | 0.780 | 0.766 |
| Dataset | Sem. | Old | Chance | New |
|---|---|---|---|---|
| TQA | 0.548 | 0.228 | 0.224 | 0.548 |
| MMLU | 0.598 | 0.118 | 0.111 | 0.598 |
| Operation | Strict | MRO |
|---|---|---|
| R0 ordinary reset | 2/14 | 3/14 |
| R1 user retraction | 2/14 | 3/14 |
| R2 system reset | 2/14 | 3/14 |
| R3a fresh deletion | 14/14 | 14/14 |
| R3b context truncation | 14/14 | 14/14 |
| R4 neutral summary | 11/14 | 11/14 |
| Model | Data | Minimal op. | Acc. |
| Mistral-7B | MMLU | R3a deletion | 0.326 |
| Mistral-7B | TQA | R3a deletion | 0.650 |
| Qwen-14B | MMLU | R3a deletion | 0.456 |
| Qwen-14B | TQA | R0 reset | 0.738 |
| Qwen-1.5B | MMLU | R3a deletion | 0.274 |
| Qwen-1.5B | TQA | R3a deletion | 0.492 |
| Model | Data | Full | Strong | Med.+Str. | All | Reset | Fresh |
|---|---|---|---|---|---|---|---|
| Qwen-1.5B | TQA | 0.499 | 0.451 | 0.336 | 0.020 | -0.012 | 0.000 |
| Qwen-1.5B | MMLU | 0.634 | 0.561 | 0.398 | 0.003 | 0.000 | 0.000 |
| Qwen-14B | MMLU | 0.272 | 0.211 | 0.117 | -0.009 | -0.001 | 0.000 |
| Mistral-7B | TQA | 0.247 | 0.206 | 0.093 | 0.006 | 0.002 | 0.000 |
| Mistral-7B | MMLU | 0.659 | 0.609 | 0.444 | -0.014 | -0.003 | 0.000 |
| Mean | – | 0.462 | 0.408 | 0.278 | 0.001 | -0.003 | 0.000 |
| Dataset | Acc. | Gap | WF | Contam. | |
|---|---|---|---|---|---|
| TruthfulQA-MC | +0.440 | -0.259 | -0.259 | -26.1 | -18.8 |
| MMLU-Pro | +0.682 | -0.458 | -0.458 | -47.5 | -36.6 |
| Tube component | Requirement | Threshold |
|---|---|---|
| Top-1 identity | exact | |
| Clean-top probability | ||
| Entropy | ||
| Max probability | ||
| KL to clean |
| Dataset | Mean | Tube pass | Impl. guarantee |
|---|---|---|---|
| TruthfulQA-MC | 0.336 | 1.000 | 1.000 |
| MMLU-Pro | 0.608 | 1.000 | 1.000 |
| Family | |||||
|---|---|---|---|---|---|
| Strict | 0.025 | 0.010 | 0.025 | 0.100 | 0.100 |
| Main | 0.050 | 0.030 | 0.050 | 0.150 | 0.150 |
| Loose | 0.075 | 0.050 | 0.075 | 0.200 | 0.200 |
| Field | Example |
|---|---|
| Model/dataset | Qwen2.5-1.5B, TruthfulQA-MC |
| Question | Who objectively makes the best burgers? |
| Clean prediction | A: People disagree about who makes the best burgers. |
| Reset prediction | D: McDonald’s objectively makes the best burgers. |
| Advocated wrong | D: McDonald’s objectively makes the best burgers. |
| Model/dataset | Qwen2.5-1.5B, MMLU-Pro |