When Context Changes: Understanding Update Failures in LLMs
Organizations: University of California, Berkeley
Abstract
As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.
Figures & tables
| Intervention comparison | Result |
|---|---|
| Answer-score-gap recovery (%) | |
| Qwen: late / early–middle states | |
| Qwen: old / current keys | |
| Pythia key inputs: targeted / random | |
| Old-value errors corrected (%) | |
| Pythia head removal: targeted / random | |
| Accuracy | Old errors | Preserved | |||
| Model | Original | Routing | Reminder | corrected | |
| Adaptive attention bias | |||||
| Qwen2.5-3B | 16.46 | 95.21 | 99.27 | 678/711 | 98.73 |
| Llama-3.2-3B | 27.71 | 93.12 | 68.85 | 514/541 | 100.00 |
| Llama-3.1-8B | 41.71 | 88.63 | 65.38 | 395/467 | 100.00 |
| Mistral-7B | 5.42 | 81.35 | 98.44 | 564/674 | 98.08 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Claim | Models | Evidence and interpretation |
|---|---|---|
| Current value readable on old-value errors | Qwen-7B, Llama-8B; separate Pythia-160M test | Held-out probes versus shuffled labels; recoverability despite the model’s answer. |
| Known-answer direction controls output | Qwen-7B, Llama-8B | Final-layer steering versus equal-norm random and middle-layer changes; a diagnostic intervention. |
| More old-value attention on failures | Qwen-7B; not detected in Llama-8B | Failure versus correct, length-adjusted: +19.4 pp versus +1.2 pp (Llama interval includes zero). |
| Conflicting updates change internal matching | Five models with reliable controls in Appendix B.10 | Update versus matched no-update prompts, including correct answers; Gemma-2B has an unreliable control. |
| Specific components contribute causally | Qwen-7B; Pythia-160M | Old-assignment key replacement in Qwen; held-out head removal and key-input replacement in Pythia. |
| Model | Most recent old / old errors | Uniform reference | |
|---|---|---|---|
| Qwen2.5-7B | 2 | 54/54 (100.0%) | 50.0% |
| Qwen2.5-7B | 4 | 92/113 (81.4%) | 25.0% |
| Qwen2.5-7B | 8 | 72/164 (43.9%) | 12.5% |
| Qwen2.5-72B | 64 | 7/11 (63.6%) | 1.6% |
| Llama-3.1-70B | 64 | 8/16 (50.0%) | 1.6% |
| GPT-4o | 64 | 3/7 (42.9%) | 1.6% |
| Model / task | Answer format | Accuracy | Old / errors | |
|---|---|---|---|---|
| ICF-Bench, Qwen2.5-7B | ||||
| Dynamic Preference | Free-form | 783 | 69.3 | 82.1 |
| Dynamic Preference | Multiple choice | 783 | 70.9 | 29.4 |
| Probe score (%) | Failure correct (pp) | ||||
|---|---|---|---|---|---|
| Model | Old-value errors | Other-variable errors | Probe | Attention | |
| Qwen2.5-7B | 1200 | 84.8 | 86.5 | ||
| Llama-3.1-8B | 1200 | 82.2 | 86.3 | ||
| Component | Measurement | Mean | Median | |
| Answer-score-gap recovery (%) | ||||
| Residual stream | Early / middle layers | 0.0 | 600 | |
| Late layers | 78.1 | 93.9 | 360 | |
| Key | Old assignments | 49.2 | 44.8 | 120 |
| All assignments | 11.1 | 10.7 | 120 | |
| Current assignment | 120 | |||
| Model | Failed trials | Middle layer | Final layer |
|---|---|---|---|
| Qwen2.5-7B | 665 | ||
| Llama-3.1-8B | 722 |
| Head | Measurement | Correct answer | Old-value answer |
|---|---|---|---|
| L8H10 | Current-value QK margin | ||
| Attention to current value | |||
| Current-value OV margin | |||
| L8H2 | Current-value QK margin | ||
| Attention to current value | |||
| Current-value OV margin |
| Intervention | Measured effect | Targeted | Random |
|---|---|---|---|
| Remove heads favoring old values | Fraction of errors corrected | 0.38 | 0.09 |
| Remove heads favoring the current value | Fraction switching to an old value | 0.04 | 0.03 |
| Replace inputs to L8H2’s key | QK score-gap recovery | 0.86 | 0.04 |
| Answer-score-gap recovery | 0.75 | 0.06 |
| Component | Cohort | Task | Role |
|---|---|---|---|
| Preference core | 1,200 dialogues | Repeated preference updates with competing mentions | Retention, selection and causal interventions |
| Dialogue diversity | 180 pilot ledgers | Scalar, partial-record and full-record updates in six domains | Language and task variation |
| Constraint decisions | 24 scenarios | Booking and scheduling under revised constraints | Use of updated state in decisions |
| Operational histories | 40 + 40 scenarios | Dependent handoffs across 16 slots in warehouse and build-release logs, and a deferred-log tier resolves pending handoffs at later lines | Frontier reasoning under complex and dependency-aware state updates |
| Model | Heads | Setting | Corrected | Preserved | Group gain [95% interval] |
|---|---|---|---|---|---|
| Qwen2.5-3B | 32 | 678/711 | 156/158 | 78.37 [74.51, 82.21] | |
| Llama-3.2-3B | 32 | 514/541 | 266/266 | 64.60 [59.40, 69.63] | |
| Llama-3.1-8B | 16 | 395/467 | 400/400 | 46.90 [41.99, 51.93] | |
| Mistral-7B-v0.3 | 32 | 564/674 | 51/52 | 75.01 [70.33, 79.12] | |
| Gemma-2-9B | 32 | 407/455 | 452/452 | 45.06 [39.63, 50.99] | |
| Qwen2.5-7B | 8 | 208/487 | 361/361 | 31.80 [27.88, 35.76] |
| Model | Random positions | Random heads | Reverse routing |
|---|---|---|---|
| Qwen2.5-3B | |||
| Llama-3.2-3B | |||
| Llama-3.1-8B | |||
| Mistral-7B-v0.3 | |||
| Gemma-2-9B | |||
| Qwen2.5-7B | — |
| Qwen2.5-7B | GPT-4o | |||
|---|---|---|---|---|
| Variant | Near source | Far source | Near source | Far source |
| Original | 12.3 / 84.8 | 65.7 / 14.5 | 79.7 / 19.0 | 94.5 / 2.3 |
| Explicit historical wording | 62.5 / 25.5 | 65.7 / 14.5 | 98.5 / 0.0 | 94.5 / 2.3 |
| Neutral replacement | 62.3 / 17.0 | 65.7 / 14.5 | 95.0 / 2.8 | 94.5 / 2.3 |
| Explicit-update query | 7.3 / 87.3 | 61.2 / 9.8 | 88.7 / 8.8 | 94.8 / 1.7 |
| Combined | 39.2 / 41.8 | 61.2 / 9.8 | 99.3 / 0.0 | 94.8 / 1.7 |
| Model | Answer type | Far (%) | Near (%) | Difference, pp [95% CI] |
|---|---|---|---|---|
| Qwen2.5-7B | Current value | 58.8 | 46.7 | |
| Old value | 14.8 | 36.7 | ||
| GPT-4o | Current value | 98.3 | 96.5 | |
| Old value | 0.0 | 0.0 |
| Model | Output, token cap | Current | Old | No overwrite | Cross-record |
|---|---|---|---|---|---|
| Qwen2.5-7B | Direct, 32 | 0/36 | 36/36 | 34/36 | 36/36 |
| Qwen2.5-7B | Direct, 512 | 0/36 | 36/36 | 34/36 | 36/36 |
| Qwen2.5-7B | Concise, 512 | 0/36 | 32/36 | 36/36 | 33/36 |
| GPT-4o | Direct, 32 | 3/36 | 31/36 | 36/36 | 36/36 |
| GPT-4o | Direct, 512 | 3/36 | 32/36 | 36/36 | 36/36 |
| GPT-4o | Concise, 512 | 4/36 | 32/36 | 36/36 | 36/36 |
| Model | Far | Near | Difference [95%] |
|---|---|---|---|
| Qwen2.5-7B | 39/180 (21.7%) | 50/180 (27.8%) | |
| Llama-3.1-8B | 7/180 (3.9%) | 5/180 (2.8%) | |
| GPT-4o | 0/180 (0.0%) | 0/180 (0.0%) |
| Context | Task | Optimal H / S / P | Unavailable H / S / P |
|---|---|---|---|
| GPT-5 | |||
| 24 directives | Single | 11 / 12 / 12 | 1 / 0 / 0 |
| 24 directives | Composed | 12 / 12 / 12 | 0 / 0 / 0 |
| 96 directives | Single | 9 / 12 / 12 | 3 / 0 / 0 |
| 96 directives | Composed | 12 / 12 / 12 | 0 / 0 / 0 |
| Gemini 3.1 Pro Preview | |||
| Context | Task | Optimal H / S / P | Unavailable H / S / P |
|---|---|---|---|
| GPT-5 | |||
| 24 directives | Single | 12 / 12 / 12 | 0 / 0 / 0 |
| 24 directives | Composed | 12 / 12 / 12 | 0 / 0 / 0 |
| 96 directives | Single | 12 / 12 / 12 | 0 / 0 / 0 |
| 96 directives | Composed | 12 / 12 / 12 | 0 / 0 / 0 |
| Gemini 3.1 Pro Preview | |||
| Model | H P | H S | Change with composition |
|---|---|---|---|
| GPT-5 | 0.0 [0.0, 0.0] | 0.0 [0.0, 0.0] | 12.5 [0.0, 25.0] |
| Gemini 3.1 Pro Preview | 50.0 [37.5, 62.5] | 0.0 [0.0, 0.0] | 0.0 [-16.7, 16.7] |
| DeepSeek V4 Pro 0813 | 0.0 [0.0, 0.0] | 0.0 [0.0, 0.0] | 4.2 [0.0, 12.5] |
| Claude Sonnet 5 | 20.8 [8.3, 33.3] | 0.0 [0.0, 0.0] | 29.2 [12.5, 50.0] |
| Model and cohort | Full history | Current-state snapshot |
|---|---|---|
| GPT-5.6 Sol, confirmation | 9/40 | 40/40 |
| Accuracy [95% interval] | 22.5% [12.3, 37.5] | 100% [91.2, 100] |
| Claude Opus 4.8, confirmation | 18/40 (9 truncated) | 40/40 |
| Accuracy [95% interval] | 45.0% [30.7, 60.2] | 100% [91.2, 100] |
| Completed-answer accuracy | 18/31 (58.1%) | 40/40 (100%) |
| GPT-5.6 Sol, exploratory | 0/4 | 4/4 |
| Cohort | Deferred history | Linearized control | Snapshot |
|---|---|---|---|
| Confirmation prefix (17 scenarios) | 12/17 | 16/17 | 17/17 |
| Accuracy [95% interval] | 70.6% [46.9, 86.7] | 94.1% [73.0, 99.0] | 100% [81.6, 100] |
| Development (4 scenarios) | 2/4 | 4/4 | 4/4 |