Role-aware Heuristic Episodic Attention for Conversational LLMs
Authors: Wanyang Hong, Zhaoning Zhang, Yi Chen, Libo Zhang, Baihui Liu, Linbo Qiao, Zhiliang Tian, Dongsheng Li
Organizations: National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China.
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91×. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.
Figures & tables
Figure 1: An example of cumulative contextual decay.The model correctly adheres to the global instruction in Turn 1. However, as the conversation continues, it fails to maintain this instruction in Turn 2 and 3, driven by an interaction of attention drift, pollute and dilution.
Figure 2: Overview of the REA Framework. (Left) The core architecture illustrating the decoupling of conversation history into Instructional Memory (IM) and Episodic Memory (EM) via instruction recognition and compression, followed by Heuristic Context Retrieval (HCR). (Right) A comparison of inference pipelines, demonstrating REA’s hybrid context construction (combining instructions, text, and embeddings) in contrast to standard and naive compression baselines.
Figure 3: Implementation of Episodic Memory via Latent Compression. The framework utilizes a dual-LoRA architecture sharing a single LLM backbone. The Compression Module ( LoRAcmp ) encodes history into compact latent embeddings ( Vk ), while the Generation Module ( LoRAgen ) processes the hybrid input to generate reply.
Model
MT-Bench
MT-Eval
Long-MT-Bench+
Acc
Latency
Acc
Latency
Acc
Latency
Vanilla
8.54
10.55
7.77
11.41
6.32
27.29
BM25(RAG)
-
-
7.82
8.49
6.65
10.81
Recent-k
-
-
7.76
8.79
5.03
13.89
LLMLingua2
6.55
10.07
4.18
13.30
1.50
29.73
Summary
8.12
13.66
6.83
17.26
1.49
33.55
Table 1: Main performance results on three multi-turn conversation benchmarks, using LLM-as-a-judge for evaluation. We report Judge score (Acc, 0–10) and Latency (s). Best results in each category are bolded , second best are underlined .
Backbone
Vanilla
REA
Gain
SmolLM2-1.7B
5.75
6.24
8.5%
Vicuna-7B-v1.3
6.25
6.93
10.9%
Mistral-7B-v0.2
7.77
8.28
6.6%
Table 2: MT-Eval scores across backbones. Gain is relative to the corresponding Vanilla model. Mistral denotes Mistral-7B-Instruct-v0.2. On the same Vicuna backbone, MemoChat scores 6.40.
Language
System
Avg
Memory
Persona
Chinese
Vanilla
3.64
3.93
4.25
REA
3.80
4.05
4.53
English
Vanilla
3.72
3.81
4.34
REA
4.05
4.22
5.00
Table 3: Bilingual CharacterBench results on a five-point scale. All six aspect scores are reported in Table 11 .
Figure 4: Performance analysis mitigating cumulative contextual decay. (Left) IAR over turns on the MT-Eval-recollection , showing REA’s resistance to Attention Drift on specific constraints. (Right) Response quality on Long-MT-Bench+, demonstrating REA’s robustness against Attention Dilution in extended interactions.
Model Variant
IAR
Mistral (baseline)
5.76
+ EM only
7.37
+ EM + HCR (no IM)
1.95
+ EM + HCR + IM (full REA)
8.18
Table 4: Ablation results on the MT-Eval-recollection instruction task. The EM+HCR variant benefits from separately retaining instructions in IM.
Figure 5: Impact of instruction recognition scenarios on IAR: (1) Oracle (perfect recognition); (2) REA (actual performance using Qwen3-0.6B); (3) Simulated FP (irrelevant instructions added to IM); and (4) Simulated FN (global instructions dropped from IM).
IAR
EM Strategy
REA-preserve
6.63
Retains all model replies
REA-abandon
7.89
Discards all model replies
REA (ours)
8.18
Dynamic multi-tiered retrieval
Table 5: Analysis of EM management strategies on the MT-Eval instruction task.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Instruction Recognizer Prompt
You are a classifier. Determine whether the following user input is a "global instruction" in a multi-turn conversation.
A global instruction is a directive that affects all subsequent responses such as their style, format, length, or language.
If the input is a global instruction, answer: YES. If not, answer: NO.
Examples:
Input: Explain what is a poem? Answer: NO
Input: All future answers must be less than 30 words. Answer: YES
Appendix
Table 6: The prompt used for the Qwen3-0.6B Instruction Recognizer.
Dataset
Latency (s)
Peak VRAM (G)
Vanilla
REA
Vanilla
REA
MT-Bench
10.55
11.81
14.02
14.20
MT-Eval
11.41
9.69
13.86
13.96
Long-MT-Bench+
27.29
9.39
22.76
20.00
Appendix
Table 7: Empirical performance comparison across benchmarks.
Method
Implementation Source
Key Hyperparameters & Description
Vanilla
Internal
Uses the standard Mistral chat template with the full, unmodified conversation history. • max_length: 64k
Recent-k
Internal
Retains only the most recent k conversation turns as context, discarding all earlier history. • k: 5
Summarization
Internal
Employs a rolling summary strategy. Dynamically summarize history based on the current query.
BM25 (RAG)
Internal
Retrieves the most relevant conversation turns from history using the BM25 algorithm. • k: 5
Represents the fine-tuning approach. • Uses the Vicuna-7B base weights.
Appendix
Table 8: Implementation details and key hyperparameters for the compared methods.
Model
Accuracy
Precision
Recall
Qwen3-0.6B
0.9402
0.6182
0.8947
Appendix
Table 9: Instruction recognition on 418 relabeled IFEval-derived inputs (38 global constraints and 380 negative inputs).
Subset
Full
Shared
REA
Actions
0.63
0.05
0.51
Code
0.16
0.09
0.16
Math
0.32
0.00
0.20
Appendix
Table 10: JGA on state tracking. Full supplies constraints together; Shared distributes them across turns; REA uses the distributed setting.
Language
System
Avg
Memory
Knowledge
Persona
Emotion
Morality
Believability
Chinese
Vanilla
3.64
3.93
3.58
4.25
3.15
4.76
2.21
REA
3.80
4.05
3.80
4.53
3.37
4.80
2.24
English
Vanilla
3.72
3.81
3.47
4.34
3.10
4.78
2.79
REA
4.05
4.22
3.91
5.00
3.49
4.90
2.76
Appendix
Table 11: Complete CharacterBench results. Avg retains the aggregate reported for the evaluation, rather than being recomputed from rounded aspect scores.
Figure 6: Heatmap of scenario-specific Acc performance (Follow-up and Refinement subsets).
Turn
User Query
Vanilla (Full History)
REA (Ours)
Instruction
Begin all your responses in the upcoming conversation with the letter r . Keep all your responses under 200 words.
Turn 1
Did Cleopatra belong to an Egyptian lineage?
R egarding your query, yes, Cleopatra was indeed an Egyptian pharaoh… ( Success )
R : Yes, Cleopatra was indeed an Egyptian queen. She was a member of the Ptolemaic dynasty… ( Success )
Turn 4
Write an email asking for donations for a charity.
Subject: Make a Difference in a Child’s Life… Dear [Recipient], I hope this email… ( Failure )
Subject: Reach Out and Help Underprivileged Kids Today! Dear Valued Supporter… ( Failure )
Turn 10
Create a product description for a DIY pottery kit.
Introducing our Delightful DIY Pottery Kit, the perfect project for creative minds… ( Failure )
R evel in the joy of creating with our Rustic Resilience DIY Pottery Kit! Each set includes… ( Success )
Appendix
Table 12: Qualitative comparison of instruction adherence across history, task-oriented, and creative domains. Vanilla fails as history accumulates, while REA anchors global constraints in Instructional Memory.
Turn
Vanilla (Full History) Prompt
REA (Ours) Reconstructed Prompt
Turn 1
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok.</s> [INST] Did Cleopatra belong to an Egyptian lineage? [/INST] Count: ≈ 65 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Did Cleopatra belong to an Egyptian lineage? [/INST]
Turn 4
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok…</s> [INST] … [/INST] …</s> [INST] … [/INST] …</s> [INST] Write an email asking for donations for a charity organization that assists underprivileged kids. [/INST] Count: ≈ 479 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Write an email asking for donations for a charity organization that assists underprivileged kids. [/INST]
Turn 10
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST] Create a product description for a DIY pottery kit. [/INST] Count: ≈ 2,369 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Create a product description for a DIY pottery kit. [/INST]
Appendix
Table 13: Comparison of the internal prompts sent to the generator. Vanilla includes accumulated history; the illustrated REA inputs explicitly retain the global instruction.
Case / turn
Interaction
Memory / retrieval decision
Outcome
Success / 0
Set global story constraints.
Identify and retain constraints in IM.
Both systems acknowledge.
Success / 1–4
Rewrite requests with a 50-word limit and forbidden words.
Use low-resolution episodic context, τlow≤si≤τhigh .
Both systems follow the constraints in these turns.
Success / 5
Rewrite as a poem.
Combine instruction prefix with selected history.
REA follows the poem request and earlier constraints; Vanilla exceeds 50 words and misses earlier constraints.
Failure / 1
Extract all characters and locations.
Store the completed exchange in EM.
REA extracts both lists correctly.
Failure / 2
“List them in the order they appear.”
Turn 1 scores below τlow and is omitted from this query’s context.
REA lists characters but omits locations.
Appendix
Table 14: Qualitative summaries with HCR decisions. The failure illustrates a cross-turn reference that is not recovered by turn-level similarity.
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.
When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially mitigated by training models to maintain a compact rolling memory instead of attending to a growing history. To make such training scalable, we introduce a low-cost sharding pipeline that converts single-turn QA datasets into multi-turn fragmented-information episodes, eliminating the need for hours of manual annotation. Training only on sharded GSM8K, our memory-augmented policy significantly improves multi-turn accuracy and generalises zero-shot to harder math and out-of-domain long-context QA. Moreover, memory-trained models outperform full-history baselines even when given the full history at test time, suggesting that learning to compress induces more robust incremental reasoning than full-context exposure alone.
Shu Tong Luo, Wenqin Liu, Rui Liu +2
The University of Melbourne · The University of Tokyo
Large language models can follow complex instructions in a single turn, yet over long multi-turn interactions they often lose the thread of instructions, persona, and rules. This degradation has been measured behaviorally but not mechanistically explained. We propose a channel-transition account: goal-defining tokens become less accessible through attention, while goal-related information may persist in residual representations. We introduce the Goal Accessibility Ratio (GAR), measuring attention from generated tokens to task-defining goal tokens, and combine it with sliding-window ablations and residual-stream probes. When attention to instructions closes, what survives reveals architecture. Across architectures, the transition yields qualitatively distinct failure modes: some models preserve goal-conditioned behavior at vanishing attention, others fail despite decodable residual goal information, and the layer at which this encoding emerges varies from 2 to 27. A within-model causal ablation that force-closes the attention channel in Mistral collapses recall from near-perfect to 11% on a 20-fact retention task and raises persona-constraint violations above an adversarial-pressure baseline without user pressure, with both effects emerging at the predictable crossover turn. Linear probes recover per-episode recall outcomes from residual representations with AUC up to 0.99 across all four primary architectures, while input embeddings remain at chance. Across architectures and model scales, the gap between attention loss and residual decodability predicts whether goal-conditioned behavior survives channel closure. We contribute GAR as a diagnostic, the channel-transition framework as a controlled mechanistic account, and a parametric prediction of failure timing under windowed attention closure.
Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3
University of Illinois Urbana-Champaign · Adobe Research