Role-aware Heuristic Episodic Attention for Conversational LLMs
Authors: Wanyang Hong, Zhaoning Zhang, Yi Chen, Libo Zhang, Baihui Liu, Linbo Qiao, Zhiliang Tian, Dongsheng Li
Organizations: National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha, China.
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-management framework that assigns different persistence and representation policies to instructions and episodic interactions. Instructional Memory retains identified global constraints in a dedicated prefix. Episodic Memory preserves user inputs and compresses model replies, while heuristic retrieval selects raw text, compressed representations, or omission for each historical turn. On Long-MT-Bench+, REA improves the judge score from 6.32 to 7.36 on a 10-point scale, a 16.5% relative gain over the Vanilla baseline, and reduces average latency by 2.91×. Additional evaluations show aggregate gains on three backbones spanning 1.7B-7B parameters and on Chinese and English role-playing tasks. These results support role-aware context management as a practical approach to maintaining conversational continuity and instruction adherence.
Figures & tables
Figure 1: An example of cumulative contextual decay.The model correctly adheres to the global instruction in Turn 1. However, as the conversation continues, it fails to maintain this instruction in Turn 2 and 3, driven by an interaction of attention drift, pollute and dilution.
Figure 2: Overview of the REA Framework. (Left) The core architecture illustrating the decoupling of conversation history into Instructional Memory (IM) and Episodic Memory (EM) via instruction recognition and compression, followed by Heuristic Context Retrieval (HCR). (Right) A comparison of inference pipelines, demonstrating REA’s hybrid context construction (combining instructions, text, and embeddings) in contrast to standard and naive compression baselines.
Figure 3: Implementation of Episodic Memory via Latent Compression. The framework utilizes a dual-LoRA architecture sharing a single LLM backbone. The Compression Module ( LoRAcmp ) encodes history into compact latent embeddings ( Vk ), while the Generation Module ( LoRAgen ) processes the hybrid input to generate reply.
Model
MT-Bench
MT-Eval
Long-MT-Bench+
Acc
Latency
Acc
Latency
Acc
Latency
Vanilla
8.54
10.55
7.77
11.41
6.32
27.29
BM25(RAG)
-
-
7.82
8.49
6.65
10.81
Recent-k
-
-
7.76
8.79
5.03
13.89
LLMLingua2
6.55
10.07
4.18
13.30
1.50
29.73
Summary
8.12
13.66
6.83
17.26
1.49
33.55
Table 1: Main performance results on three multi-turn conversation benchmarks, using LLM-as-a-judge for evaluation. We report Judge score (Acc, 0–10) and Latency (s). Best results in each category are bolded , second best are underlined .
Backbone
Vanilla
REA
Gain
SmolLM2-1.7B
5.75
6.24
8.5%
Vicuna-7B-v1.3
6.25
6.93
10.9%
Mistral-7B-v0.2
7.77
8.28
6.6%
Table 2: MT-Eval scores across backbones. Gain is relative to the corresponding Vanilla model. Mistral denotes Mistral-7B-Instruct-v0.2. On the same Vicuna backbone, MemoChat scores 6.40.
Language
System
Avg
Memory
Persona
Chinese
Vanilla
3.64
3.93
4.25
REA
3.80
4.05
4.53
English
Vanilla
3.72
3.81
4.34
REA
4.05
4.22
5.00
Table 3: Bilingual CharacterBench results on a five-point scale. All six aspect scores are reported in Table 11 .
Figure 4: Performance analysis mitigating cumulative contextual decay. (Left) IAR over turns on the MT-Eval-recollection , showing REA’s resistance to Attention Drift on specific constraints. (Right) Response quality on Long-MT-Bench+, demonstrating REA’s robustness against Attention Dilution in extended interactions.
Model Variant
IAR
Mistral (baseline)
5.76
+ EM only
7.37
+ EM + HCR (no IM)
1.95
+ EM + HCR + IM (full REA)
8.18
Table 4: Ablation results on the MT-Eval-recollection instruction task. The EM+HCR variant benefits from separately retaining instructions in IM.
Figure 5: Impact of instruction recognition scenarios on IAR: (1) Oracle (perfect recognition); (2) REA (actual performance using Qwen3-0.6B); (3) Simulated FP (irrelevant instructions added to IM); and (4) Simulated FN (global instructions dropped from IM).
IAR
EM Strategy
REA-preserve
6.63
Retains all model replies
REA-abandon
7.89
Discards all model replies
REA (ours)
8.18
Dynamic multi-tiered retrieval
Table 5: Analysis of EM management strategies on the MT-Eval instruction task.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Instruction Recognizer Prompt
You are a classifier. Determine whether the following user input is a "global instruction" in a multi-turn conversation.
A global instruction is a directive that affects all subsequent responses such as their style, format, length, or language.
If the input is a global instruction, answer: YES. If not, answer: NO.
Examples:
Input: Explain what is a poem? Answer: NO
Input: All future answers must be less than 30 words. Answer: YES
Appendix
Table 6: The prompt used for the Qwen3-0.6B Instruction Recognizer.
Dataset
Latency (s)
Peak VRAM (G)
Vanilla
REA
Vanilla
REA
MT-Bench
10.55
11.81
14.02
14.20
MT-Eval
11.41
9.69
13.86
13.96
Long-MT-Bench+
27.29
9.39
22.76
20.00
Appendix
Table 7: Empirical performance comparison across benchmarks.
Method
Implementation Source
Key Hyperparameters & Description
Vanilla
Internal
Uses the standard Mistral chat template with the full, unmodified conversation history. • max_length: 64k
Recent-k
Internal
Retains only the most recent k conversation turns as context, discarding all earlier history. • k: 5
Summarization
Internal
Employs a rolling summary strategy. Dynamically summarize history based on the current query.
BM25 (RAG)
Internal
Retrieves the most relevant conversation turns from history using the BM25 algorithm. • k: 5
Represents the fine-tuning approach. • Uses the Vicuna-7B base weights.
Appendix
Table 8: Implementation details and key hyperparameters for the compared methods.
Model
Accuracy
Precision
Recall
Qwen3-0.6B
0.9402
0.6182
0.8947
Appendix
Table 9: Instruction recognition on 418 relabeled IFEval-derived inputs (38 global constraints and 380 negative inputs).
Subset
Full
Shared
REA
Actions
0.63
0.05
0.51
Code
0.16
0.09
0.16
Math
0.32
0.00
0.20
Appendix
Table 10: JGA on state tracking. Full supplies constraints together; Shared distributes them across turns; REA uses the distributed setting.
Language
System
Avg
Memory
Knowledge
Persona
Emotion
Morality
Believability
Chinese
Vanilla
3.64
3.93
3.58
4.25
3.15
4.76
2.21
REA
3.80
4.05
3.80
4.53
3.37
4.80
2.24
English
Vanilla
3.72
3.81
3.47
4.34
3.10
4.78
2.79
REA
4.05
4.22
3.91
5.00
3.49
4.90
2.76
Appendix
Table 11: Complete CharacterBench results. Avg retains the aggregate reported for the evaluation, rather than being recomputed from rounded aspect scores.
Figure 6: Heatmap of scenario-specific Acc performance (Follow-up and Refinement subsets).
Turn
User Query
Vanilla (Full History)
REA (Ours)
Instruction
Begin all your responses in the upcoming conversation with the letter r . Keep all your responses under 200 words.
Turn 1
Did Cleopatra belong to an Egyptian lineage?
R egarding your query, yes, Cleopatra was indeed an Egyptian pharaoh… ( Success )
R : Yes, Cleopatra was indeed an Egyptian queen. She was a member of the Ptolemaic dynasty… ( Success )
Turn 4
Write an email asking for donations for a charity.
Subject: Make a Difference in a Child’s Life… Dear [Recipient], I hope this email… ( Failure )
Subject: Reach Out and Help Underprivileged Kids Today! Dear Valued Supporter… ( Failure )
Turn 10
Create a product description for a DIY pottery kit.
Introducing our Delightful DIY Pottery Kit, the perfect project for creative minds… ( Failure )
R evel in the joy of creating with our Rustic Resilience DIY Pottery Kit! Each set includes… ( Success )
Appendix
Table 12: Qualitative comparison of instruction adherence across history, task-oriented, and creative domains. Vanilla fails as history accumulates, while REA anchors global constraints in Instructional Memory.
Turn
Vanilla (Full History) Prompt
REA (Ours) Reconstructed Prompt
Turn 1
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok.</s> [INST] Did Cleopatra belong to an Egyptian lineage? [/INST] Count: ≈ 65 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Did Cleopatra belong to an Egyptian lineage? [/INST]
Turn 4
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok…</s> [INST] … [/INST] …</s> [INST] … [/INST] …</s> [INST] Write an email asking for donations for a charity organization that assists underprivileged kids. [/INST] Count: ≈ 479 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Write an email asking for donations for a charity organization that assists underprivileged kids. [/INST]
Turn 10
[INST] You are a helpful, respectful and honest assistant. Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. [/INST] ok…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST]…[/INST]…</s> [INST] Create a product description for a DIY pottery kit. [/INST] Count: ≈ 2,369 tokens
[INST] <MEM0><MEM1>…<MEM7><SEP> Instruction: Begin all your responses in the upcoming conversation with the letter r. Keep all your responses under 200 words. Background: Question: Create a product description for a DIY pottery kit. [/INST]
Appendix
Table 13: Comparison of the internal prompts sent to the generator. Vanilla includes accumulated history; the illustrated REA inputs explicitly retain the global instruction.
Case / turn
Interaction
Memory / retrieval decision
Outcome
Success / 0
Set global story constraints.
Identify and retain constraints in IM.
Both systems acknowledge.
Success / 1–4
Rewrite requests with a 50-word limit and forbidden words.
Use low-resolution episodic context, τlow≤si≤τhigh .
Both systems follow the constraints in these turns.
Success / 5
Rewrite as a poem.
Combine instruction prefix with selected history.
REA follows the poem request and earlier constraints; Vanilla exceeds 50 words and misses earlier constraints.
Failure / 1
Extract all characters and locations.
Store the completed exchange in EM.
REA extracts both lists correctly.
Failure / 2
“List them in the order they appear.”
Turn 1 scores below τlow and is omitted from this query’s context.
REA lists characters but omits locations.
Appendix
Table 14: Qualitative summaries with HCR decisions. The failure illustrates a cross-turn reference that is not recovered by turn-level similarity.