When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.
Figures & tables
Figure 1: Overview of a multi-turn retrieval simulation, shown at turn 3. The user simulator has just revealed the next clue ( u3 ). The retriever is then queried with either all user turns so far ( u1+u2+u3 ) or only the latest ( u3 ), depending on the retrieval policy. The LLM then answers from this retrieved evidence together with the conversation history, and the response classifier flags its output as an answer attempt. The text after the Answer: marker is scored against the gold answer. Here the predicted date is wrong, so the assistant receives a score of 0 F1.
Condition
Assistant Input
FULL
The original question, in one turn q
CONCAT
All shards as a list, in one turn u1u2u3
SHARDED
One shard per turn u1u2u3
RECAP
SHARDED , then one final turn restating every shard u1u2u3u1u2u3
SNOWBALL
SHARDED , with each turn repeating all earlier shards u1u1u2u1u2u3
Table 1: Information-delivery conditions. Each box represents one user turn.
FULL
SHARDED (history)
SHARDED (current)
Mean gap
Retriever
HP
2W
MS-2
MS-3
MS-4
HP
2W
MS-2
MS-3
MS-4
HP
2W
MS-2
MS-3
MS-4
Hist.
Curr.
HippoRAG2
75.9
69.0
49.0
37.6
21.6
68.0
57.9
40.5
34.0
18.6
65.8
51.7
39.2
29.3
17.7
+6.8
+9.9
RAPTOR
70.2
63.0
42.2
35.8
20.7
65.4
52.3
32.4
34.0
18.6
60.4
46.4
30.4
29.3
17.0
+5.8
+9.7
HippoRAG
66.6
71.5
42.3
27.7
19.9
62.6
61.0
38.3
25.4
16.8
60.5
54.3
37.1
25.6
16.1
+4.8
+6.9
Dense
71.6
60.1
37.8
37.0
20.5
65.5
48.6
31.0
33.9
17.9
60.4
45.8
28.3
29.3
17.3
+6.0
+9.2
ToG-2
67.0
52.3
46.1
34.9
21.5
62.0
48.2
40.0
32.1
21.3
62.5
49.2
34.7
26.1
17.1
+3.7
+6.4
Table 2: F1 ( ×100 ) under FULL and SHARDED, averaged over ten LLMs. Red shading indicates a loss and blue a gain relative to FULL; intensity reflects the absolute difference in F1 points. Rows are ordered by mean FULL score.
Figure 2: FULL vs. SHARDED F1 (points) under history queries. Each cell averages over the five datasets. Rows are ordered by mean FULL score across the seven retrieval methods. The five symbols beneath each value summarize dataset-level paired 95% bootstrap intervals for Δ=\textscFULL−\textscSHARDED : ∙ entirely above zero, † entirely below zero, and ∘ containing zero. Intervals excluding zero indicate a statistically significant difference at the 5% level.
FULL − CONCAT
CONCAT − SHARDED
Total
Retriever
FULL
HP
2W
MS-2
MS-3
MS-4
Mean
HP
2W
MS-2
MS-3
MS-4
Mean
FULL − SH.
BM25
36.5
+14.5
+21.6
+10.5
+5.8
+6.2
+11.7
-8.0
-6.4
-5.5
-5.1
-0.9
-5.2
+6.5
Dense
45.4
+8.3
+13.0
+7.7
+5.2
+1.9
+7.2
-2.3
-1.5
-0.9
-2.1
+0.8
-1.2
+6.0
RAPTOR
46.4
+7.6
+15.5
+10.0
+4.7
+2.6
+8.1
-2.8
-4.9
-0.1
-2.9
-0.5
-2.2
+5.8
HippoRAG2
50.6
+7.7
+9.3
+10.1
+6.2
+2.2
+7.1
+0.2
+1.8
-1.7
-2.6
+0.9
-0.3
+6.8
HippoRAG
45.6
+2.0
+4.3
+0.2
+0.6
-0.3
+1.4
+2.0
+6.2
+3.8
+1.8
+3.4
+3.4
+4.8
Table 3: Decomposition of the FULL – SHARDED F1 gap ( ×100 ) under HISTORY queries. All cells average over ten LLMs. FULL, Mean, and Total columns also average over the five datasets.
Single turn
SHARDED, history query
SHARDED, current-turn query
Retriever
FULL
CONCAT
per query
final turn
cumulative
per query
final turn
cumulative
BM25
45.6
18.6
19.5
30.3
35.3
22.4
35.5
51.7
Dense
62.6
51.9
35.3
54.4
58.9
29.9
43.3
65.6
HippoRAG2
70.7
63.0
40.9
64.6
69.6
35.0
52.5
74.8
HippoRAG
61.7
65.1
32.7
63.2
65.0
27.7
50.1
65.4
LightRAG
40.9
41.4
24.4
40.1
46.0
19.4
30.8
44.8
Table 4: Supporting-passage Recall@5 ( ×100 ), averaged over ten LLMs and five datasets. For SHARDED , per query is the mean over retrieval turns, final turn uses the last retrieval, and cumulative is the union of every turn’s top-five passages. Underlined entries exceed the row’s FULL recall. RAPTOR is omitted because its retrieval units consist of normal passages and their summary nodes.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
HotpotQA Who wrote the novel that the movie directed by Stanley Kubrick that was sampled in the album “Where Blood and Fire Bring Rest” was based on ? [Stephen King]
I’m trying to find out who the author of a novel is
it’s related to a movie directed by Stanley Kubrick
and it was sampled in the album ’Where Blood and Fire Bring Rest’
the novel was the basis for that movie
HotpotQA The world’s greatest Super-Heroes anthology showcased one of four superheroes known for speaking the phrase “SHAZAM” , what was their name ? [Captain Marvel]
I’m trying to find out the name of a superhero
Appendix
Table 5: Examples of shard construction across the five evaluated datasets. Highlighting marks the span of the FULL question that each user turn carries, and the coloured square on a turn matches its span. The order of the colored spans in the FULL question may differ from the order of the turns below, because shards are revealed according to their dependencies.
Figure 3: Distribution of user turns in the constructed shard sets. Each square is one of the 150 evaluated questions in a dataset.
Example According to the 2010 census, what was the population of the city after which the vice president, in April 1813, was named?
Turn 1 (intent)
I’m trying to find out the population of a city
Turn 2
it was based on the 2010 census 11 distinct / 15
Turn 2 (simulated)
✓ The population info is from the 2010 census .
✓ the population I mean was based on the 2010 census .
✓ The population was based on the 2010 census .
✓ It was based on the 2010 census .
Appendix
Table 6: Five realizations of one HotpotQA clue, sampled from 15 conversations with the GPT-4o-mini across 3 different retrieval configurations
Figure 4: Closed-book performance across different simulation conditions. The displayed models in Panel (b) are 5.6 Luna and the largest evaluated member of each open-weight model series: Llama-3.3-70B, Gemma-4-31B, Qwen3.6-27B, and Qwen3-32B.
Closed-book BEST F1
Relative to FULL
FULL
CONCAT
SHARDED
CONCAT/FULL
SHARDED drop (%)
Datasets, pooled over all ten LLMs
HotpotQA
.313
.292
.274
0.93
12%
2Wiki
.373
.320
.298
0.86
20%
MuSiQue-2hop
.192
.167
.153
0.87
20%
MuSiQue-3hop
.159
.145
.142
0.91
11%
Appendix
Table 7: Complete closed-book results. F1 is reported on a 0–1 scale.
Figure 5: Recovery of the FULL – SHARDED performance gap under RECAP and SNOWBALL , using the CURRENT retrieval query. Recovery is normalized such that 0% corresponds to SHARDED and 100% to FULL . Negative values indicate performance below SHARDED .
Figure 6: Mean F1 ( ×100 ) under all five conditions, averaged over the three LLMs and five datasets. The upper panel uses CURRENT retrieval queries and the lower panel uses HISTORY . Colors distinguish the two consolidation interventions.
Figure 7: Averaged F1 ( ×100 ) under FULL and both SHARDED policies. Rows are ordered by FULL score.
Figure 8: Aptitude and unreliability under multi-turn setting. (a) Without retrieval, aptitude falls while unreliability remains the same. With retrieval, aptitude holds but unreliability increases sharply. (b) Per-model and per-retriever breakdown, ordered by mean FULL score.
Full
Sharded
A
U
P
A
U
P
ΔA
ΔU
LLM Assistants (35 retrieval configurations each)
Llama-3.1-8B-Inst
42.1
23.6
29.9
40.5
29.1
25.3
-1.7
+5.4
Qwen3-4B
41.4
10.6
36.0
40.2
18.3
30.7
-1.3
+7.8
Qwen3-14B
44.0
9.9
39.0
37.5
15.5
29.5
-6.6
+5.7
Qwen3-8B
45.5
12.6
39.1
42.3
19.1
32.5
-3.2
+6.5
Appendix
Table 8: Aptitude A , unreliability U , and mean F1 P ( ×100 ) under FULL and SHARDED ( HISTORY ). Shading marks the change in U .
Dataset
Pairs
FULL
SHARDED
Gap
HotpotQA
21,107
78.5
67.0
11.5 ∙
2Wiki
17,578
77.5
68.4
9.0 ∙
MuSiQue-2hop
14,255
64.3
50.7
13.6 ∙
MuSiQue-3hop
1,723
57.9
35.3
22.5 ∙
MuSiQue-4hop
213
47.0
28.4
18.6
Appendix
Table 9: Paired F1 under complete supporting-source coverage across datasets ( HISTORY , six retrievers, RAPTOR excluded). ∙ : the paired bootstrap 95% interval of the gap lies above zero.
Figure 9: Paired F1 by LLM under complete supporting-source coverage. FULL – SHARDED difference reported on the right.
Token F1 (BEST)
LLM judge (BEST)
Retriever
Full
Sharded
Gap
Full
Sharded
Gap
HippoRAG2
53.3
47.4
+6.0
55.7
50.2
+5.5
RAPTOR
48.2
43.3
+4.9
50.3
45.4
+4.8
HippoRAG
47.2
44.5
+2.7
49.0
47.4
+1.6
Dense
47.1
41.5
+5.6
49.2
43.2
+6.0
ToG-2
46.0
44.5
+1.4
47.9
47.1
+0.8
Appendix
Table 10: F1 scores ( × 100) and binary judge accuracy under the HISTORY query policy, pooled over Llama-3.3-70B, Gemma-4-31B, and Qwen3-32B and the five datasets. Shading marks the size of the gap.
Figure 10: FULL – SHARDED gaps are consistent across metrics. Each point represents one model–dataset–retriever combination, comparing the gap under token F1 with the same gap under the binary judge.
Retriever
Indexing
Retrieval
Generation
Knowledge type
Index content
Query input
Granularity
Model context
NONE
Closed-book
—
—
—
Parametric only
BM25 ( Robertson & Zaragoza, 2009 )
Plain text
Frozen passages
Lexical terms
Passage
5 retrieved passages
Dense TextRAG ( Xiao et al., 2024 )
Plain text
Passage embeddings
Dense query
Passage
5 retrieved passages
HippoRAG ( Gutierrez et al., 2024 )
Phrase graph
Phrases, passages
Query phrases
Entity, passage
5 retrieved passages
HippoRAG2 ( Gutiérrez et al., 2025 )
Fact graph, passages
Facts, entities, passages
Query-to-fact
Fact, entity, passage
5 retrieved passages
Appendix
Table 11: Implementation details of the evaluated retrieval families.
Questions
Parent corpus
Corpus
Evaluated
Donor
Documents
Passages
Approx. tokens
HotpotQA
150
850
1,000
9,772
907,305
2WikiMultiHopQA
150
850
1,000
6,324
466,381
MuSiQue (2/3/4-hop)
450
550
1,000
11,693
966,093
Appendix
Table 12: Parent corpora used for retrieval.
Display name
Model ID / revision
∣θ∣
Access
GPT-4o-mini
gpt-4o-mini-2024-07-18
–
OpenAI API
5.6 Luna
gpt-5.6-luna
–
OpenAI API
Qwen3-4B
Qwen/Qwen3-4B
4B
Local vLLM
Qwen3-8B
Qwen/Qwen3-8B
8B
Local vLLM
Qwen3-14B
Qwen/Qwen3-14B
14B
Local vLLM
Qwen3-32B
Qwen/Qwen3-32B
32B
Local vLLM
Appendix
Table 13: Evaluated LLMs with versions and access details.
Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns. Yet recent work shows that LLMs perform far worse in this multi-turn setting than in a single turn with same information being available at once, a phenomenon termed "Lost-in-Conversation." However, bridging this gap effectively remains an open problem. Here we introduce Found in Conversation (FiC), a training framework where a model teaches itself to find and recover its single-turn competence given underspecified multi-turn prompts. We develop View-Asymmetric Self-Distillation, which distills across two views of the same task information--single-turn view for the teacher, multi-turn view for the student--transferring strong single-turn behavior into weak multi-turn behavior. This requires no stronger external teacher, which is unavailable as even frontier LLMs exhibit this gap. Across model families (Llama, Qwen, Phi, and OLMo) and sizes (3B-14B), FiC recovers at least 92% of single-turn performance and reaches 100% on two Llama backbones, yielding more efficient and helpful multi-turn conversations with single-turn capabilities intact.
We introduce 5ting, our system for the SemEval2026 Task 8 (MTRAGEval), which evaluates multi-turn Retrieval Augmented Generation (RAG) systems. Multi turn RAG involves context drift, under specification, and hallucination risk. Our system combines BGE-M3 dense retrieval with FAISS indexing, dual-query merged retrieval, and LLM based reranking, followed by role separated generation constrained to retrieved evidence. The retriever achieved nDCG@5 = 0.4719 in Task A, while the end to end system ranked in Task C with a harmonic score of 0.5597 and RL_F = 0.7692.
Thien-Qua-T-Nguyen, Chi Hoang, Nguyen Tran +3
University of Information Technology, Ho Chi Minh City, Vietnam · 2Vietnam National University Ho Chi Minh City, Ho Chi Minh City, Vietnam
Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history, rewritten queries, or unstructured summaries, making it difficult to recover the specific prior reasoning steps and evidence required for follow-up queries. Our key insight is to align conversational memory with retrieval by representing dialogue context as sub-question-level reasoning traces. Building on this insight, we introduce MuMu-QA, a benchmark for multi-turn multi-hop RAG with explicit cross-turn sub-question dependency annotations, and CMT-RAG, a complementary memory framework for this setting. At each turn, CMT-RAG employs a state-space trace generator, whose recurrent state serves as runtime memory, to incorporate recent conversational context and decompose the current query into structured trace drafts containing retrieval-oriented sub-questions and dependencies on earlier traces. It then grounds these drafts with retrieved evidence and stores them as persistent memory traces in a session-level DAG, enabling future turns to efficiently recover relevant prior reasoning and evidence. Experiments on MuMu-QA and corpus-level RAG benchmarks show that CMT-RAG consistently outperforms five categories of RAG baselines in answer accuracy.
Lang Zhou, Yingjian Chen, Shuxuan Li +2
Sun Yat-sen University · Shenzhen Loop Area Institude · The University of Hong Kong