Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-k selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.
Figures & tables
Figure 1 : The paper’s central result: coverage of three policies on HotpotQA fullwiki as measured pool redundancy grows under controlled injection. Nearest-neighbor selection, the framework default, collapses past the shaded crossover bracket (measured redundancy between 0.0012 and 0.0030 , Section 6.2 ). Diversifying hurts on clean pools and the decision rule of Section 4.3 tracks the upper envelope of both.
Figure 2 : ( 2(a) ) The open lune of x and y (shaded) is the intersection of the two open balls of radius d(x,y) centered at x and y ; z1 lies inside it and obstructs the pair {x,y} , so that pair is not an edge of the RNG , while z2 does not. ( 2(b) ) The RNG of a planar point set, with an edge between exactly the pairs whose lune is empty.
Figure 3 : One candidate pool, three computed top- 5 selections (filled blue = selected, open = unselected, ■=q ). ( 3(a) ) k -NN spends the whole budget on a tight cluster of near-duplicates; ( 3(b) ) the RNG-Score keeps one cluster representative and disperses to nearby distinct points; ( 3(c) ) MMR pushes further out onto less relevant points.
rank
k -NN
RNG-Score ( γ=0 )
MMR ( λ=0.5 )
1
1925 Birthday Honours
1925 Birthday Honours
1925 Birthday Honours
2
1915 Birthday Honours
David Lloyd George
Edward VII
3
1924 Birthday Honours
1915 Birthday Honours
George Champagne
4
1926 Birthday Honours
Henry V, Holy Roman Emperor
Julius Meinl V
5
1927 Birthday Honours
George V
A. Thynn, Marquess of Bath
Table 1: A HotpotQA fullwiki query needing two pieces of evidence ( 1925 Birthday Honours and George V , in bold).
Method
Recall@5
NDCG@5
α -NDCG@5
S-Recall@5
APD
Vendi
EM
F1
k -NN
74.8
73.9
76.5
78.7
0.450
3.08
50.8
63.6
MMR ( λ=0.3 )
45.3
53.3
55.4
48.1
0.630
4.00
36.4
47.8
MMR ( λ=0.5 )
55.6
60.3
62.5
58.6
0.574
3.73
41.6
53.5
MMR ( λ=0.7 )
69.2
70.0
72.5
72.8
0.505
3.38
47.6
59.9
MMR ( λ=0.9 )
74.1
73.5
76.1
78.0
0.463
3.16
50.3
62.9
Maxmin
42.9
51.7
53.8
45.7
0.625
3.85
36.0
47.4
Table 2: Controlled reranking on HotpotQA (MDR top-100 pools, bge-m3 , m=100 , k=5 ; clean level of the sweep of Section 6.2 , on its grids). The four left metrics are objectives (best in bold, second best underlined); APD and Vendi are descriptive spread statistics, not objectives; EM and F1: one deterministic Qwen3.8-27B run, seed- 0 test split ( 5,924 queries).
Method
SciFact
FiQA
TREC-COVID
ArguAna
Touché
k -NN
64.1 / 64.1
41.7 / 46.7
66.1 / 70.2
38.3 / 38.3
25.2 / 27.8
MMR
52.9/52.9
29.6/33.0
37.0/40.3
0.2/0.2
10.9/12.9
Maxmin
49.4/49.4
27.2/30.2
46.7/52.0
0.9/0.9
9.2/11.7
Greedy DPP
56.0/56.0
30.9/34.4
39.8/43.8
1.0/1.0
11.0/13.2
RNG-Score
63.4 / 63.5
41.7 / 46.8
65.6 / 69.6
26.4 / 26.4
17.8 / 19.6
Table 3: NDCG@10 / α -NDCG@10 on five BEIR tasks ( bge-m3 , k=10 ; MMR at λ=0.5 ; best in bold, second best underlined). SciFact, FiQA and TREC-COVID are seed-replicated; conclusions on these small query sets are exploratory (Appendix D ). ArguAna and Touché are single-run.
Dataset
Method (post-CE)
Recall@5
NDCG@5
EM
F1
HotpotQA
CE top- k
89.6
87.9
41.0
53.3
CE+ MMR
84.6
83.4
38.2
50.0
CE+ DPP
80.0
80.8
36.2
47.7
CE+ RNG-Score
89.6
87.9
41.0
53.3
2WikiMultiHopQA
CE top- k
90.6
90.3
31.3
37.9
CE+ MMR
71.0
76.3
25.0
29.6
Table 4: Cross-encoder pipelines, k=5 ( MMR at its best λ=0.7 ; CE+ RNG-Score is the tuned dissimilarity variant of Appendix B ; the non-HotpotQA datasets are single-run). Best per dataset and column in bold; a tie occupies both ranks, so cells below a tied best are not underlined.
ρ=0 (clean)
ρ=1 (redundant)
Method
best
worst
best
worst
max. downside
minimax regret
k -NN
78.7
52.6
22.0
22.2
MMR
78.0
48.1
74.6
53.5
30.6
6.3
Maxmin
45.7
45.7
50.6
50.6
33.0
33.6
Greedy DPP
69.1
69.1
72.1
72.1
9.6
10.0
RNG-Score
79.0
74.8
65.2
52.8
3.9
9.5
Table 5: Misspecification downside ( 5.1 ) and common-oracle minimax regret ( 5.2 ) on HotpotQA, with the best and worst S-Recall@5 over each method’s grid on clean and heavily redundant pools (the k -NN row of the downside column is constructed against the best tuned method per regime and is not directly comparable to the others). Lower is better in the last two columns. Query-bootstrap 95% intervals (computed, like the regret column, from query-level test means, hence a few tenths off the displayed seed-mean cells): downside MMR [30.2,31.5] , Maxmin [32.6,34.0] , DPP [9.2,10.3] , RNG-Score [3.5,4.4] ; regret k -NN [21.6,22.8] , MMR [5.9,6.8] , Maxmin [32.9,34.3] , DPP [9.6,10.6] , RNG-Score [8.9,10.0] . Best per column in bold, second best underlined.
Method
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
0
.0006
.0012
.0030
.0054
.012
.042
.147
k -NN
79.0
77.3
76.9
71.4
62.1
54.5
53.0
52.7
MMR ( λ∗ )
78.2
77.0
76.9
73.2
73.3
73.5
73.9
74.9
Maxmin
45.7
46.1
46.2
46.7
47.1
47.5
48.6
51.8
Greedy DPP
69.2
69.5
69.9
70.3
70.6
70.8
71.3
72.4
RNG-Score ( γ∗ )
79.2
77.8
77.8
76.5
74.7
71.7
68.1
65.4
Table 6: Redundancy-injection sweep, HotpotQA ( bge-m3 , k=5 ); S-Recall@5, operating points tuned per ρ . Second header row: measured pool redundancy ( 4.1 ). Best per column in bold, second best underlined.
Figure 4 : The HotpotQA injection sweep under three encoders: S-Recall vs measured near-duplicate pair fraction ( 4.1 ), on common axes (log-scaled with the clean level kept on-axis).
Figure 5 : The BEIR injection sweeps ( bge-m3 , k=10 , three seeds): S-Recall@10 vs measured near-duplicate pair fraction ( 4.1 ), on common axes. Single-evidence SciFact ( 5(a) ) stays flat across the sweep; multi-evidence FiQA-2018 ( 5(b) ) reproduces the multi-hop collapse-and-crossover pattern.
Sweep
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
HotpotQA, bge-m3
+0.2
−0.3
−0.3
−0.2
−0.2
−0.2
−0.2
−0.2
HotpotQA, Qwen3-Embedding-4B
+0.3
±
±
−0.2
−0.1
−0.1
−0.1
−0.1
HotpotQA, all-MiniLM-L6-v2
+0.3
+0.3
+0.2
+0.2
+0.2
±
−0.1
−0.1
SciFact, bge-m3
−0.2
±
±
−0.3
±
±
±
±
FiQA-2018, bge-m3
−0.2
±
−0.2
−0.2
−0.2
−0.2
−0.2
−0.2
Table 7: Validation-selected RNG-Score margins γ∗ per injection level: modal value across the three seeds where the seeds agree in sign, ± where they do not.
Method
HotpotQA
2Wiki
MuSiQue
SciFact
NQ-Open
TREC-COVID
hops
2
2
2–4
1
1
sat.
red .75
.001
.011
.015
.002
.005
.006
k -NN
78.1 /75.0
88.5 /87.4
74.1/73.2
77.2 /73.2
92.1/89.7
21.8 /23.1
Maxmin
46.5/48.3
74.7/83.6
56.9/69.5
57.8/62.2
86.8/87.6
14.9/21.9
Greedy DPP
69.5/70.9
87.2/ 94.4
76.2/ 86.1
73.3/74.1
91.6/92.0
18.6/25.9
MMR ∗
77.6/ 76.6
88.2/91.5
76.3 /85.0
77.0 / 76.1
92.8 / 93.0
21.4 / 27.2
Table 8: Chunk-overlap sweep ( bge-m3 , k=5 ; k=10 for SciFact and TREC-COVID; 2Wiki abbreviates 2WikiMultiHopQA). Each cell reports the row method’s S-Recall@ k on the non-overlapping pool / at overlap 0.75 (tuned operating points per level); the header rows give each dataset’s hop count and measured redundancy ( 4.1 ) at overlap 0.75 . Best per dataset in bold, second best underlined (ties both bold and occupying both ranks). The NQ-Open sweep is single-run; conclusions on the small SciFact and TREC-COVID query sets are exploratory (Appendix D ); saturated TREC-COVID (mean 493 relevant documents per query) has no hop count.
Figure 6 : Oracle headroom: percentage of test queries on which the all-methods oracle beats k -NN (HotpotQA, 2Wiki and MuSiQue at k=5 ; SciFact, FiQA-2018 and TREC-COVID at k=10 ; three seeds throughout). ( 6(a) ) Injection sweeps vs injected fraction ρ (log-scaled axis with the clean level kept on-axis); ( 6(b) ) Chunk-overlap sweeps vs overlap.
Figure 7 : The decision rule on the HotpotQA sweep. ( 7(a) ) Pooled validation S-Recall@5 as the trigger threshold τ varies continuously. ( 7(b) ) The frozen rule per injection level (log-scaled ρ axis, ticks at the sweep levels).
Selector
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
pool
always k -NN
78.7
77.0
76.7
71.1
62.0
54.5
52.9
52.6
65.7
always D∗
72.8
72.7
73.0
73.0
73.1
73.3
73.7
74.6
73.3
rule
78.7
76.9
76.0
72.9
73.1
73.2
73.5
74.3
74.8
trigger %
0.7
3.3
15.5
81.0
87.8
90.0
91.4
92.9
Reference bounds (not deployable policies)
per-level tuned
78.9
77.6
77.5
76.2
74.5
73.3
73.7
74.6
75.8
Table 9: The frozen rule ( τ=h=2 , D∗=\textscMMR(0.7) ) per injection level, S-Recall@5. “Per-level tuned” knows the regime; the oracle is the all-methods per-query bound of Section 6.3 ; the trigger rate is the fraction of queries on which the rule diversifies. Among the three policies, best in bold, second best underlined.
Transfer target
clean Δ
heavy Δ
pooled Δ
Different encoder (injection)
Qwen3-Embedding-4B
−0.3
+22.5
+9.8
all-MiniLM-L6-v2
−0.2
+13.6
+5.9
Different k (injection)
HotpotQA ( k=10 )
−0.0
+24.1
+5.7
Different mechanism (chunk overlap)
Table 10: Transfer of the frozen rule ( τ=min{h,k} with gold h , an oracle analysis; D∗=\textscMMR(0.7) ) to further sweeps with no retuning, grouped by what each sweep changes relative to the tuning sweep ( bge-m3 HotpotQA injection). Each Δ is the rule’s S-Recall@ k minus k -NN’s at the same sweep level, averaged over seeds (positive favors the rule): clean Δ is the worst Δ over the near-clean levels ( ρ≤0.05 on injection sweeps, overlap ≤0.25 on chunking sweeps), heavy Δ the Δ at the heaviest level ( ρ=1 , respectively overlap 0.75 ) and pooled Δ the difference with all levels pooled.
Figure 8 : Exact match vs injected redundancy on ( 8(a) ) HotpotQA and ( 8(b) ) 2WikiMultiHopQA, generator flan-t5-base . Log-scaled ρ axis with ticks at the sweep levels.
HotpotQA
2WikiMultiHopQA
Method
ρ=0
ρ=1
ρ=0
ρ=1
k -NN
50.8 / 63.6
41.7/53.6
63.4 / 71.8
49.5/56.5
MMR ( λ=0.7 )
47.6/59.9
47.8 / 60.5
59.3/67.5
55.5 / 63.5
rule ( τ=h )
50.8 / 63.6
47.8 / 60.4
63.4 / 71.9
55.4 / 63.4
Table 11: Answer quality (EM/F1) at the clean and heaviest injection levels, generator Qwen3.8-27B (a single run per dataset). Best in bold, second best underlined.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Mechanism
MMR ∗
Maxmin
Greedy DPP
RNG-Score
HotpotQA
injection
+23.0 ( +22.5 )
+32.4 ( +31.8 )
+29.5 ( +28.9 )
+12.5 ( +12.1 )
HotpotQA
chunk overlap
+2.1 ( +1.8 )
+4.8 ( +4.4 )
+4.4 ( +4.0 )
+3.0 ( +2.7 )
2WikiMultiHopQA
injection
+22.0 ( +21.7 )
+27.4 ( +26.8 )
+24.3 ( +23.9 )
+12.1 ( +11.8 )
2WikiMultiHopQA
chunk overlap
+4.3 ( +4.0 )
+9.9 ( +9.5 )
+8.3 ( +7.9 )
+6.0 ( +5.7 )
MuSiQue
injection
+23.8 ( +23.0 )
+31.6 ( +30.3 )
+27.5 ( +26.6 )
+17.3 ( +16.5 )
MuSiQue
chunk overlap
+9.6 ( +8.7 )
+13.5 ( +12.2 )
+10.8 ( +9.9 )
+9.9 ( +9.1 )
Appendix
Table 12: Primary test: heavy-minus-clean change of each method’s paired S-Recall effect against k -NN, in percentage points (one-sided 95% lower bound in parentheses). All entries are significant after Holm adjustment except the one marked † .
Dataset
red 0
k -NN
Maxmin
DPP
MMR ∗
RNG-Score ∗
2WikiMultiHopQA
0.0006
85.7
69.3
84.4
85.5
85.8
MuSiQue
0.0011
69.2
48.4
71.3
71.3
71.7
NQ-Open
0.0007
92.1
86.8
91.6
92.8
92.9
Appendix
Table 13: Single-stage reranking on the natural pools of the remaining QA datasets (S-Recall@5, bge-m3 , single run each; red 0 = measured pool redundancy ( 4.1 ); tuned operating points λ=0.9/0.9/0.7 and γ=−0.3/0.1/−0.1 for 2WikiMultiHopQA/MuSiQue/NQ-Open). Best in bold, second best underlined.
Retrieval-augmented generation (RAG) systems typically rely on a single retriever and a single set of hyperparameters, despite facing highly heterogeneous queries that range from simple factoid questions to complex multi-hop reasoning. We propose a method that automatically selects a small, diverse subset of retrievers (a portfolio) from a large pool of candidates, to cover different regions of the target query distribution. We formalize this setting via an expected best-of-k objective over the query distribution and show that it admits an efficient portfolio construction algorithm with near-optimal guarantees. Across multiple QA benchmarks, our learned portfolios and router pipeline consistently outperform single-retriever and naive multi-retriever baselines on both retrieval metrics and answer quality. In addition, compared to inference-time hyperparameter tuning approaches, fixed portfolios enable parallel retrieval and LLM calls, achieving comparable (and sometimes better) accuracy with substantially lower latency and token cost.
Miltiadis Stouras, Vincent Cohen-Addad, Silvio Lattanzi +1
Retrieval augmented generation (RAG) depends critically on the quality and granularity of retrieved evidence. Large retrieval units preserve context but often introduce irrelevant content, which can dilute answer bearing evidence and worsen long context utilization. Fine-grained units are more compact, but they may be difficult to retrieve reliably because short chunks can lack semantic, lexical, or bridging cues needed to match the query. We propose Uncertainty-aware Multi-Granularity RAG (UMG-RAG), a training-free hybrid retrieval framework that treats chunk granularity as query-specific reliability estimation. Instead of training a new retriever or modifying the generator, UMG-RAG uses existing dense and sparse retrievers as complementary experts across multiple chunk granularities. For each query, it converts each expert-granularity score list into an evidence distribution, estimates reliability from distribution entropy, and fuses candidates according to query-specific semantic, lexical, and granularity confidence. We further introduce UMGP-RAG, a parent promotion variant that uses fine-grained hits to locate relevant evidence while returning broader non-redundant parent chunks for local coherence. Experiments on question answering benchmarks show that uncertainty-aware fusion and parent promotion improve generation quality while maintaining a lightweight, plug-and-play retrieval pipeline.
Hoin Jung, Xiaoqian Wang
Elmore Family School of Electrical and Computer Engineering Purdue University West Lafayette, IN 47907
Large Language Models (LLMs) have made query reformulation ubiquitous in modern retrieval and Retrieval-Augmented Generation (RAG) pipelines, enabling the generation of multiple semantically equivalent query variants. However, executing the full pipeline for every reformulation is computationally expensive, motivating selective execution: can we identify the best query variant before incurring downstream retrieval and generation costs? We investigate Query Performance Prediction (QPP) as a mechanism for variant selection across ad-hoc retrieval and end-to-end RAG. Unlike traditional QPP, which estimates query difficulty across topics, we study intra-topic discrimination - selecting the optimal reformulation among competing variants of the same information need. Through large-scale experiments on TREC-RAG using both sparse and dense retrievers, we evaluate pre- and post-retrieval predictors under correlation- and decision-based metrics. Our results reveal a systematic divergence between retrieval and generation objectives: variants that maximize ranking metrics such as nDCG often fail to produce the best generated answers, exposing a "utility gap" between retrieval relevance and generation fidelity. Nevertheless, QPP can reliably identify variants that improve end-to-end quality over the original query. Notably, lightweight pre-retrieval predictors frequently match or outperform more expensive post-retrieval methods, offering a latency-efficient approach to robust RAG.
Negar Arabzadeh, Andrew Drozdov, Michael Bendersky +1
UC Berkeley · Berkeley, CA, United States · Databricks +1