Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-k selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.
Figures & tables
Figure 1 : The paper’s central result: coverage of three policies on HotpotQA fullwiki as measured pool redundancy grows under controlled injection. Nearest-neighbor selection, the framework default, collapses past the shaded crossover bracket (measured redundancy between 0.0012 and 0.0030 , Section 6.2 ). Diversifying hurts on clean pools and the decision rule of Section 4.3 tracks the upper envelope of both.
Figure 2 : ( 2(a) ) The open lune of x and y (shaded) is the intersection of the two open balls of radius d(x,y) centered at x and y ; z1 lies inside it and obstructs the pair {x,y} , so that pair is not an edge of the RNG , while z2 does not. ( 2(b) ) The RNG of a planar point set, with an edge between exactly the pairs whose lune is empty.
Figure 3 : One candidate pool, three computed top- 5 selections (filled blue = selected, open = unselected, ■=q ). ( 3(a) ) k -NN spends the whole budget on a tight cluster of near-duplicates; ( 3(b) ) the RNG-Score keeps one cluster representative and disperses to nearby distinct points; ( 3(c) ) MMR pushes further out onto less relevant points.
rank
k -NN
RNG-Score ( γ=0 )
MMR ( λ=0.5 )
1
1925 Birthday Honours
1925 Birthday Honours
1925 Birthday Honours
2
1915 Birthday Honours
David Lloyd George
Edward VII
3
1924 Birthday Honours
1915 Birthday Honours
George Champagne
4
1926 Birthday Honours
Henry V, Holy Roman Emperor
Julius Meinl V
5
1927 Birthday Honours
George V
A. Thynn, Marquess of Bath
Table 1: A HotpotQA fullwiki query needing two pieces of evidence ( 1925 Birthday Honours and George V , in bold).
Method
Recall@5
NDCG@5
α -NDCG@5
S-Recall@5
APD
Vendi
EM
F1
k -NN
74.8
73.9
76.5
78.7
0.450
3.08
50.8
63.6
MMR ( λ=0.3 )
45.3
53.3
55.4
48.1
0.630
4.00
36.4
47.8
MMR ( λ=0.5 )
55.6
60.3
62.5
58.6
0.574
3.73
41.6
53.5
MMR ( λ=0.7 )
69.2
70.0
72.5
72.8
0.505
3.38
47.6
59.9
MMR ( λ=0.9 )
74.1
73.5
76.1
78.0
0.463
3.16
50.3
62.9
Maxmin
42.9
51.7
53.8
45.7
0.625
3.85
36.0
47.4
Table 2: Controlled reranking on HotpotQA (MDR top-100 pools, bge-m3 , m=100 , k=5 ; clean level of the sweep of Section 6.2 , on its grids). The four left metrics are objectives (best in bold, second best underlined); APD and Vendi are descriptive spread statistics, not objectives; EM and F1: one deterministic Qwen3.8-27B run, seed- 0 test split ( 5,924 queries).
Method
SciFact
FiQA
TREC-COVID
ArguAna
Touché
k -NN
64.1 / 64.1
41.7 / 46.7
66.1 / 70.2
38.3 / 38.3
25.2 / 27.8
MMR
52.9/52.9
29.6/33.0
37.0/40.3
0.2/0.2
10.9/12.9
Maxmin
49.4/49.4
27.2/30.2
46.7/52.0
0.9/0.9
9.2/11.7
Greedy DPP
56.0/56.0
30.9/34.4
39.8/43.8
1.0/1.0
11.0/13.2
RNG-Score
63.4 / 63.5
41.7 / 46.8
65.6 / 69.6
26.4 / 26.4
17.8 / 19.6
Table 3: NDCG@10 / α -NDCG@10 on five BEIR tasks ( bge-m3 , k=10 ; MMR at λ=0.5 ; best in bold, second best underlined). SciFact, FiQA and TREC-COVID are seed-replicated; conclusions on these small query sets are exploratory (Appendix D ). ArguAna and Touché are single-run.
Dataset
Method (post-CE)
Recall@5
NDCG@5
EM
F1
HotpotQA
CE top- k
89.6
87.9
41.0
53.3
CE+ MMR
84.6
83.4
38.2
50.0
CE+ DPP
80.0
80.8
36.2
47.7
CE+ RNG-Score
89.6
87.9
41.0
53.3
2WikiMultiHopQA
CE top- k
90.6
90.3
31.3
37.9
CE+ MMR
71.0
76.3
25.0
29.6
Table 4: Cross-encoder pipelines, k=5 ( MMR at its best λ=0.7 ; CE+ RNG-Score is the tuned dissimilarity variant of Appendix B ; the non-HotpotQA datasets are single-run). Best per dataset and column in bold; a tie occupies both ranks, so cells below a tied best are not underlined.
ρ=0 (clean)
ρ=1 (redundant)
Method
best
worst
best
worst
max. downside
minimax regret
k -NN
78.7
52.6
22.0
22.2
MMR
78.0
48.1
74.6
53.5
30.6
6.3
Maxmin
45.7
45.7
50.6
50.6
33.0
33.6
Greedy DPP
69.1
69.1
72.1
72.1
9.6
10.0
RNG-Score
79.0
74.8
65.2
52.8
3.9
9.5
Table 5: Misspecification downside ( 5.1 ) and common-oracle minimax regret ( 5.2 ) on HotpotQA, with the best and worst S-Recall@5 over each method’s grid on clean and heavily redundant pools (the k -NN row of the downside column is constructed against the best tuned method per regime and is not directly comparable to the others). Lower is better in the last two columns. Query-bootstrap 95% intervals (computed, like the regret column, from query-level test means, hence a few tenths off the displayed seed-mean cells): downside MMR [30.2,31.5] , Maxmin [32.6,34.0] , DPP [9.2,10.3] , RNG-Score [3.5,4.4] ; regret k -NN [21.6,22.8] , MMR [5.9,6.8] , Maxmin [32.9,34.3] , DPP [9.6,10.6] , RNG-Score [8.9,10.0] . Best per column in bold, second best underlined.
Method
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
0
.0006
.0012
.0030
.0054
.012
.042
.147
k -NN
79.0
77.3
76.9
71.4
62.1
54.5
53.0
52.7
MMR ( λ∗ )
78.2
77.0
76.9
73.2
73.3
73.5
73.9
74.9
Maxmin
45.7
46.1
46.2
46.7
47.1
47.5
48.6
51.8
Greedy DPP
69.2
69.5
69.9
70.3
70.6
70.8
71.3
72.4
RNG-Score ( γ∗ )
79.2
77.8
77.8
76.5
74.7
71.7
68.1
65.4
Table 6: Redundancy-injection sweep, HotpotQA ( bge-m3 , k=5 ); S-Recall@5, operating points tuned per ρ . Second header row: measured pool redundancy ( 4.1 ). Best per column in bold, second best underlined.
Figure 4 : The HotpotQA injection sweep under three encoders: S-Recall vs measured near-duplicate pair fraction ( 4.1 ), on common axes (log-scaled with the clean level kept on-axis).
Figure 5 : The BEIR injection sweeps ( bge-m3 , k=10 , three seeds): S-Recall@10 vs measured near-duplicate pair fraction ( 4.1 ), on common axes. Single-evidence SciFact ( 5(a) ) stays flat across the sweep; multi-evidence FiQA-2018 ( 5(b) ) reproduces the multi-hop collapse-and-crossover pattern.
Sweep
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
HotpotQA, bge-m3
+0.2
−0.3
−0.3
−0.2
−0.2
−0.2
−0.2
−0.2
HotpotQA, Qwen3-Embedding-4B
+0.3
±
±
−0.2
−0.1
−0.1
−0.1
−0.1
HotpotQA, all-MiniLM-L6-v2
+0.3
+0.3
+0.2
+0.2
+0.2
±
−0.1
−0.1
SciFact, bge-m3
−0.2
±
±
−0.3
±
±
±
±
FiQA-2018, bge-m3
−0.2
±
−0.2
−0.2
−0.2
−0.2
−0.2
−0.2
Table 7: Validation-selected RNG-Score margins γ∗ per injection level: modal value across the three seeds where the seeds agree in sign, ± where they do not.
Method
HotpotQA
2Wiki
MuSiQue
SciFact
NQ-Open
TREC-COVID
hops
2
2
2–4
1
1
sat.
red .75
.001
.011
.015
.002
.005
.006
k -NN
78.1 /75.0
88.5 /87.4
74.1/73.2
77.2 /73.2
92.1/89.7
21.8 /23.1
Maxmin
46.5/48.3
74.7/83.6
56.9/69.5
57.8/62.2
86.8/87.6
14.9/21.9
Greedy DPP
69.5/70.9
87.2/ 94.4
76.2/ 86.1
73.3/74.1
91.6/92.0
18.6/25.9
MMR ∗
77.6/ 76.6
88.2/91.5
76.3 /85.0
77.0 / 76.1
92.8 / 93.0
21.4 / 27.2
Table 8: Chunk-overlap sweep ( bge-m3 , k=5 ; k=10 for SciFact and TREC-COVID; 2Wiki abbreviates 2WikiMultiHopQA). Each cell reports the row method’s S-Recall@ k on the non-overlapping pool / at overlap 0.75 (tuned operating points per level); the header rows give each dataset’s hop count and measured redundancy ( 4.1 ) at overlap 0.75 . Best per dataset in bold, second best underlined (ties both bold and occupying both ranks). The NQ-Open sweep is single-run; conclusions on the small SciFact and TREC-COVID query sets are exploratory (Appendix D ); saturated TREC-COVID (mean 493 relevant documents per query) has no hop count.
Figure 6 : Oracle headroom: percentage of test queries on which the all-methods oracle beats k -NN (HotpotQA, 2Wiki and MuSiQue at k=5 ; SciFact, FiQA-2018 and TREC-COVID at k=10 ; three seeds throughout). ( 6(a) ) Injection sweeps vs injected fraction ρ (log-scaled axis with the clean level kept on-axis); ( 6(b) ) Chunk-overlap sweeps vs overlap.
Figure 7 : The decision rule on the HotpotQA sweep. ( 7(a) ) Pooled validation S-Recall@5 as the trigger threshold τ varies continuously. ( 7(b) ) The frozen rule per injection level (log-scaled ρ axis, ticks at the sweep levels).
Selector
ρ=0
0.025
0.05
0.1
0.15
0.25
0.5
1.0
pool
always k -NN
78.7
77.0
76.7
71.1
62.0
54.5
52.9
52.6
65.7
always D∗
72.8
72.7
73.0
73.0
73.1
73.3
73.7
74.6
73.3
rule
78.7
76.9
76.0
72.9
73.1
73.2
73.5
74.3
74.8
trigger %
0.7
3.3
15.5
81.0
87.8
90.0
91.4
92.9
Reference bounds (not deployable policies)
per-level tuned
78.9
77.6
77.5
76.2
74.5
73.3
73.7
74.6
75.8
Table 9: The frozen rule ( τ=h=2 , D∗=\textscMMR(0.7) ) per injection level, S-Recall@5. “Per-level tuned” knows the regime; the oracle is the all-methods per-query bound of Section 6.3 ; the trigger rate is the fraction of queries on which the rule diversifies. Among the three policies, best in bold, second best underlined.
Transfer target
clean Δ
heavy Δ
pooled Δ
Different encoder (injection)
Qwen3-Embedding-4B
−0.3
+22.5
+9.8
all-MiniLM-L6-v2
−0.2
+13.6
+5.9
Different k (injection)
HotpotQA ( k=10 )
−0.0
+24.1
+5.7
Different mechanism (chunk overlap)
Table 10: Transfer of the frozen rule ( τ=min{h,k} with gold h , an oracle analysis; D∗=\textscMMR(0.7) ) to further sweeps with no retuning, grouped by what each sweep changes relative to the tuning sweep ( bge-m3 HotpotQA injection). Each Δ is the rule’s S-Recall@ k minus k -NN’s at the same sweep level, averaged over seeds (positive favors the rule): clean Δ is the worst Δ over the near-clean levels ( ρ≤0.05 on injection sweeps, overlap ≤0.25 on chunking sweeps), heavy Δ the Δ at the heaviest level ( ρ=1 , respectively overlap 0.75 ) and pooled Δ the difference with all levels pooled.
Figure 8 : Exact match vs injected redundancy on ( 8(a) ) HotpotQA and ( 8(b) ) 2WikiMultiHopQA, generator flan-t5-base . Log-scaled ρ axis with ticks at the sweep levels.
HotpotQA
2WikiMultiHopQA
Method
ρ=0
ρ=1
ρ=0
ρ=1
k -NN
50.8 / 63.6
41.7/53.6
63.4 / 71.8
49.5/56.5
MMR ( λ=0.7 )
47.6/59.9
47.8 / 60.5
59.3/67.5
55.5 / 63.5
rule ( τ=h )
50.8 / 63.6
47.8 / 60.4
63.4 / 71.9
55.4 / 63.4
Table 11: Answer quality (EM/F1) at the clean and heaviest injection levels, generator Qwen3.8-27B (a single run per dataset). Best in bold, second best underlined.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Mechanism
MMR ∗
Maxmin
Greedy DPP
RNG-Score
HotpotQA
injection
+23.0 ( +22.5 )
+32.4 ( +31.8 )
+29.5 ( +28.9 )
+12.5 ( +12.1 )
HotpotQA
chunk overlap
+2.1 ( +1.8 )
+4.8 ( +4.4 )
+4.4 ( +4.0 )
+3.0 ( +2.7 )
2WikiMultiHopQA
injection
+22.0 ( +21.7 )
+27.4 ( +26.8 )
+24.3 ( +23.9 )
+12.1 ( +11.8 )
2WikiMultiHopQA
chunk overlap
+4.3 ( +4.0 )
+9.9 ( +9.5 )
+8.3 ( +7.9 )
+6.0 ( +5.7 )
MuSiQue
injection
+23.8 ( +23.0 )
+31.6 ( +30.3 )
+27.5 ( +26.6 )
+17.3 ( +16.5 )
MuSiQue
chunk overlap
+9.6 ( +8.7 )
+13.5 ( +12.2 )
+10.8 ( +9.9 )
+9.9 ( +9.1 )
Appendix
Table 12: Primary test: heavy-minus-clean change of each method’s paired S-Recall effect against k -NN, in percentage points (one-sided 95% lower bound in parentheses). All entries are significant after Holm adjustment except the one marked † .
Dataset
red 0
k -NN
Maxmin
DPP
MMR ∗
RNG-Score ∗
2WikiMultiHopQA
0.0006
85.7
69.3
84.4
85.5
85.8
MuSiQue
0.0011
69.2
48.4
71.3
71.3
71.7
NQ-Open
0.0007
92.1
86.8
91.6
92.8
92.9
Appendix
Table 13: Single-stage reranking on the natural pools of the remaining QA datasets (S-Recall@5, bge-m3 , single run each; red 0 = measured pool redundancy ( 4.1 ); tuned operating points λ=0.9/0.9/0.7 and γ=−0.3/0.1/−0.1 for 2WikiMultiHopQA/MuSiQue/NQ-Open). Best in bold, second best underlined.