Query Expansion (QE) techniques have long been widely used in Information Retrieval (IR) to address the vocabulary mismatch problem. They remain relevant in modern retrieval systems, including those based on large language models (LLMs). However, no single QE method consistently outperforms others across all queries. This work seeks to explain the variation in QE performance through two complementary perspectives. The first is the concept of an Ideal Expanded Query (IEQ)--a hypothetical query that maximizes retrieval effectiveness with a downstream BM25 retrieval model. The second is a separability perspective, which quantifies how distinctly relevant and non-relevant documents are scored for a given expanded query using Cohen's (d). We develop a separability measure and practical formulations to approximate the IEQ and investigate how these factors relate to retrieval effectiveness. Extensive experiments on the TREC Robust collection, TREC DL 2019-2022 passage collections, and TREC DL 2019-2020 document collections reveal several interesting patterns. In particular, we find that expanded queries that are closer to the ideal expanded query tend to achieve higher retrieval effectiveness. We further show that the separability of relevant and non-relevant documents provides a complementary perspective for understanding QE performance.
Figures & tables
Method
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
QRocchio
0.6945
0.8096
0.6916
0.5383
QIEQDFO
0.8652
0.9240
0.8752
0.8372
QIEQLS(λ=0.1)
0.9510 †
0.9667 †
0.8296 †
0.7456 †
QIEQLS(λ=1)
0.8975 †
0.9520 †
0.8747
0.7415 †
Truncated to 200 terms
QIEQLS(λ=0.1)
0.8757
0.9601
0.8241
0.7081
Table 2 . Retrieval effectiveness (MAP) achieved by the ideal queries generated using DFO and LS, on the Robust, DL19-20 Passage, DL19-20 Document, and DL21-22 Passage collections. For reference, the initial Rocchio vector generated using full relevance assessments is also included. The highest MAP obtained for each collection is shown in bold. Superscript † denotes statistically significant difference with DFO, using t-test ( p < 0.05).
Figure 1 . Per-query scatter plots of AP versus the number of relevant documents for three different ideal-query formulations on the Robust collection. The Spearman rank correlation coefficient rs between these two variables is reported in each plot. AP vs number of reldocs plot Per-query scatter plots of AP versus the number of relevant documents for different ideal-query formulations on the Robust collection. The Spearman rank correlation coefficient rs between the two variables is reported in each plot.
Dataset
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
QIEQDFO∩QIEQLS(λ=0.1)
99.34 [0.5659]
115.71 [0.6477]
89.33 [0.5315]
130.33 [0.6416]
QIEQDFO∩QIEQLS(λ=1)
131.55 [0.6767]
133.52 [0.7755]
100.99 [0.6027]
153.67 [0.7421]
QIEQLS(λ=0.1)∩QIEQLS(λ=1)
155.19 [0.8659]
162.65 [0.9103]
176.42 [0.9336]
164.12 [0.8847]
Table 3 . Average term overlap between ideal queries, computed across all queries. For each ideal query variant, only the top 200 positively weighted terms are considered. Cosine similarity between the ideal queries is reported in square brackets.
Dataset
Pearson ρ
Spearman rs
Robust
0.6783
0.4497
DL19-20 Passage
0.8823
0.8349
DL19-20 Document
0.7749
0.5311
DL21-22 Passage
0.8834
0.8512
Table 4 . Mean Pearson and Spearman correlation coefficients between the similarity of synthetically generated expanded queries to the ideal query QIEQLS(λ=0.1) , and their retrieval effectiveness, across all collections.
Figure 2 . Box plot of Pearson and Spearman correlations between the similarity of synthetically generated expanded queries to the ideal query QIEQLS(λ=0.1) and their retrieval effectiveness across all collections. box plot for synthetic queries Box plot of Pearson and Spearman correlations across the four different collections.
Figure 3 . Y-axis shows the AP variation and the X-axis plots the angle from the ideal query vector for the synthetically generated versions of the query ‘black bear attacks’ (Robust qid: 336) . ap vs theta Y-axis shows the AP variation and the X-axis plots the angle from the ideal query vector for the synthetically generated expansion terms of the query \emph{`black bear attacks' (Robust qid: 336)}.
Method
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
ρ
rs
ρ
rs
ρ
rs
ρ
rs
QIEQDFO
0.4366
0.4185
0.4762
0.4488
0.4245
0.4030
0.6774
0.5413
QIEQLS(λ=0.1)
0.5997
0.5770
0.4110
0.4097
0.4260
0.4390
0.4916
0.4360
QIEQLS(λ=1)
0.6053
0.5716
0.5078
0.4965
0.4680
0.4685
0.6395
0.5459
Table 5 . Mean Pearson and Spearman correlation coefficients between the similarity of real expanded queries to different ideal query formulations and the AP of the expanded queries on the Robust, DL19-20 Passage, DL19-20 Document, and DL21-22 Passage collections.
Dataset
Pearson’s ρ
Spearman’s rs
Robust
0.7849
0.7359
DL19-20 Passage
0.6504
0.6417
DL19-20 Document
0.6503
0.6241
DL21-22 Passage
0.7024
0.5971
Table 6 . Mean correlation (Pearson’s ρ and Spearman’s rs ) between the Separability measure and actual AP across all expanded queries on the Robust, DL19-20 Passage, DL19-20 Document, and DL21-22 Passage collections.
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
Method
Avg.
Max.
Avg.
Max.
Avg.
Max.
Avg.
Max.
QIEQDFO
0.3790
0.4608
0.6770
0.6693
0.4146
0.3556
0.5043
0.5378
QIEQLS(λ=0.1)
0.7393
0.7150
0.7379
0.6344
0.4547
0.4507
0.3059
0.2698
QIEQLS(λ=1)
0.7421
0.6915
0.7656
0.6400
0.4961
0.4647
0.3811
0.3617
Table 7 . Pearson correlation between aggregate AP and similarity for expanded queries. Three different ideal query representations are considered. For each collection, correlations are reported for the average and maximum similarity. The highest correlation within each collection and aggregation setting is shown in bold.
Dataset
ρ
rs
Robust
0.9689
0.9448
DL19-20 Passage
0.8560
0.8263
DL19-20 Document
0.8252
0.8203
DL21-22 Passage
0.7774
0.6767
Table 8 . Correlation (Pearson’s ρ and Spearman rs ) between restricted AP and actual AP across all expanded variants of a query, averaged over all queries in a collection.
Figure 4 . Trajectories along the great circles constructed from QIEQ and each Qexp , along which synthetic queries are generated. great circle trajectories Trajectories along the great circles constructed from QIEQ and each Qexp, along which synthetic queries are generated.
Figure 5 . Variation in Restricted AP along great circles originating from QIEQLS(λ=0.1) and passing through 500 synthetic queries generated from expanded queries produced by RM3, Bo1, KL, CEQE, HyDE, and HyDE-Rocchio.
Figure 6 . Query-specific variation (only for qid 730539 of the DL19-20 Document collection) of Restricted AP along 90 great circles (originating from QIEQLS(λ=0.1) and passing through 90 real expanded versions, generated using RM3, Bo1, KL, CEQE, HyDE, and HyDE-Rocchio methods, at different parameter settings). ap vs theta The angular deviation from the ideal query vector for the synthetically generated EQs is plotted along the X-axis, while the Y-axis corresponds to Restricted AP
Method
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
QIEQDFO
10.02∘
11.15∘
9.86∘
9.49∘
QIEQLS(λ=0.1)
3.75∘
5.42∘
4.04∘
4.92∘
QIEQLS(λ=1)
5.69∘
7.30∘
4.58∘
7.93∘
Average ΔAP
0.2215
0.2853
0.2624
0.1750
Table 9 . Average Δθ and ΔAP for all collections.
Figure 7 . Per-query variation in AP across different expanded queries (RM3, Bo1, KL, CEQE, HyDE and HyDE-Rocchio) for the benchmark queries in the DL19–20 Passage collection.
Figure 8 . Per-query variation in angular distance from QIEQLS(λ=1) across different expanded queries (RM3, Bo1, KL, CEQE, HyDE and HyDE-Rocchio) for the benchmark queries in the DL19–20 Passage collection.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 . Per-query scatter plots of AP versus the number of relevant documents for different ideal-query formulations on the DL19-20 passage collection. The Spearman rank correlation coefficient rs between the two variables is reported in each plot.
Figure 10 . Per-query scatter plots of AP versus the number of relevant documents for different ideal-query formulations on the DL19-20 document collection. The Spearman rank correlation coefficient rs between the two variables is reported in each plot.
Figure 11 . Per-query scatter plots of AP versus the number of relevant documents for different ideal-query formulations on the DL21-22 passage collection. The Spearman rank correlation coefficient rs between the two variables is reported in each plot. One outlier query is removed.
Figure 12 . Relationship between the MAP values and λ for LS based ideal queries. map vs lambda plot Relationship between the MAP values and λ for LS based ideal queries.
Figure 13 . Relationship between the average overlap with DFO based ideal query the LS based ideal queries generated with different λ . avg overlap vs lambda plot Relationship between the MAP values and λ for LS based ideal queries.
Method
Robust
DL19-20 Passage
DL19-20 Document
DL21-22 Passage
QIEQDFO
15.79
15.95
17.74
9.65
QIEQLS(λ=0.1)
14.37
6.42
9.20
10.38
QIEQLS(λ=1)
11.44
8.89
9.13
8.76
AP
8.66
7.07
6.76
6.85
Appendix
Table 10 . BQV/WQV ratios for both angular distances and AP for all collections. Larger BQV/WQV ratios indicate that the separation between query groups is substantially greater than the variability within each group.
LLM-based query expansion improves retrieval by enriching the original query with additional context. Yet most methods remain generation-driven, producing plausible pseudo-documents or expansions without checking how the target corpus responds. This can introduce retrieval drift, amplify misleading vocabulary, or miss terms that distinguish relevant from non-relevant documents. We argue that effective expansion requires retrieval-grounded feedback, not just single-pass generation or unverified iteration. We introduce ADORE (ADapt, Observe, Relevance Evaluate), an iterative framework that turns retrieval outcomes into feedback for the next expansion. At each round, an LLM generates pseudo-passages, a retriever exposes the corpus response, and a relevance assessor evaluates retrieved documents against the original query. These judgments identify what to reinforce, what remains undercovered, and what to suppress. Across TREC Deep Learning, BEIR, and BRIGHT, ADORE consistently outperforms strong query expansion baselines with notable improvements across nearly all evaluation settings, improving average nDCG@10 by 24.5% over BM25 and 3.6% over the strongest prior query expansion method on BEIR, and by 122.9% over BM25 and 9.2% over the best query expansion baseline on BRIGHT. Our code and data are publicly available.
Amin Bigdeli, Negar Arabzadeh, Radin Hamidi Rad +3
University of Waterloo · University of California, Berkeley · Mila – Quebec AI Institute +1
Lexical retrieval (BM25) captures exact keyword matches and weights terms by corpus-wide significance, but it is blind to the semantic vocabulary gap: when a relevant document phrases an answer differently from the query, BM25 never retrieves it, and no amount of downstream reranking or fusion can recover a document that was never in the candidate set. We present Cross-Encoder Query Expansion (CE-QE), which reads the per-token relevance attributions of a cross-encoder applied to top semantic search results, selects the terms the cross-encoder treats as decisive, and appends them to the BM25 query. Unlike classical pseudo-relevance feedback, which reuses BM25's own (possibly wrong) top results, CE-QE seeds expansion from the semantic retriever's results, avoiding self-reinforcing query drift. Unlike recent generative query expansion (HyDE, Query2doc), which prompts a large language model to hallucinate text from its parametric knowledge, every CE-QE expansion term is copied verbatim from a retrieved passage, so it cannot introduce vocabulary the corpus does not contain, and its only added cost is attribution extraction on a cross-encoder a hybrid pipeline already runs for reranking. On seven BEIR datasets, CE-QE improves lexical recall substantially where query and answer vocabulary diverge (e.g., NQ Recall@100 from 0.32 to 0.47), and its score-fusion variant (SESF) beats cross-encoder score fusion by 2.5% on Recall@100 and beats SPLADEv2 and ColBERTv2 by 5.3% and 4.6% on nDCG@10, while leaving the underlying BM25 index completely unmodified.
Adam Kahirov, Umesh Deshpande, Swaminathan Sundararaman
Retrieval-augmented generation (RAG) systems rely on retrieval modules to ground large language model (LLM) outputs. LLM-based query expansion enriches retrieval with document-like passages, but evaluations of hybrid retrieval often fuse fixed top-L prefixes of dense and sparse rankings. Because L controls cross-channel contributions and ranking access, it can alter measured expansion gains. We therefore evaluate complete-list effectiveness and record per-channel replay stopping depths required to certify the ordered top-K. This changes the design: because both rankings determine the fused result, their query constructions should be coordinated rather than designed independently. We present DESA (Dense Expansion and Sparse Anchoring), which shares generated references across channels but specializes their integration. Orthogonal residual expansion adds new semantic directions to the dense query, whereas score-product anchoring reorders the original sparse support without admitting expansion-only matches. The same references thus play complementary roles: Dense expands; Sparse anchors. Across seven BEIR datasets, DESA improves nDCG@10 and Recall@20 over the unexpanded query by 3.82% and 2.38%, while reducing dense and sparse replay stopping depths by 36.90% and 36.56%.
Chunran Zhang
School of Computing and Artificial Intelligence Southwest Jiaotong University Chengdu, China