Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion
Authors: Xinkai Du, Chao Lv, Yalin Sun, Quanjie Han, Lei Yao, Maosong Sun
Organizations: Beijing Wanlian Zhilian Technology Corporation Limited, Beijing, China · Department of Computer Science and Technology, Tsinghua University, Beijing, China · Sunshine Digital Intelligence Tech Co., Ltd., Beijing, China
Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semantic space mismatch between queries and retrieved contexts. We propose Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. Our approach leverages the complementary strengths of generative and retrieval-based knowledge through a multistage process that enhances both relevance and accuracy. We evaluate KASB on three popular open-domain Question Answering datasets to demonstrate the effectiveness of our approach.
Figures & tables
Figure 1: Overview of the KASB Framework. A finetuned generator is derived via the DPO algorithm on the training set to produce a more effective generative context generator, which then serves in the test phase. During inference, the finetuned generator first generates query-aligned generative contexts G ; meanwhile, the evidence (representing the complete retrieved knowledge base) is processed through the generated context-based selector to identify relevant retrieved contexts C , which are then fused with the generated contexts [C;G] for final answer generation.
Datasets
Train
Dev
TriviaQA
75675
8750
NQ
91334
10039
WebQ
3906
-
Table 1: Datasets statistics.
Methods
TriviaQA
NQ
WebQ
Single Knowledge
Retri-Only
62.2
48.6
48.3
Gen-Only
67.5
50.3
41.6
Two Knowledge
HyDE
72.3
53.4
54.6
COMBO
74.6
54.2
53.0
Table 2: Exact match scores on test dataset.
Figure 2: Average retrieval exact match score obtained by finetuned vs original LLM generator on generated contexts.
Model
TriviaQA
NQ
WebQ
KASB
75.4
59.6
61.2
w/o GK
69.9
54.9
56.3
w/o RK
73.2
56.8
59.5
w/o DPO
74.6
57.7
59.8
Table 4: Question answering performance (Exact Match). GK and RK denote generated knowledge and retrieved knowledge, respectively. The combined knowledge results are computed by union of single knowledge sources.
Selection
TriviaQA
NQ
WebQ
Avg.
Fixed- k , k=5
73.2
58.2
59.2
63.5
Fixed- k , k=10
73.8
58.5
59.8
64.0
Fixed- k , k=20
74.2
58.9
59.7
64.3
Fixed- k , k=50
73.5
58.1
59.4
63.7
Adaptive k⋆
75.4
59.6
61.2
65.4
Table 5: Fixed- k versus adaptive evidence selection. EM is reported on the three QA datasets; Avg. is the mean across datasets.
Figure 3: Parameter sensitivity analysis showing the effect of DPO coefficient β and initial candidate pool size on EM performance across datasets.
Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallucinations, and improving factuality. However, existing systems rely on generating natural language queries at each hop and maintaining a strict architectural separation between retriever and generator, preventing them from leveraging the full representational capacity of the LLM. We propose \textbf{LAnR} (Latent Abstraction for RAG), a unified framework in which a single LLM jointly performs encoding, retrieval, and generation entirely within its own latent space. Rather than generating textual queries, LAnR produces dense retrieval vectors from the hidden states of a designated \texttt{[PRED]} token and uses them to match against encoded document representations from the same model. Furthermore, LAnR adaptively decides when sufficient evidence has been retrieved using a lightweight MLP control head over those same hidden states, eliminating both the separate retriever and explicit token-level stopping reasoning. This design is motivated by our empirical observation that answer token entropy reliably signals retrieval sufficiency. Extensive experiments on six QA benchmarks spanning single-hop and multi-hop settings demonstrate that LAnR outperforms existing RAG methods, while achieving improved inference efficiency through reduced number of retrieval calls and tighter model integration.
Retrieval-Augmented Generation (RAG) systems depend critically on document chunking quality for retrieving relevant context. Fixed chunking segments documents into uniform units irrespective of semantics or user intent, producing a precision-recall trade-off unresolvable by tuning chunk size alone. Semantic and agentic methods partially address these limitations but do not integrate user queries at the chunking stage. We present Query-Adaptive Semantic Chunking (QASC), which dynamically constructs chunks by integrating queries into segmentation through three mechanisms: cosine similarity scoring between sentence and query embeddings to identify seed sentences, contextual window expansion around seeds to preserve coherence, and chunk-level score aggregation to ensure holistic relevance. We evaluate QASC on 100 technical documents across 200 queries spanning four types, comparing against fixed chunking at five granularities, recursive splitting, semantic chunking, and agentic chunking. QASC achieves an F1-score of 0.85, a relative improvement of 18-27% over fixed chunking and 8-12% over semantic and agentic alternatives. Ablation studies confirm each component contributes meaningfully. Human evaluation by three annotators (Cohen kappa = 0.82) corroborates that QASC produces more relevant and coherent chunks than existing methods.
Retrieval-augmented generation (RAG) typically treats context selection as ranking chunks against a single query embedding. This assumption breaks down for complex queries, such as multi-hop or ambiguous questions, where top-k selection tends to over-cover one semantic aspect while ignoring critical sub-questions. We propose GeoRAG, which recasts context selection as Information Demand Coverage Optimization. GeoRAG builds a multi-dimensional demand distribution through diverse sub-query generation and reverse-validation weighting, then selects context by minimizing the Sinkhorn-Wasserstein distance between this demand distribution and the coverage of the selected set. The resulting demand-weighted facility-location objective is monotone submodular, giving a 1−1/e greedy guarantee, which we approximate with a Sinkhorn-based marginal-gain surrogate. The method is unsupervised, training-free, and retrieval-agnostic. We further show that single-point, query-proximity scorers cannot cover multi-modal demands, exposing a structural limit of ranking-based selection. On six open-domain QA benchmarks, GeoRAG improves exact match (EM) by +6.5 to +7.5 points over top-k truncation (up to +9.7 on HotpotQA and ASQA) and outperforms strong baselines including MMR, DPP, BGE-Reranker, SMART-RAG, and AdaGReS, with stable gains across context budgets and sub-query generators.
Bingxue Zhang, Jianying Jia, Feida Zhu
University of Shanghai for Science and Technology, Shanghai, China · Singapore Management University, Singapore