RAG systems are increasingly used to summarize what large collections of documents say. A user asks "What do people think about X?" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.
Figures & tables
Figure 1: WARP pipeline. Stage 1 : semantic retrieval returns a Top- N candidate pool. Stage 1.5 : deficit-aware pool expansion recovers under-represented sentiment poles via re-retrieval ( Entity-Gated or Adaptive Expansion ; § 3.3 ; self-bypasses when the pool is already balanced). Stage 2 : W1 re-ranking selects the final k documents whose empirical opinion distribution matches the population target Ppop ( W1 Minimizer for dense pools, W1 -MMR for sparse pools, or WassRank OT as a tuning-free fallback; § 3.4 ). Pool density drives the variant choice (decision matrix in Table 3 ).
Figure 2: W1 re-rankers (no re-retrieval): domain-matched variants reduce distributional error at least 43% vs. Top- k (gray dashed) — 43% Seller Forums ( W1 -MMR), 80% Yelp ( W1 Minimizer), 69% OpinRank ( W1 Minimizer); see Table 1 . Curves converge at k∗=3 – 5 documents (black dot; Appendix C.3 ) as most entities’ opinion mass concentrates in 2–3 of 7 bins. N=200 .
Seller Forums
Yelp Hotels
OpinRank Cars
Latency
Method
W1↓
EM%
W1↓
EM%
W1↓
EM%
(ms)
Top- k (baseline)
10.33
42.9
13.60
40.1
13.33
27.0
0
MMR *
10.31
43.5
8.67
14.2
9.64
11.3
12
DPP ***
11.69
41.2
5.63
78.8
7.38
69.2
45
OpinionMMR **
10.44
43.5
8.59
12.4
10.02
10.6
12
KL (JS) Minimizer *
8.87
51.4
3.00
93.8
4.54
85.8
9
Table 1: Main results: Diversity baselines and calibrated methods vs WARP re-rankers. W1↓ : distributional distance to the population opinion target. EM% ↑ : fraction of selected documents matching the queried entity. EG/AE = pool expansion strategies (Entity-Gated / Adaptive Expansion). N=200 candidates, k=20 output. Paired Wilcoxon vs. Top- k : ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 . Bold: best per column. † Shared FAISS index across all OpinRank entities (EM=82.9%); entity-specific indices give W1=0.95 with 100% EM (Appendix F.3 ).
Approach
W/L/T
Fair%
Win%
Random
16/36/104
31%
10%
MMR ∗
39/8/109
83%
25%
DPP ∗
54/7/95
89%
35%
WassRank OT ∗
66/28/62
70%
42%
W 1 -MMR ∗
44/10/102
81%
28%
W 1 Minimizer ∗
51/8/97
86%
33%
Table 2: Generation evaluation ( k=5 ): 5-judge majority vote, position-controlled blind pairwise vs. Top- k ( N=156 queries). Fair% = W/(W+L). Win% = W/ N . All approaches significant at p<0.001 (sign test) except Random. k=10 in Appendix G.1.3 .
Pool Condition
Method
ms
Complexity
Dense (EM ≥ 85%)
W1 Minimizer
9
O(∣Ce∣⋅k⋅m)
Sparse (baseline EM ∼ 40%)
W1 -MMR
154
O(N⋅k⋅m)
Variable / unknown
EG + W1 Min
13
O(∣Ce∣⋅k⋅m)
Table 3: Deployment decision matrix. Pool density determines algorithm choice; all methods SLA-compliant ( N=200 , k=20 , m=7 bins). ∣Ce∣ = entity-matched candidates. Noise tolerance details in Appendix F.1 .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Sentiment
Intensity
SI Score
Positive
High / Med / Low
+30 / +20 / +10
Neutral
Any
0
Mixed
—
0
Negative
High / Med / Low
−30 / −20 / −10
Appendix
Table 4: Sentiment-Intensity (SI) mapping to the 7-bin ordinal scale. Mixed labels (praise-and-criticism co-occurring in one review) collapse to the neutral bin; they account for 4–8% of extractions across domains and W1 Minimizer never crosses the Top- k baseline on Yelp even at 50% label corruption (§ F.1 ), so this pooling is bounded in impact. Multi-issue axes require multi-marginal transport — outside our current scope (§Limitations).
Dataset
Domain
Source
Size
Entities
Queries
Search
Embedding
Chunking
Labeling
Seller Forums
E-comm. seller
Public
∼ 8K
6
36
Hybrid
Titan V2 ( Amazon Web Services, 2024a ) (1024d)
300 tok / 20% overlap
LLM-enriched metadata
Yelp Hotels
Hospitality
Yelp Open
14K
10
60
Semantic
MiniLM-L6 (384d)
Whole review
LLM-extracted (Claude)
OpinRank Cars
Automotive
OpinRank
13K
10
60
Semantic
MiniLM-L6 (384d)
Whole review
LLM-extracted
Appendix
Table 5: Dataset summary across three evaluation domains.
Type
Idx
Seller Forums
Yelp Hotels
OpinRank Cars
breadth
0
What do sellers think about {entity}?
What do guests think about {entity}?
What do owners think about {entity}?
breadth
1
Summarize seller opinions on {entity}.
Summarize guest opinions on {entity}.
Summarize owner opinions on {entity}.
polar
0
What are the biggest complaints about {entity}?
What are the biggest complaints about {entity}?
What are the biggest complaints about {entity}?
polar
1
What positive experiences have sellers had with {entity}?
What positive experiences have guests had with {entity}?
What do owners praise most about {entity}?
segment
0
How do small sellers vs large sellers feel about {entity}?
How do business travelers vs leisure travelers feel about {entity}?
How do commuters vs enthusiasts feel about {entity}?
segment
1
Do new sellers and experienced sellers differ in their views of {entity}?
Do solo guests and families differ in their views of {entity}?
Do new buyers and long-term owners differ in their views of {entity}?
Appendix
Table 6: Query templates per domain. Each template is instantiated with every selected entity (6 for Seller Forums, 10 for Yelp/OpinRank), producing 6 questions per entity (2 breadth, 2 polar, 2 segment).
Symbol
Meaning
d
A candidate document
e
Target entity
k
Output size (documents returned)
N
Candidate pool size
SI(d)
Sentiment-intensity label of d
Ppop
Ground-truth population distribution
Appendix
Table 7: Notation reference.
Component
Local FAISS
Managed Bedrock Knowledge Base
Retrieval
< 50 ms
∼ 1.5 s
Re-retrieval †
< 5 ms
∼ 100 ms
Re-ranking (p99)
< 310 ms
< 310 ms
End-to-end
< 365 ms
∼ 1.9 s
† Only fires when pool expansion (Entity-Gated / Adaptive) detects a deficit.
Appendix
Table 8: Per-component latency breakdown. Re-ranking is hardware-only (no model inference); retrieval and re-retrieval depend on the vector-store backend. p99 for re-ranking; typical/median for the rest.
Domain
Method
Queries
p50 (ms)
p95 (ms)
p99 (ms)
Max (ms)
Seller Forums
W1 Minimizer
36
0.1
38.7
45.5
48.0
Seller Forums
W1 -MMR
36
156.8
163.1
166.2
167.4
Seller Forums
EG + W1 Min
36
0.6
211.0
216.2
218.8
Seller Forums
DPP
36
47.8
49.1
51.9
53.0
Yelp Hotels
W1 Minimizer
60
65.0
270.8
289.9
295.3
Yelp Hotels
W1 -MMR
60
81.6
284.7
301.6
303.5
Appendix
Table 9: Per-query re-ranking latency percentiles ( N=96 queries: 36 Seller Forums + 60 Yelp Hotels, single-threaded on the hardware above). p99 drives production SLAs; max reported for tail auditing. All WARP variants stay under 310 ms at p99.
Yelp
OpinRank
Method
W1
EM%
W1
EM%
KL (JS)-MMR
2.51
38.2%
3.32
36.1%
W1 -MMR
4.02
19.4%
3.41
21.8%
KL (JS) Minimizer
3.00
93.8%
4.54
85.8%
W1 Minimizer
2.76
93.8%
4.14
85.8%
Appendix
Table 10: Entity-match confound in hybrid variants. KL(JS)-MMR retrieves ∼ 1.7–2 × more entity-matched documents than W1 -MMR on both Yelp and OpinRank. In the controlled Minimizer comparison (same entity filter, EM equalized) W1 wins on both domains — the KL(JS)-MMR W1 advantage in the uncontrolled comparison traces to the EM gap, not to metric superiority.
Approach
EM%
W1
Δ vs Top-K
Δ vs EF+Top-K
Top K (no filter)
40.1
13.60
—
—
EntityFilter + Top-K
100.0
7.38
−45.7%
—
EntityFilter + MMR
100.0
6.99
−48.6%
−5.3%
EntityFilter + DPP
100.0
5.64
−58.5%
−23.6%
W1 Minimizer
93.8
2.76
−79.7%
−62.6%
W1 -MMR
19.4
4.02
−70.4%
−45.5%
Appendix
Table 11: Entity-filter control experiment (Yelp Hotels, k=20 ). Entity filtering alone reduces W1 by 45.7%; W1 Minimizer achieves 62.6% additional reduction beyond the entity-filter control.
Figure 3: W1 Minimizer convergence vs. output size k . Median convergence point k∗ (where ΔW1<0.5 per additional document): Seller Forums k∗=3 , Yelp k∗=4 , OpinRank k∗=5 .
Table 14: Benjamini–Hochberg FDR correction ( α=0.05 ) on paired-Wilcoxon p -values for the W1 family. Every W1 -family algorithm survives on every domain (14/14). Reduction is per-query paired reduction relative to Top- k ; aggregate reductions in Table 1 weight queries differently. ∗∗∗padj<10−3 , ∗∗padj<10−2 , ∗padj<0.05 .
Dataset
Method
#Entities
Mean ΔW1
95% CI
Yelp
W1 Minimizer
10
10.85
[8.26,13.37]
Yelp
EG + W1 Min
10
12.09
[8.72,15.37]
Yelp
W1 -MMR
10
9.58
[6.40,12.80]
OpinRank
W1 Minimizer
9
9.65
[6.92,12.34]
OpinRank
W1 -MMR
9
10.22
[6.48,14.17]
Seller Forums
W1 Minimizer
4
5.32
[0.67,10.38]
Appendix
Table 15: Entity-clustered bootstrap (10,000 resamples). Mean improvement in W1 relative to Top- k ; 95% CIs computed by resampling entities with replacement. All CIs exclude zero.
Entity
N reviews
W1(star-Ppop,LLM-Ppop)
Booking Experience
1,372
3.66
Breakfast
788
1.44
Breakfast Quality
637
4.15
Business Center
142
3.23
Casino
208
1.36
Loyalty Program
607
1.57
Appendix
Table 16: Star-derived vs. LLM-derived Ppop per Yelp entity. Mean W1 gap: 2.88. Bounds: W1(WARP,star-Ppop)≤5.64 , W1(Top-k,star-Ppop)≥10.72 ; worst-case reduction ≥47.4% .
Figure 4: Dirichlet perturbation robustness across all three domains. Dashed red line = Top- k baseline. At ε=0.2 , W1 Minimizer remains 72% better than Top- k on Yelp.
Figure 5: Mislabel sensitivity: W1 vs. label error rate. Dashed red line = Top- k baseline. W1 Minimizer never crosses Top- k even at 50% error on Yelp.
Method
Seller W1
Yelp W1
OpinRank W1
Top- k (baseline)
10.33
13.60
13.33
Top- k + 5% Noise
10.07
13.00
12.97
Reduction
2.5%
4.4%
2.7%
Significance
ns ( p>0.05 , all domains)
W1 Minimizer
8.84
2.76
4.14
Reduction
14.4%
79.7%
69.0%
Appendix
Table 17: Noise control: 5% Gaussian perturbation of retrieval scores yields negligible W1 improvement (2–5%, p>0.05 ), confirming that our 43–89% distributional improvements ( W1 -MMR on sparse pools; W1 Minimizer / EG+ W1 Min on dense) represent genuine distributional optimization.
Figure 6: λ -sensitivity for W1 -MMR (left) and OpinionMMR (right). W1 -MMR is stable at λ≤0.5 and degrades at λ=0.9 . OpinionMMR shows higher W1 across all domains with limited sensitivity to λ .
Figure 7: W1 vs. pool size N ( k=20 ). N=100 suffices for dense pools; W1 -MMR benefits from larger pools on sparse data.
Figure 8: W1(P^pop,Ppop∗) vs. entity-matched documents sampled. With 50–100 documents per entity, estimation error is comparable to or below the re-ranking improvement magnitude itself.
Judge
Consensus Tie Rate ( k=5 )
Consensus Tie Rate ( k=10 )
Claude Sonnet 4
46%
46%
Mistral Large 3
48%
61%
DeepSeek v3.2
63%
71%
Llama 3.3 70B
73%
86%
Amazon Nova Pro
79%
86%
Appendix
Table 18: Consensus tie rates across the 5-judge panel. Higher tie rates indicate stronger positional bias (position-swap converts inconsistent preferences to ties).
Judge (vs. others)
k=5κd
k=5 Agree%
k=10κd
k=10 Agree%
Claude vs. others
0.68
88.2%
0.86
93.5%
Llama 3.3 vs. others
0.89
96.0%
0.97
98.8%
Mistral L3 vs. others
0.89
95.8%
0.95
97.8%
DeepSeek vs. others
0.93
97.4%
0.95
97.9%
Nova Pro vs. others
0.88
95.3%
0.94
97.3%
Appendix
Table 19: Mean directional κd per judge vs. all others (fairness dimension). Non-Anthropic judges agree at κd≥0.95 at k=10 ; at k=5 , κd=0.88 – 0.93 .
k=5
k=10
k=20
Approach
W/L/T
Fair%
Win%
W/L/T
Fair%
Win%
W/L/T
Fair%
Win%
Random
16/36/104
31%
10%
12/40/104
23%
8%
3/26/127
10%
2%
MMR ∗∗
39/8/109
83%
25%
21/9/126
70%
13%
5/12/139
29%
3%
OpinionMMR ∗∗
50/11/95
82%
32%
25/8/123
76%
16%
7/14/135
33%
4%
DPP ∗∗
54/7/95
89%
35%
42/6/108
88%
27%
25/3/128
89%
16%
WassRank OT ∗∗
66/28/62
70%
42%
75/28/53
73%
48%
33/3/120
92%
21%
Appendix
Table 20: 5-judge majority-vote generation results (cross-domain, 156 questions) at k∈{5,10,20} . All calibration methods achieve p<0.001 at k=5 ; retain significance at k=10 ( p≤0.021 ) and k=20 ( p≤0.004 ). Random fails at all k values. At k=20 , MMR/OpinionMMR Fair% collapse (29%/33%) as semantic diversity saturates; distributional methods maintain or grow their advantage; WassRank OT rises 70% → 73% → 92%. Fair% = W/(W+L). Win% = W/ N .
k=5
k=10
Approach
Prop. (A)
Cov. (B)
Prop. (A)
Cov. (B)
Stratified (oracle)
35%
—
—
—
W1 Minimizer
34%
38%
30%
27%
DPP
26%
35%
24%
25%
W1 -MMR
23%
36%
18%
17%
OpinionMMR
21%
30%
13%
19%
Appendix
Table 21: Judge prompt ablation: fairness win rate (%) under proportional vs. coverage definitions (3 domains, 156 questions per approach). W1 Minimizer dominates under proportional accuracy; the gap narrows under coverage. Stratified sampling (oracle) confirms the Prompt A criterion tracks the ground-truth distribution, not any specific retrieval strategy — W1 Minimizer matches oracle Fair% within 1 pp at k=5 .
Approach
Info% (A)
Info% (B)
Grnd% (A)
Grnd% (B)
W1 Minimizer
44%
40%
28%
24%
DPP
38%
40%
27%
24%
Random
19%
18%
15%
8%
Appendix
Table 22: Control dimensions remain stable across prompt variants ( k=5 , Δ = Prompt B − Prompt A). Shifts ≤ 4 pp confirm the fairness prompt change does not contaminate other evaluation axes.
Domain
Method
N
W/L/T
Fair%
Win%
Seller Forums
W1 Minimizer
36
10/9/17
53%
28%
KL (JS) Minimizer
36
7/10/19
41%
19%
Yelp
W1 Minimizer
60
39/4/17
91%
65%
KL (JS) Minimizer
60
31/8/21
79%
52%
OpinRank
W1 Minimizer
60
29/8/23
78%
48%
KL (JS) Minimizer
60
35/10/15
78%
58%
Appendix
Table 23: W1 Minimizer vs. KL(JS) Minimizer, generation evaluation at k=5 (156 queries, single-judge Claude Sonnet 4, position-controlled, blind pairwise vs. Top- k ). W1 leads KL(JS) by +7 pp cross-domain; both significantly beat Top- k ( p<0.001 ).
This position paper argues that Retrieval-Augmented Generation (RAG) systems exhibit a factual bias-optimizing for epistemic uncertainty reduction while ignoring the aleatoric uncertainty inherent in opinion-rich content. This misalignment demands a paradigm shift in RAG system design. A survey of 34 major RAG benchmarks reveals that only one addresses opinion synthesis, confirming that the bias is structural and embedded in datasets, retrieval-generation objectives, and evaluation metrics alike. Beyond technical limitations, this bias poses risks to transparent and accountable AI. Namely, echo chamber effects that amplify dominant viewpoints, which can lead to opinion manipulation and under-representation of minority voices. We formalize the problem through the lens of uncertainty quantification, showing that factual queries should minimize posterior entropy while opinion queries must preserve it. We derive a unified objective over coverage, fidelity, and fairness using the Wasserstein distance. As an existence proof, we present Opinion-Aware RAG (O-RAG), an architecture featuring LLM-based opinion extraction and entity-linked opinion metadata. We evaluate it across two domains -- e-commerce seller forums and public hotel reviews. Experiments demonstrate 18-48% reduction in Wasserstein distance to corpus-level sentiment distributions, +26.8% sentiment diversity, and +42.7% entity match rate. Human evaluators preferred opinion-enriched generation 79.2% of the time. We propose a research agenda and argue that as RAG systems increasingly mediate access to information, their ability to represent diverse perspectives is of the essence.
Retrieval-augmented generation (RAG) systems typically rely on a single retriever and a single set of hyperparameters, despite facing highly heterogeneous queries that range from simple factoid questions to complex multi-hop reasoning. We propose a method that automatically selects a small, diverse subset of retrievers (a portfolio) from a large pool of candidates, to cover different regions of the target query distribution. We formalize this setting via an expected best-of-k objective over the query distribution and show that it admits an efficient portfolio construction algorithm with near-optimal guarantees. Across multiple QA benchmarks, our learned portfolios and router pipeline consistently outperform single-retriever and naive multi-retriever baselines on both retrieval metrics and answer quality. In addition, compared to inference-time hyperparameter tuning approaches, fixed portfolios enable parallel retrieval and LLM calls, achieving comparable (and sometimes better) accuracy with substantially lower latency and token cost.
Miltiadis Stouras, Vincent Cohen-Addad, Silvio Lattanzi +1
Retrieval-Augmented Generation (RAG) supplements a language model's input with retrieved documents, yet most RAG pipelines inherit retrieval components designed for human readers. How retrieved content should be represented when the consumer is a large language model (LLM) rather than a human is less well understood. Recent work has proposed transformations of retrieved content and identified properties that affect generation, but each examines a single transformation or property in isolation, leaving open which features of a document's representation matter most. We address this with a controlled comparison: holding retrieval fixed, we vary only the representation of retrieved documents, comparing an original baseline against thirteen transformations spanning selection, summarisation, and reformulation, in query-dependent and query-independent variants. Across these fourteen representations we measure question-answering accuracy for four generators, and for each representation we also measure answer retention: whether a known answer-bearing document still supports its answer after transformation. We find that answer retention is the primary determinant of generator accuracy; notably, when retention is high, a representation's wording, structure, length, and query-dependence have limited effect. This suggests that accuracy gains attributed to specific mechanisms in prior work may be partly explained by how well those mechanisms preserve answer-bearing content, an attribution that cannot be settled without controlling for retention.
Jonathan J Ross, Bevan Koopman, Anton van der Vegt +1
The University of Queensland · CSIRO / The University of Queensland