RAG systems are increasingly used to summarize what large collections of documents say. A user asks "What do people think about X?" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.
Figures & tables
Figure 1: WARP pipeline. Stage 1 : semantic retrieval returns a Top- N candidate pool. Stage 1.5 : deficit-aware pool expansion recovers under-represented sentiment poles via re-retrieval ( Entity-Gated or Adaptive Expansion ; § 3.3 ; self-bypasses when the pool is already balanced). Stage 2 : W1 re-ranking selects the final k documents whose empirical opinion distribution matches the population target Ppop ( W1 Minimizer for dense pools, W1 -MMR for sparse pools, or WassRank OT as a tuning-free fallback; § 3.4 ). Pool density drives the variant choice (decision matrix in Table 3 ).
Figure 2: W1 re-rankers (no re-retrieval): domain-matched variants reduce distributional error at least 43% vs. Top- k (gray dashed) — 43% Seller Forums ( W1 -MMR), 80% Yelp ( W1 Minimizer), 69% OpinRank ( W1 Minimizer); see Table 1 . Curves converge at k∗=3 – 5 documents (black dot; Appendix C.3 ) as most entities’ opinion mass concentrates in 2–3 of 7 bins. N=200 .
Seller Forums
Yelp Hotels
OpinRank Cars
Latency
Method
W1↓
EM%
W1↓
EM%
W1↓
EM%
(ms)
Top- k (baseline)
10.33
42.9
13.60
40.1
13.33
27.0
0
MMR *
10.31
43.5
8.67
14.2
9.64
11.3
12
DPP ***
11.69
41.2
5.63
78.8
7.38
69.2
45
OpinionMMR **
10.44
43.5
8.59
12.4
10.02
10.6
12
KL (JS) Minimizer *
8.87
51.4
3.00
93.8
4.54
85.8
9
Table 1: Main results: Diversity baselines and calibrated methods vs WARP re-rankers. W1↓ : distributional distance to the population opinion target. EM% ↑ : fraction of selected documents matching the queried entity. EG/AE = pool expansion strategies (Entity-Gated / Adaptive Expansion). N=200 candidates, k=20 output. Paired Wilcoxon vs. Top- k : ∗∗∗p<0.001 , ∗∗p<0.01 , ∗p<0.05 . Bold: best per column. † Shared FAISS index across all OpinRank entities (EM=82.9%); entity-specific indices give W1=0.95 with 100% EM (Appendix F.3 ).
Approach
W/L/T
Fair%
Win%
Random
16/36/104
31%
10%
MMR ∗
39/8/109
83%
25%
DPP ∗
54/7/95
89%
35%
WassRank OT ∗
66/28/62
70%
42%
W 1 -MMR ∗
44/10/102
81%
28%
W 1 Minimizer ∗
51/8/97
86%
33%
Table 2: Generation evaluation ( k=5 ): 5-judge majority vote, position-controlled blind pairwise vs. Top- k ( N=156 queries). Fair% = W/(W+L). Win% = W/ N . All approaches significant at p<0.001 (sign test) except Random. k=10 in Appendix G.1.3 .
Pool Condition
Method
ms
Complexity
Dense (EM ≥ 85%)
W1 Minimizer
9
O(∣Ce∣⋅k⋅m)
Sparse (baseline EM ∼ 40%)
W1 -MMR
154
O(N⋅k⋅m)
Variable / unknown
EG + W1 Min
13
O(∣Ce∣⋅k⋅m)
Table 3: Deployment decision matrix. Pool density determines algorithm choice; all methods SLA-compliant ( N=200 , k=20 , m=7 bins). ∣Ce∣ = entity-matched candidates. Noise tolerance details in Appendix F.1 .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Sentiment
Intensity
SI Score
Positive
High / Med / Low
+30 / +20 / +10
Neutral
Any
0
Mixed
—
0
Negative
High / Med / Low
−30 / −20 / −10
Appendix
Table 4: Sentiment-Intensity (SI) mapping to the 7-bin ordinal scale. Mixed labels (praise-and-criticism co-occurring in one review) collapse to the neutral bin; they account for 4–8% of extractions across domains and W1 Minimizer never crosses the Top- k baseline on Yelp even at 50% label corruption (§ F.1 ), so this pooling is bounded in impact. Multi-issue axes require multi-marginal transport — outside our current scope (§Limitations).
Dataset
Domain
Source
Size
Entities
Queries
Search
Embedding
Chunking
Labeling
Seller Forums
E-comm. seller
Public
∼ 8K
6
36
Hybrid
Titan V2 ( Amazon Web Services, 2024a ) (1024d)
300 tok / 20% overlap
LLM-enriched metadata
Yelp Hotels
Hospitality
Yelp Open
14K
10
60
Semantic
MiniLM-L6 (384d)
Whole review
LLM-extracted (Claude)
OpinRank Cars
Automotive
OpinRank
13K
10
60
Semantic
MiniLM-L6 (384d)
Whole review
LLM-extracted
Appendix
Table 5: Dataset summary across three evaluation domains.
Type
Idx
Seller Forums
Yelp Hotels
OpinRank Cars
breadth
0
What do sellers think about {entity}?
What do guests think about {entity}?
What do owners think about {entity}?
breadth
1
Summarize seller opinions on {entity}.
Summarize guest opinions on {entity}.
Summarize owner opinions on {entity}.
polar
0
What are the biggest complaints about {entity}?
What are the biggest complaints about {entity}?
What are the biggest complaints about {entity}?
polar
1
What positive experiences have sellers had with {entity}?
What positive experiences have guests had with {entity}?
What do owners praise most about {entity}?
segment
0
How do small sellers vs large sellers feel about {entity}?
How do business travelers vs leisure travelers feel about {entity}?
How do commuters vs enthusiasts feel about {entity}?
segment
1
Do new sellers and experienced sellers differ in their views of {entity}?
Do solo guests and families differ in their views of {entity}?
Do new buyers and long-term owners differ in their views of {entity}?
Appendix
Table 6: Query templates per domain. Each template is instantiated with every selected entity (6 for Seller Forums, 10 for Yelp/OpinRank), producing 6 questions per entity (2 breadth, 2 polar, 2 segment).
Symbol
Meaning
d
A candidate document
e
Target entity
k
Output size (documents returned)
N
Candidate pool size
SI(d)
Sentiment-intensity label of d
Ppop
Ground-truth population distribution
Appendix
Table 7: Notation reference.
Component
Local FAISS
Managed Bedrock Knowledge Base
Retrieval
< 50 ms
∼ 1.5 s
Re-retrieval †
< 5 ms
∼ 100 ms
Re-ranking (p99)
< 310 ms
< 310 ms
End-to-end
< 365 ms
∼ 1.9 s
† Only fires when pool expansion (Entity-Gated / Adaptive) detects a deficit.
Appendix
Table 8: Per-component latency breakdown. Re-ranking is hardware-only (no model inference); retrieval and re-retrieval depend on the vector-store backend. p99 for re-ranking; typical/median for the rest.
Domain
Method
Queries
p50 (ms)
p95 (ms)
p99 (ms)
Max (ms)
Seller Forums
W1 Minimizer
36
0.1
38.7
45.5
48.0
Seller Forums
W1 -MMR
36
156.8
163.1
166.2
167.4
Seller Forums
EG + W1 Min
36
0.6
211.0
216.2
218.8
Seller Forums
DPP
36
47.8
49.1
51.9
53.0
Yelp Hotels
W1 Minimizer
60
65.0
270.8
289.9
295.3
Yelp Hotels
W1 -MMR
60
81.6
284.7
301.6
303.5
Appendix
Table 9: Per-query re-ranking latency percentiles ( N=96 queries: 36 Seller Forums + 60 Yelp Hotels, single-threaded on the hardware above). p99 drives production SLAs; max reported for tail auditing. All WARP variants stay under 310 ms at p99.
Yelp
OpinRank
Method
W1
EM%
W1
EM%
KL (JS)-MMR
2.51
38.2%
3.32
36.1%
W1 -MMR
4.02
19.4%
3.41
21.8%
KL (JS) Minimizer
3.00
93.8%
4.54
85.8%
W1 Minimizer
2.76
93.8%
4.14
85.8%
Appendix
Table 10: Entity-match confound in hybrid variants. KL(JS)-MMR retrieves ∼ 1.7–2 × more entity-matched documents than W1 -MMR on both Yelp and OpinRank. In the controlled Minimizer comparison (same entity filter, EM equalized) W1 wins on both domains — the KL(JS)-MMR W1 advantage in the uncontrolled comparison traces to the EM gap, not to metric superiority.
Approach
EM%
W1
Δ vs Top-K
Δ vs EF+Top-K
Top K (no filter)
40.1
13.60
—
—
EntityFilter + Top-K
100.0
7.38
−45.7%
—
EntityFilter + MMR
100.0
6.99
−48.6%
−5.3%
EntityFilter + DPP
100.0
5.64
−58.5%
−23.6%
W1 Minimizer
93.8
2.76
−79.7%
−62.6%
W1 -MMR
19.4
4.02
−70.4%
−45.5%
Appendix
Table 11: Entity-filter control experiment (Yelp Hotels, k=20 ). Entity filtering alone reduces W1 by 45.7%; W1 Minimizer achieves 62.6% additional reduction beyond the entity-filter control.
Figure 3: W1 Minimizer convergence vs. output size k . Median convergence point k∗ (where ΔW1<0.5 per additional document): Seller Forums k∗=3 , Yelp k∗=4 , OpinRank k∗=5 .
Table 14: Benjamini–Hochberg FDR correction ( α=0.05 ) on paired-Wilcoxon p -values for the W1 family. Every W1 -family algorithm survives on every domain (14/14). Reduction is per-query paired reduction relative to Top- k ; aggregate reductions in Table 1 weight queries differently. ∗∗∗padj<10−3 , ∗∗padj<10−2 , ∗padj<0.05 .
Dataset
Method
#Entities
Mean ΔW1
95% CI
Yelp
W1 Minimizer
10
10.85
[8.26,13.37]
Yelp
EG + W1 Min
10
12.09
[8.72,15.37]
Yelp
W1 -MMR
10
9.58
[6.40,12.80]
OpinRank
W1 Minimizer
9
9.65
[6.92,12.34]
OpinRank
W1 -MMR
9
10.22
[6.48,14.17]
Seller Forums
W1 Minimizer
4
5.32
[0.67,10.38]
Appendix
Table 15: Entity-clustered bootstrap (10,000 resamples). Mean improvement in W1 relative to Top- k ; 95% CIs computed by resampling entities with replacement. All CIs exclude zero.
Entity
N reviews
W1(star-Ppop,LLM-Ppop)
Booking Experience
1,372
3.66
Breakfast
788
1.44
Breakfast Quality
637
4.15
Business Center
142
3.23
Casino
208
1.36
Loyalty Program
607
1.57
Appendix
Table 16: Star-derived vs. LLM-derived Ppop per Yelp entity. Mean W1 gap: 2.88. Bounds: W1(WARP,star-Ppop)≤5.64 , W1(Top-k,star-Ppop)≥10.72 ; worst-case reduction ≥47.4% .
Figure 4: Dirichlet perturbation robustness across all three domains. Dashed red line = Top- k baseline. At ε=0.2 , W1 Minimizer remains 72% better than Top- k on Yelp.
Figure 5: Mislabel sensitivity: W1 vs. label error rate. Dashed red line = Top- k baseline. W1 Minimizer never crosses Top- k even at 50% error on Yelp.
Method
Seller W1
Yelp W1
OpinRank W1
Top- k (baseline)
10.33
13.60
13.33
Top- k + 5% Noise
10.07
13.00
12.97
Reduction
2.5%
4.4%
2.7%
Significance
ns ( p>0.05 , all domains)
W1 Minimizer
8.84
2.76
4.14
Reduction
14.4%
79.7%
69.0%
Appendix
Table 17: Noise control: 5% Gaussian perturbation of retrieval scores yields negligible W1 improvement (2–5%, p>0.05 ), confirming that our 43–89% distributional improvements ( W1 -MMR on sparse pools; W1 Minimizer / EG+ W1 Min on dense) represent genuine distributional optimization.
Figure 6: λ -sensitivity for W1 -MMR (left) and OpinionMMR (right). W1 -MMR is stable at λ≤0.5 and degrades at λ=0.9 . OpinionMMR shows higher W1 across all domains with limited sensitivity to λ .
Figure 7: W1 vs. pool size N ( k=20 ). N=100 suffices for dense pools; W1 -MMR benefits from larger pools on sparse data.
Figure 8: W1(P^pop,Ppop∗) vs. entity-matched documents sampled. With 50–100 documents per entity, estimation error is comparable to or below the re-ranking improvement magnitude itself.
Judge
Consensus Tie Rate ( k=5 )
Consensus Tie Rate ( k=10 )
Claude Sonnet 4
46%
46%
Mistral Large 3
48%
61%
DeepSeek v3.2
63%
71%
Llama 3.3 70B
73%
86%
Amazon Nova Pro
79%
86%
Appendix
Table 18: Consensus tie rates across the 5-judge panel. Higher tie rates indicate stronger positional bias (position-swap converts inconsistent preferences to ties).
Judge (vs. others)
k=5κd
k=5 Agree%
k=10κd
k=10 Agree%
Claude vs. others
0.68
88.2%
0.86
93.5%
Llama 3.3 vs. others
0.89
96.0%
0.97
98.8%
Mistral L3 vs. others
0.89
95.8%
0.95
97.8%
DeepSeek vs. others
0.93
97.4%
0.95
97.9%
Nova Pro vs. others
0.88
95.3%
0.94
97.3%
Appendix
Table 19: Mean directional κd per judge vs. all others (fairness dimension). Non-Anthropic judges agree at κd≥0.95 at k=10 ; at k=5 , κd=0.88 – 0.93 .
k=5
k=10
k=20
Approach
W/L/T
Fair%
Win%
W/L/T
Fair%
Win%
W/L/T
Fair%
Win%
Random
16/36/104
31%
10%
12/40/104
23%
8%
3/26/127
10%
2%
MMR ∗∗
39/8/109
83%
25%
21/9/126
70%
13%
5/12/139
29%
3%
OpinionMMR ∗∗
50/11/95
82%
32%
25/8/123
76%
16%
7/14/135
33%
4%
DPP ∗∗
54/7/95
89%
35%
42/6/108
88%
27%
25/3/128
89%
16%
WassRank OT ∗∗
66/28/62
70%
42%
75/28/53
73%
48%
33/3/120
92%
21%
Appendix
Table 20: 5-judge majority-vote generation results (cross-domain, 156 questions) at k∈{5,10,20} . All calibration methods achieve p<0.001 at k=5 ; retain significance at k=10 ( p≤0.021 ) and k=20 ( p≤0.004 ). Random fails at all k values. At k=20 , MMR/OpinionMMR Fair% collapse (29%/33%) as semantic diversity saturates; distributional methods maintain or grow their advantage; WassRank OT rises 70% → 73% → 92%. Fair% = W/(W+L). Win% = W/ N .
k=5
k=10
Approach
Prop. (A)
Cov. (B)
Prop. (A)
Cov. (B)
Stratified (oracle)
35%
—
—
—
W1 Minimizer
34%
38%
30%
27%
DPP
26%
35%
24%
25%
W1 -MMR
23%
36%
18%
17%
OpinionMMR
21%
30%
13%
19%
Appendix
Table 21: Judge prompt ablation: fairness win rate (%) under proportional vs. coverage definitions (3 domains, 156 questions per approach). W1 Minimizer dominates under proportional accuracy; the gap narrows under coverage. Stratified sampling (oracle) confirms the Prompt A criterion tracks the ground-truth distribution, not any specific retrieval strategy — W1 Minimizer matches oracle Fair% within 1 pp at k=5 .
Approach
Info% (A)
Info% (B)
Grnd% (A)
Grnd% (B)
W1 Minimizer
44%
40%
28%
24%
DPP
38%
40%
27%
24%
Random
19%
18%
15%
8%
Appendix
Table 22: Control dimensions remain stable across prompt variants ( k=5 , Δ = Prompt B − Prompt A). Shifts ≤ 4 pp confirm the fairness prompt change does not contaminate other evaluation axes.
Domain
Method
N
W/L/T
Fair%
Win%
Seller Forums
W1 Minimizer
36
10/9/17
53%
28%
KL (JS) Minimizer
36
7/10/19
41%
19%
Yelp
W1 Minimizer
60
39/4/17
91%
65%
KL (JS) Minimizer
60
31/8/21
79%
52%
OpinRank
W1 Minimizer
60
29/8/23
78%
48%
KL (JS) Minimizer
60
35/10/15
78%
58%
Appendix
Table 23: W1 Minimizer vs. KL(JS) Minimizer, generation evaluation at k=5 (156 queries, single-judge Claude Sonnet 4, position-controlled, blind pairwise vs. Top- k ). W1 leads KL(JS) by +7 pp cross-domain; both significantly beat Top- k ( p<0.001 ).