Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-5 to top-30, DEX-Comp at 16× compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by 4.4×--23.7×. Evaluations across additional datasets and backbones further demonstrate its generalization.
Figures & tables
Figure 1: (a) Representative case. Standard distillation reproduces an incorrect answer from the full-context teacher, whereas Pure Distillation (PD) and PD with Hard Exploration (PD+HE) answer correctly. (b) Outcome distribution. Under greedy decoding on a fixed subset of teacher-failed questions, PD reduces the frequency of reproducing teacher errors, while HE further increases correct answers. (c) Sampling performance. On the same questions, PD improves pass@ 1 and pass@ 3 over standard distillation, with further gains from HE.
Figure 2: Overview of DEX-Comp. (a) Soft compression. A compressor turns each document into a shorter sequence of soft embeddings by selecting the final-layer hidden states at compression-token positions. The embeddings can be precomputed offline and consumed by the decoder at inference time. (b) Pure Distillation. On teacher-correct queries ( S+ ), KL distillation aligns the compressed model’s response distribution with that of a frozen RAG teacher reading the full retrieved documents. (c) Hard Exploration. Initialized from Stage I, the compressed model trains on teacher-failed queries ( S− ) using outcome-based GRPO, directly rewarding answer correctness rather than teacher agreement.
Method
Comp. Rate
NQ
TriviaQA
HotpotQA
ASQA
PopQA
Avg.
CEM
LLM
CEM
LLM
CEM
LLM
CEM
LLM
CEM
LLM
CEM
LLM
Top-5
Full context
-
58.65
73.67
90.18
89.04
46.45
55.57
68.99
77.00
58.53
57.36
64.56
70.53
LLMLingua-2
4×
49.74
63.59
84.41
83.19
36.43
45.50
59.49
65.30
39.07
39.70
53.83
59.46
xRAG
128×
37.36
51.46
78.06
78.54
27.66
35.27
43.25
50.63
28.15
29.12
42.90
49.00
ICAE
4×
41.59
56.33
77.78
79.27
28.75
40.48
47.68
60.65
38.28
40.24
46.82
55.39
Table 1: Results with Mistral-7B at top- 5/15/30 . Best and second-best scores at each depth are bold and underlined. † : retrained at the corresponding top- k because matching checkpoints are unavailable.
Table 3: Top- 30 ablation. Dataset scores and Acc are LLM-judged accuracy (%). RL subsets contain teacher-correct (Success), teacher-failed (Failure), or both (Mixed) questions.
Configuration
NQ
TriviaQA
HotpotQA
ASQA
PopQA
Avg.
GPU-hours
Max. RL batch
Full context
73.67
89.04
55.57
77.00
57.36
70.53
–
–
Full context + RL
80.23
92.36
65.00
82.28
64.02
76.78
38
16
PD ( 16× )
71.04
88.25
54.72
74.18
51.41
67.92
1.5
–
PD + HE ( 16× )
76.52
89.97
57.55
79.96
58.70
72.54
13
64
Table 4: Top- 5 LLM-judged accuracy (%) and training cost on eight GPUs. Max. RL batch is the largest per-GPU batch fitting in memory; GPU-hours count total GPU training time.
Figure 3: Resilience Rate (top) and Boost Rate (bottom) across retrieval depths. Subscripts denote the training top- k .
Figure 4: Effects of (a) compression rate and (b) training depth, with the other fixed.
Figure 5: Generalization across (a) backbone families and (b) out-of-distribution datasets.
Figure 7: PISCO compression embeddings learned by standard distillation, for the same passage as Figure 6 . (a) Cosine similarity between embeddings and document tokens. (b) Document tokens appearing among the top- 10 logit-lens vocabulary projections of the memory embeddings. Colors identify embeddings within each model.
Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views. We first show that modern compressors have unstable gains over simple truncation and can add substantial inference-time latency. We then propose Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store. REVA maps token-level attention to readable word units, aggregates importance across repeated document accesses, and renders budget-specific plain-text views that preserve document order and the standard RAG interface. Across four representative benchmarks and modern LLMs, REVA improves generation quality by 1.0--5.8 points over existing advances, while reducing compression overhead by a factor of 5.3 to 15.6, adding less than 40 ms of latency.
Tuan Nguyen, Qiran Hu, Banruo Liu +3
VinUni-Illinois Smart Health Center, VinUniversity, Vietnam · University of Illinois Urbana-Champaign, USA
Retrieval-Augmented Generation (RAG) has become essential for knowledge-intensive question answering, yet scaling RAG pipelines remains challenging due to the prohibitive computational cost of processing lengthy retrieved contexts. Existing compression approaches face a fundamental trade-off: hard compression methods operate online in a query-aware fashion but achieve only modest compression rates and typically require fine-tuning the generative model, while soft compression methods attain higher ratios but rely on costly offline encoding that is entirely agnostic to the input query. To bridge this gap, we introduce RAGOCR, a novel framework that compresses retrieved documents into compact visual representations conditioned on the input query. To further balance compression rate and information fidelity, we introduce a query-aware dynamic resolution mechanism that adaptively allocates visual granularity based on each document's estimated relevance and complexity: highly relevant passages are rendered at higher resolution to preserve fine-grained details, while peripheral documents are aggressively compressed at lower resolution. Experiments on five QA benchmarks using the MedOmniKB retrieval corpus demonstrate that RAGOCR surpasses naive RAG by over 15% in accuracy while requiring only one-eighth the number of input tokens, and consistently outperforms both hard and soft compression baselines across varying retrieval depths.
Jiayang Yu, Jialun Zhong, Lei Zou
Wangxuan Institute of Computer Technology, Peking University Beijing, China
Retrieval-Augmented Generation (RAG) compression papers often evaluate a compressor on one to three readers and treat the compressed evidence layer as evaluation-neutral. We show this assumption is false: fixed compression can raise average accuracy while hiding reader upgrades and reversing model rankings. Across 20 readers and ten domain-method settings over four QA benchmarks and one summarization benchmark, compression gain decreases with reader baseline (nine of ten settings significant, p < 0.05). Generic summarization flips 31% of pairwise model rankings on LongMemEval-S, and a fixed HotpotQA compressor hides 80% of the raw upgrade from Qwen 7B to GPT-4.1-mini. Two opposing forces explain this paradox: compression rescues weak readers by removing noise they cannot filter, and harms strong readers by dropping details they would have used. The pattern appears across structured compilation, generic summarization, three trained compressor families, query-focused summarization, and an external audit of nine published compression papers. We release ragscale, a toolkit built on 177,000 row-level compression transitions, so any compression paper can audit reader scaling with three readers in one day.