cs.CLSep 4, 2026

Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG

Authors: Shuyu Guo, Shuo Zhang, Zhaochun Ren

Organizations: Shandong University Qingdao, China · Bloomberg London, United Kingdom · Leiden University Leiden, The Netherlands

Abstract

Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-55 to top-3030, DEX-Comp at 16×16\times compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by 4.4×4.4\times--23.7×23.7\times. Evaluations across additional datasets and backbones further demonstrate its generalization.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

    Sep 11, 2026Tuan Nguyen, Qiran Hu, Banruo Liu +3Retrieval-Augmented Generation PipelinesInteraction History

  2. RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation

    Aug 1, 2026Jiayang Yu, Jialun Zhong, Lei ZouKnowledge-Based Visual Question AnsweringLingdt-Vl-Ocr

  3. Fixed RAG Compression Collapses Measured Reader Scaling

    Jun 20, 2026Sugam Panthi, Rabab AbdelfattahData Compression MethodsQuestion-Answering Benchmarks