cs.LGOct 7, 2026

EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation

Authors: Inbasekaran S

Organizations: SRM Institute of Science and Technology, India

Abstract

Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regularized head pooling respects grouped-query attention while exposing a quantitative trade-off between specialization and worst-head coverage. We derive a mixture-to-head deletion envelope, a computable upper bound on feasible token removal, and a finite-sample observer guarantee that remains valid when the pruning layer is selected adaptively. We then establish a conditional transformer perturbation bound with explicit sufficient Lipschitz constants and a first-token decision-margin corollary. A counterexample shows why shallow observations alone cannot imply an unconditional future-output guarantee. The systems analysis distinguishes query-head unions, physical page allocation, and KV-transfer payload, and gives an arithmetic break-even condition for pruning. This manuscript is theoretical in scope: it defines the procedure, its assumptions, and its formal limits, but does not report measured acceleration or task-accuracy preservation. Experiments are reserved for subsequent validation of the assumptions, approximation tightness, and end-to-end resource trade-offs.

Explore similar work

CardsList
  1. Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

    May 15, 2026Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy +1Model ActivationsTransformer Encoder

  2. A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

    Oct 6, 2026Davis Wertheimer, Haochen Shen, Ahan Gupta +6Large Language Model CompressionKey-Value Cache Compression

  3. Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling

    Apr 20, 2026Yujie Chen, Tailai Chen, Yifeng Gao +4PrefillEfficient Long-Context Inference