cs.LGSep 28, 2026

Output-aware Residual Stream Pruning for Large Language Models

Authors: Chayne Thrash, Kevin Chen, Soheil Kolouri

Organizations: Department of Computer Science Vanderbilt University Nashville, TN 37235, USA

Abstract

Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Sparsity Whisperer

    Aug 6, 2026Linghao Kong, Inimai Subramanian, Micah Adler +3SparsityStructured Pruning

  2. LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

    Sep 11, 2026Sankar Behera, Dhruv Singh, Anshika Agnihotri +3Large Language Model CompressionUnstructured Pruning

  3. Forward-Free LLM Depth Pruning via Weight Redundancy

    Sep 9, 2026Vincent-Daniel Yun, Woosang LimUnstructured PruningLarge Language Model Compression