cs.CLOct 7, 2025

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

Authors: Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang

Organizations: University of California-Santa Barbara, USA · Amazon, USA

Abstract

Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.

Explore similar work

CardsList
  1. LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression

    Jul 3, 2026Zhuowen Liu, Longkun Hao, Shiyu Feng +3Large Language Model CompressionSingular Value Decomposition

  2. LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models

    May 30, 2026Haiyu Wang, Yutong Wang, Leshu Li +2Large Language Model CompressionLow-Rank Structure