cs.CLSep 30, 2026

Sequential Functional Structured Tucker Compression for Large Language Model Attentions

Authors: Jiangfeng Chen, Xinyu Wang, Tianshuo Yan, Hanwei Wu, Xiao-Wen Chang, Yang Zhang, Lei Ding

Organizations: University of Manitoba · McGill University · Simpleway · The University of Hong Kong · McMaster University

Abstract

Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A general tensor-structured compression scheme for efficient large language models

    May 25, 2026Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang +3Large Language Model CompressionEfficient Large Language Model

  2. Compressing Sequences in the Latent Embedding Space: KK-Token Merging for Large Language Models

    Apr 16, 2026Zihao Xu, John Harvill, Ziwei Fan +3Large Language Model CompressionTask-Aware Compression

  3. CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning

    Jul 20, 2026Zhiren Gong, Zihao Zeng, Zijie Wang +3Structured PruningLarge Language Model Compression