cs.LGSep 30, 2026

IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs

Authors: Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara

Organizations: AImageLab, University of Modena and Reggio Emilia, Italy

Abstract

We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    May 18, 2026Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-ThiemeUnstructured PruningFeed-Forward

  2. DarwinLM: Evolutionary Structured Pruning of Large Language Models

    Feb 11, 2025Shengkun Tang, Oliver Sieberling, Eldar Kurtic +2Large Language Model CompressionUnstructured Pruning

  3. Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

    Aug 13, 2026Palaash Goel, Ayan Sengupta, Akshay Nambi +1Large Language Model CompressionStructured Pruning